Walder#

Walder declares each route in a YAML file shaped like an OpenAPI document: the route carries a query, the sources to run it against, and a view for every response type. Comunica does the querying, sending each pattern to every listed source and joining on the shared variables. Comunica treats a source as a SPARQL endpoint when the source says so with sd:endpoint, or when its URL ends in /sparql, as the two Fuseki snapshots do here.

from helper import call

A simple request#

Looking up a DOI returns the article’s title from OpenCitations Meta. Every route serves its HTML view by default, so the call asks for JSON-LD instead.

call(
    "http://localhost:8089/articles?doi=10.1007/s11192-022-04367-w",
    headers={"Accept": "application/ld+json"},
)
curl -H 'Accept: application/ld+json' 'http://localhost:8089/articles?doi=10.1007/s11192-022-04367-w' 

# 200 OK

[
  {
    "@id": "https://w3id.org/oc/meta/br/061202127149",
    "http://purl.org/dc/terms/title": [
      {
        "@value": "Identifying And Correcting Invalid Citations Due To DOI Errors In Crossref Data"
      }
    ]
  }
]

The join#

This route adds OpenCitations Index to its sources. Comunica runs the one query over both endpoints and joins on the article, so a single response carries the title from Meta and the thirty cited entities from Index. The join happens inside the query: Walder cannot join what two separate queries return.

call(
    "http://localhost:8089/article-references?doi=10.1007/s11192-022-04367-w",
    headers={"Accept": "application/ld+json"},
    max_lines=20,
)
curl -H 'Accept: application/ld+json' 'http://localhost:8089/article-references?doi=10.1007/s11192-022-04367-w' 
# 200 OK

[
  {
    "@id": "https://w3id.org/oc/meta/br/061202127149",
    "http://purl.org/dc/terms/title": [
      {
        "@value": "Identifying And Correcting Invalid Citations Due To DOI Errors In Crossref Data"
      }
    ],
    "http://purl.org/spar/cito/cites": [
      {
        "@id": "https://w3id.org/oc/meta/br/062501777134"
      },
      {
        "@id": "https://w3id.org/oc/meta/br/061302130520"
      },
      {
        "@id": "https://w3id.org/oc/meta/br/062601255589"
      },
      {
        "@id": "https://w3id.org/oc/meta/br/061302130471"
... (82 more lines)

Output#

Walder negotiates HTML, JSON-LD, Turtle, N-Triples, and N-Quads, and the HTML view decides what a browser shows.

A json-ld-frame beside the query shapes the JSON instead: the frame names the terms and keeps the nodes that carry a title, so the citations become a plain list of IRIs under the article. Walder puts the document under the name of the query, here data.

call(
    "http://localhost:8089/articles?doi=10.1007/s11192-022-04367-w",
    headers={"Accept": "text/turtle"},
)
curl -H 'Accept: text/turtle' 'http://localhost:8089/articles?doi=10.1007/s11192-022-04367-w' 

# 200 OK

<https://w3id.org/oc/meta/br/061202127149> <http://purl.org/dc/terms/title> "Identifying And Correcting Invalid Citations Due To DOI Errors In Crossref Data".
call(
    "http://localhost:8089/article-references?doi=10.1007/s11192-022-04367-w",
    headers={"Accept": "application/json"},
    max_lines=24,
)
curl -H 'Accept: application/json' 'http://localhost:8089/article-references?doi=10.1007/s11192-022-04367-w' 
# 200 OK

{
  "data": {
    "@context": {
      "title": "http://purl.org/dc/terms/title",
      "cites": {
        "@id": "http://purl.org/spar/cito/cites",
        "@type": "@id"
      }
    },
    "@graph": [
      {
        "@id": "https://w3id.org/oc/meta/br/061202127149",
        "title": "Identifying And Correcting Invalid Citations Due To DOI Errors In Crossref Data",
        "cites": [
          "https://w3id.org/oc/meta/br/062501777134",
          "https://w3id.org/oc/meta/br/061302130520",
          "https://w3id.org/oc/meta/br/062601255589",
          "https://w3id.org/oc/meta/br/061302130471",
          "https://w3id.org/oc/meta/br/06903303973",
          "https://w3id.org/oc/meta/br/06902330758",
          "https://w3id.org/oc/meta/br/06250648394",
          "https://w3id.org/oc/meta/br/061403569058",
          "https://w3id.org/oc/meta/br/061503593762",
          "https://w3id.org/oc/meta/br/061402111914",
... (25 more lines)

Pagination#

The query behind /references ends with LIMIT ?limit OFFSET ?offset, and Walder fills both variables from the request: it reads page and limit, then multiplies them into the offset. The response carries the window alone, without Link headers and without a total, so a client cannot tell where the list ends.

call(
    "http://localhost:8089/references?doi=10.1007/s11192-022-04367-w&page=1&limit=5",
    headers={"Accept": "application/ld+json"},
)
curl -H 'Accept: application/ld+json' 'http://localhost:8089/references?doi=10.1007/s11192-022-04367-w&page=1&limit=5' 
# 200 OK

[
  {
    "@id": "https://w3id.org/oc/meta/br/061202127149",
    "http://purl.org/spar/cito/cites": [
      {
        "@id": "https://w3id.org/oc/meta/br/061303572746"
      },
      {
        "@id": "https://w3id.org/oc/meta/br/061402111914"
      },
      {
        "@id": "https://w3id.org/oc/meta/br/061402112592"
      },
      {
        "@id": "https://w3id.org/oc/meta/br/061403569058"
      },
      {
        "@id": "https://w3id.org/oc/meta/br/061403572753"
      }
    ]
  }
]

Versioning#

Not supported.

API description#

Walder serves no specification, yet its configuration file is already an OpenAPI 3.0 document: paths, parameters, and responses are the standard ones, while queries, sources, and views sit in x-walder- extensions. The file itself is rendered below.

from pathlib import Path

import yaml
from helper import embed_swagger

spec = yaml.safe_load(Path("walder/config.yaml").read_text())
embed_swagger(spec, base_url="http://localhost:8089/")

Consumer authentication#

Not supported, in either direction: Walder never challenges its own clients, and a source is a bare URL, so it carries no credentials for the endpoint behind it.

Endpoint authentication#

Not supported.

Operations#

GET only.

call("http://localhost:8089/articles?doi=10.1007/s11192-022-04367-w", method="POST", max_lines=0)
call("http://localhost:8089/articles?doi=10.1007/s11192-022-04367-w", method="PUT", max_lines=0)
call("http://localhost:8089/articles?doi=10.1007/s11192-022-04367-w", method="DELETE", max_lines=0)
curl -X POST 'http://localhost:8089/articles?doi=10.1007/s11192-022-04367-w' 

# 404 Not Found


... (10 more lines)
curl -X PUT 'http://localhost:8089/articles?doi=10.1007/s11192-022-04367-w' 

# 404 Not Found


... (10 more lines)
curl -X DELETE 'http://localhost:8089/articles?doi=10.1007/s11192-022-04367-w' 

# 404 Not Found


... (10 more lines)

Caching#

Walder leaves the default caching of Comunica on unless it starts with --no-cache. That cache keeps what Comunica fetches to identify a source, such as the service description of a SPARQL endpoint, and not the answer to a query: Walder reaches OpenCitations Meta through a proxy that records every request it forwards, and two identical calls send the same SPARQL request twice.

from helper import count_endpoint_requests

count_endpoint_requests("http://localhost:8089/articles?doi=10.1007/s11192-022-04367-w", headers={'Accept': 'application/ld+json'})
First call: 200 OK, SPARQL requests that reached the endpoint: 1
Second call: 200 OK, SPARQL requests that reached the endpoint: 1