Datastreams¶
This page documents practical usage of the datastream classes in httk.core.datastream.
It covers both text and byte data and shows how to write functions that accept flexible input types.
Overview¶
Datastream support is split into two parallel families:
text:
backends:
TextstreamFile,TextstreamFilename,TextstreamString,TextstreamRequest,TextstreamURLviews:
TextstreamFileView,TextstreamFilenameView,TextstreamStringView,TextstreamRequestView,TextstreamURLViewaccepted union:
TextstreamLike
bytes:
backends:
BytestreamFile,BytestreamFilename,BytestreamBytes,BytestreamRequest,BytestreamURLviews:
BytestreamFileView,BytestreamFilenameView,BytestreamBytesView,BytestreamRequestView,BytestreamURLViewaccepted union:
BytestreamLike
In normal user code, you usually accept *Like and normalize immediately to one view.
Textstream¶
Common Calling Patterns¶
import urllib.request
from pathlib import Path
from httk.core.datastream import TextstreamStringView
# filename (str)
TextstreamStringView("README.md")
# filename (Path)
TextstreamStringView(Path("README.md"))
# already-open text file object: mainly useful when other code you do not
# control opened the file for you; when starting from a filename, pass it
# directly instead (as in the first example)
with open("README.md", "r") as f:
TextstreamStringView(f)
# raw string content (disambiguate with kind="content")
TextstreamStringView("line1\nline2\n", kind="content")
# remote content via a urllib request object
TextstreamStringView(urllib.request.Request("https://example.com/data.txt"))
# remote content via a URL string (auto-recognized by its scheme; kind="url" also forces it)
TextstreamStringView("https://example.com/data.txt")
Example: String-Oriented Function¶
from httk.core import TextstreamStringView
from httk.core import TextstreamLike
def header_text(slike: TextstreamLike, **hints: object) -> str:
text = TextstreamStringView(slike, **hints)
return text.splitlines()[0] if text else ""
Use TextstreamStringView when the algorithm naturally wants complete in-memory string data.
Example: Streaming Function¶
from httk.core import TextstreamFileView
from httk.core import TextstreamLike
def count_nonempty_lines(slike: TextstreamLike, **hints: object) -> int:
stream = TextstreamFileView(slike, **hints)
return sum(1 for line in stream if line.strip())
Use TextstreamFileView when line-by-line processing is natural and you do not want eager full materialization.
Textstream Notes¶
TextstreamStringViewis eager: it reads remaining stream content immediately.A bare
strwhose scheme ishttp,https,ftp, orfileis treated as a URL; any otherstrdefaults to filename resolution. Passkind="content"for literal content orkind="filename"/kind="url"to force an interpretation.TextstreamFilenameViewrequires an underlying name; it raisesTypeErrorwhen no filename exists.Remote text is decoded using the
encodinghint if given, else the HTTP Content-Type charset, else utf-8.TextstreamFilenamereads local files as utf-8 by default; passencodingto override.See Remote Content (Request / URL) for shared remote-fetch behavior and Compressed Content for transparent decompression.
Bytestream¶
Common Calling Patterns¶
import urllib.request
from pathlib import Path
from httk.core import BytestreamBytesView
# filename (str)
BytestreamBytesView("payload.bin")
# filename (Path)
BytestreamBytesView(Path("payload.bin"))
# already-open binary file object: mainly useful when other code you do not
# control opened the file for you; when starting from a filename, pass it
# directly instead (as in the first example)
with open("payload.bin", "rb") as f:
BytestreamBytesView(f)
# raw bytes or bytearray
BytestreamBytesView(b"\x00\x01\x02")
BytestreamBytesView(bytearray([0, 1, 2]))
# remote content via a urllib request object
BytestreamBytesView(urllib.request.Request("https://example.com/payload.bin"))
# remote content via a URL string (auto-recognized by its scheme; kind="url" also forces it)
BytestreamBytesView("https://example.com/payload.bin")
Example: Bytes-Oriented Function¶
import hashlib
from httk.core import BytestreamBytesView, BytestreamLike
def digest_payload(blike: BytestreamLike, **hints: object) -> str:
payload = BytestreamBytesView(blike, **hints)
return hashlib.sha256(payload).hexdigest()
Example: Chunked Streaming Function¶
from httk.core import BytestreamFileView, BytestreamLike
def first_chunk(blike: BytestreamLike, size: int = 4096, **hints: object) -> bytes:
stream = BytestreamFileView(blike, **hints)
return stream.read(size)
Bytestream Notes¶
BytestreamBytesViewis eager: it reads remaining stream content immediately.BytestreamFilenameViewrequires an underlying name and raisesTypeErrorif unavailable.A bare
strwhose scheme ishttp,https,ftp, orfileis treated as a URL; any otherstrdefaults to a filename.For explicit interpretation when needed, pass
kind="filename",kind="file",kind="content",kind="request", orkind="url".See Remote Content (Request / URL) for shared remote-fetch behavior and Compressed Content for transparent decompression.
Remote Content (Request / URL)¶
Both families can fetch remote content through Python’s built-in urllib.request:
A
urllib.request.Requestobject is unambiguous and is accepted directly anywhere a*Likeis accepted. Use aRequestwhen you need headers, a method, or a request body; it is passed tourllib.request.urlopenas-is.A URL passed as a plain
stris auto-recognized when its scheme is one ofhttp,https,ftp, orfile. A schemeless string still means a filename (or, withkind="content", literal content). Passkind="url"to force URL interpretation of a string, orkind="filename"to force a scheme’d string to be treated as a filename (for example, so a local file literally named like a URL is not fetched over the network).
Remote backends fetch lazily: the connection is opened on first read, not when the backend or view is
created. Note that unwrap() also opens the connection, since it returns the underlying response object.
An optional timeout hint (in seconds) is forwarded to urlopen.
TextstreamRequestView/TextstreamURLView and their byte counterparts are the URL-facing analogues of
*FilenameView: they present the source location of a backend rather than its data. A *URLView is a
str holding the URL; a *RequestView is a genuine urllib.request.Request (preserving headers/data when
built from a request backend) and can be passed to any code that expects one. Symmetrically to
*FilenameView raising TypeError for backends with no filename, these views raise TypeError for
backends with no underlying URL — and remote backends have no name, so *FilenameView raises for them.
import urllib.request
from httk.core import TextstreamRequestView
req = TextstreamRequestView("https://example.com/data.txt", kind="url")
# req is a urllib.request.Request and can be handed to code that expects one:
with urllib.request.urlopen(req) as resp:
...
Compressed Content¶
Compression is an orthogonal layer below the backends: it turns a compressed byte stream into
an uncompressed one (and, for text, before decoding), independently of where the bytes come from.
The same codecs therefore apply to filenames, open files, raw bytes, and remote responses, so a
data.json.gz filename or a gzipped URL loads with no extra ceremony. The stdlib codecs gzip,
bzip2, xz, and lzma are built in.
A compression hint (parallel to kind) selects how a codec is chosen:
"auto"— use the filename extension if it is a known compression suffix, otherwise sniff the leading magic bytes."detect"— always sniff the magic bytes, ignoring the extension."extension"— decide from the name only; never sniff."none"— no decompression.a registered codec name (e.g.
"gzip") — force that codec; an unknown name raisesValueError.
Defaults depend on the source: filename-based backends default to "extension" (a data.json.gz
name decompresses, a compressed file with a plain name does not unless you ask); all other
byte-producing sources — open files, raw bytes, Request, and URL strings — default to "auto".
Resolution is lazy: like the rest of the stream layer, the extension check or magic sniff only runs
on the first read, and sniffing never consumes data. Compression does not apply to text-native
sources (an already-open text stream or a literal string); for those, only the no-op modes are
accepted and a codec name or "detect" raises ValueError.
import gzip
from httk.core import BytestreamBytesView, DataLoader
# Transparent: extension recognized, decompressed on read.
DataLoader("symmetry", "data/spacegroups.json.gz")
# In-memory gzip is sniffed by default ("auto").
BytestreamBytesView(gzip.compress(b"payload")) # -> b"payload"
# Force or disable decompression explicitly.
BytestreamBytesView("blob.dat", compression="gzip")
BytestreamBytesView("archive.gz", compression="none") # raw compressed bytes
Register additional codecs (for example, a third-party zstd) with register_compression;
known_compressions() lists the registered names.
import io
from httk.core import CompressionCodec, register_compression
register_compression(
CompressionCodec(
name="zstd",
extensions=(".zst",),
magics=(b"\x28\xb5\x2f\xfd",),
open_stream=lambda stream: io.BytesIO(...), # return a decompressed binary stream
)
)
Archives (.tar.*, .zip), write-side compression, and HTTP Content-Encoding negotiation are
out of scope for this layer.
Loading files by type¶
httk.core.load(filename) selects a loader for a file and calls it. Capability
modules register loaders with register_loader, naming the file extensions
and/or exact basenames they handle:
from httk.core import load
from httk.core.register import register_loader, known_extensions, known_filenames
# A stand-in for a real loader (httk-io registers the CIF and POSCAR loaders).
def _demo_loader(filename, **kwargs):
return {"loaded": filename}
register_loader(
name="demo",
loader=_demo_loader,
extensions=(".demo",),
filenames=("DEMOCAR",),
)
assert ".demo" in known_extensions()
assert "democar" in known_filenames() # basenames are stored lower-cased
Dispatch strips at most one recognized compression suffix (.gz, .bz2,
.xz, .lzma) to obtain an inner name, then matches that name’s extension
first and its exact basename second (both case-insensitively). The loader always
receives the original filename, so it can open the still-compressed bytes
through the datastream layer for transparent decompression:
from httk.core import load
from httk.core.register import register_loader
def _demo_loader(filename, **kwargs):
return {"loaded": filename}
register_loader(name="demo", loader=_demo_loader, extensions=(".demo",), filenames=("DEMOCAR",))
# By extension, transparently through a compression suffix:
assert load("/data/sample.demo.bz2") == {"loaded": "/data/sample.demo.bz2"}
# By exact basename (an extension-less file), original path preserved:
assert load("/data/DEMOCAR.gz") == {"loaded": "/data/DEMOCAR.gz"}
An unrecognized file raises a clear ValueError listing the known extensions
and filenames.