Usage#
Basic I/O#
import polyxios as px
# Read any supported format
mesh = px.read("brain.vtk")
# Inspect
print(mesh.vertices.shape) # (n_verts, 3)
print(len(mesh.element_types)) # number of elements
# Write to a different format
px.write(mesh, "brain.ply")
px.write(mesh, "brain.vtp")
Files, buffers and streams#
Anything with a read or a write works where a path does, so a mesh
never has to touch disk:
import io
buf = io.BytesIO()
px.write(mesh, buf, fmt=".ply") # fmt= names the format
buf.seek(0)
same = px.read(buf, fmt=".ply")
with open("brain.vtk", "rb") as fh:
mesh = px.read(fh) # a named handle needs no fmt=
Two rules hold everywhere:
a handle polyxios is given is read or written where it stands, and is never closed - the caller keeps control of its own file;
a buffer with no file name has no extension to infer a format from, so
fmt=is required for one;open()gives a handle a name, and that is enough.
A source has to offer a binary read and a destination a binary
write; everything else a codec reaches for - readline, seek,
seekable - is supplied for a handle that lacks it. A handle that cannot
seek is still read, except by the formats that need to measure a file before
parsing it and by the extensions several formats share, both of which say so
rather than guess.
Reading lazily needs a real file behind the handle wherever the arrays that
come back view the file itself - binary PLY and binary VTK - since mmap
maps a file descriptor: an io.BytesIO raises LazyReadError rather
than quietly loading eagerly. It also needs the handle to stand at the start
of that file - a mapping addresses a file from byte zero - so a handle
part-way into one is refused for a lazy read and read eagerly without
complaint. A format whose lazy mode only skips work, such as binary STL
skipping vertex deduplication, copies what it reads and so takes a buffer
like any other read. TetGen is the one format a buffer cannot carry - a mesh
is a .node and an .ele file found beside each other by name.
Text formats are written with \n line endings on every platform, so the
bytes a path receives and the bytes a buffer receives are the same ones.
Compressed files#
gzip is handled by the same layer, for every format at once:
mesh = px.read("brain.vol.gz") # decompressed on the way in
px.write(mesh, "brain.vtk.gz") # compressed on the way out
Reading looks at the content: a file compressed without being renamed
reads just as well as one ending in .gz, and a .gz name over plain
bytes is read as the plain file it is. The whole four-byte gzip header is
what decides, not the two magic bytes alone - those open one file in every
65536 by chance, and inside a headerless binary format they are an ordinary
coordinate’s low bytes. Writing looks at the name, since
an output file has no content to inspect yet - a destination ending .gz
is compressed, and nothing else is. A nameless buffer has no name to end in
.gz, so fmt= says it there instead:
buf = io.BytesIO()
px.write(mesh, buf, fmt=".obj.gz") # compressed into the buffer
.gz names the compression rather than the format wherever it appears, in
a file name and in fmt= alike, so ".obj.gz" picks the same codec
brain.obj.gz does.
A member does not have to run to the end of what it sits in. Reading stops where the member stops, so a mesh compressed into the middle of an archive is read without the archive’s own bytes after it turning into an error; members written back to back are still read as the one stream they spell.
The compressed output carries no timestamp and no embedded file name, so the
same mesh always produces the same bytes. Lazy reads are the one thing gzip
takes away, in the formats that map: mmap maps a file as it is stored, so
a format whose lazy read hands back arrays viewing the mapping raises
LazyReadError over a compressed file rather than handing back compressed
bytes. A lazy read that copies what it reads - binary STL’s, which skips
vertex deduplication and nothing else - takes a compressed file like any
other read.
TetGen is outside all of this for the same reason it cannot take a buffer: it
opens its own sibling files rather than going through the layer that unwraps
gzip. A .node.gz is refused with a message saying so, rather than read as
text or written uncompressed under a name promising otherwise.
A compressed handle is the one exception to “left where the codec left it”: a decompressor reads ahead, so where it stops says nothing about how much of the mesh was consumed. The handle is put back at the front of the member instead, which is the only position over compressed bytes that means anything to whatever reads next.
Format-specific options#
px.write(mesh, "brain.vtk", binary=True)
px.write(mesh, "brain.ply", binary=True, endian="little")
Every codec’s own options are listed on its page under Supported formats.
Entity numbering#
A PolyData indexes its vertices and elements densely from
zero, always. Most formats agree, but a hand-edited Abaqus deck, a Nastran bulk
data file, a Gmsh .msh, a FLAC3D grid and a Kratos .mdpa all let their
author number entities freely - neither dense nor zero-based, and in any order.
Reading such a file has to renumber, and renumbering on its own is lossy: the
ids are how the author’s other files talk about this mesh, a load case naming
GRID 7000001 or a report keyed on element 4001.
So the reader records what the file said and the writer puts it back:
mesh = px.read("model.bdf")
mesh.vertex_attrs["original_ids"] # the GRID ids the deck spelled
mesh.element_attrs["original_ids"] # the element ids
px.write(mesh, "model.msh") # written back under those numbers
The key is deliberately not format-prefixed: an id says something about the
entity, not about the file it came from, so a mesh read from a .bdf and
written as .msh keeps its numbering. Three rules hold across every codec
that carries ids:
Read renumbers. Connectivity, tags and attributes all speak dense zero-based indices whatever the file spelled.
Read remembers only what the index does not already say. A file numbering
1..nin order records nothing at all - the writer’s own renumbering reproduces it exactly. The key’s absence means “no numbering worth keeping”, not “numbered from one”.Write asks. Stored ids win when they can still be written; otherwise the writer renumbers from one, with a warning.
“Can still be written” is checked at the point of writing rather than trusted,
because a transform moves a mesh out from under its ids without either one
being wrong: merge concatenates two meshes that each numbered from one, so
the ids collide, and splitting a quad leaves two triangles carrying one id.
Either would spell a file with two entities answering to one number, which no
solver loads, so the ids have to be positive, unique and one per entity to
survive.
A format that numbers densely by construction - Medit, OFF, STL, the VTK
family - has no id of its own to record, and ignores the key on write. Whether
it carries one is a separate question, and comes down to whether it has a
general attribute channel: PLY and the VTK family write original_ids back
out as an ordinary data array, so a mesh read from a .bdf, welded, and
written as .vtu keeps the ids of the vertices that survived. Tecplot has
the channel on one side only - its variables are per-vertex, so vertex ids
travel and element ids are dropped with the rest of the element attributes.
Medit, OFF, STL and OBJ have no such channel at all and drop the key with
every other attribute, silently. Where the channel is float - a legacy VTK
data array, a Tecplot variable, a PLY face property - the ids come back as
whole float64 values rather than int64.
Two-dimensional meshes#
A PolyData holds three coordinate columns, always. A format
that spells two - a bamg .mesh, an NDIME= 2 SU2 case, a 2-D MFEM mesh -
is padded with z=0 on the way in rather than kept narrow, so every consumer
can index vertices[:, 2] without first asking what the file said.
Padding alone is lossy in one direction: nothing downstream could tell a plane
written in two dimensions from one written in three that happens to sit at
z=0, so a round trip would widen the file. The reader records the fact
instead:
mesh = px.read("plate.su2") # NDIME= 2
mesh.vertices.shape # (n, 3), the z column all zeros
mesh.global_attrs["was_2d"] # True
px.write(mesh, "plate.mesh") # an MFEM mesh of two columns
Three rules, and every 2-D-capable codec follows them:
Read pads. Vertices are
(n, 3)float64 whatever the file declared.Read remembers. A file that declared two dimensions sets
global_attrs["was_2d"] = True. A three-dimensional file sets nothing: the key’s absence means “not known to be two-dimensional”, not “3-D”.Write asks. The flag decides how many columns go out - while the mesh is still flat. Coordinates outrank it: a third coordinate that reached the mesh after the read is data, so a mesh that has left the plane is written in three dimensions with a warning rather than flattened in silence.
The key is deliberately not format-prefixed the way tecplot_title is: it
says something about the mesh, not about the file it came from. A plane read
from a 2-D Netgen .vol - which has no two-dimensional spelling of its own
to write back - still lands as an NDIME= 2 SU2 case.
Medit, Medit binary, SU2, Netgen, TetGen, MFEM, DOLFIN, Tecplot, Abaqus and WKT all record the flag on the way in; every one of those but Netgen restores it on the way out. A format with no two-dimensional spelling at all - OBJ, the VTK family - keeps writing three columns and ignores the flag.
Two columns sometimes constrain the rest of the file, and the writer keeps it
consistent rather than emitting one no reader loads: an Abaqus deck of
two-column node cards is written under the planar element cards (CPS3,
CPS4, T2D2), since Abaqus takes a node’s dimensionality from the
element referencing it, and a flat mesh of solid cells - a tetrahedron is one
however flat it lies - keeps its third column in every format that reads the
node count per element from a separate number.
Several meshes in one file#
polyxios.read() hands back one PolyData, always. A
function that sometimes returns a sequence puts the branch in every caller, so
a file that holds several meshes is treated as what it is - an index of
sub-files, holding no geometry of its own - and reading one raises
UnsupportedFormatError naming the format.
The several live in polyxios.helper:
from polyxios import helper
whole = helper.read_multiblock("case.pvtu") # every piece, merged
blocks = helper.read_blocks("case.vtm") # one PolyData per sub-file
Both read .vtm, .pvtu, .pvtp, .pvtr, .pvts and .pvti,
as well as a .vtp holding a <vtkMultiBlockDataSet>. An index naming
another index is followed and read flat; a sub-file that is missing or
unreadable is skipped with a line on the polyxios logger, so a partially
downloaded companion directory still loads; and a reference resolving outside
the index file’s own directory raises PermissionError rather than reading
it.
helper.read_blocks is the one to reach for when the blocks mean different
things - a fluid domain and a solid one, one part per block - since
merge() joins them into a mesh that no longer says
which was which.
Where to go next#
Supported formats - the twenty-seven supported formats, one page each
Lazy loading - reading files larger than RAM
Transforms - filtering, cleaning and merging meshes
Command line interface (pxios) - the
pxioscommand linePlugin system - teaching polyxios a new format