Changelog

All notable changes to this project are documented here. The format follows Keep a Changelog and the project uses Semantic Versioning.

Unreleased

1.0.0 - 2026-10-02

The first stable release. The core packages' public API is the one frozen in 1.0.0-rc.1, and from now on it follows semantic versioning: no breaking changes before 2.0. New since 1.0.0-rc.1:

  • five add-on packages, released with AvroSharp at the same version: AvroSharp.Confluent for Confluent.Kafka with Confluent Schema Registry, AvroSharp.KafkaFlow, AvroSharp.Azure.SchemaRegistry, and AvroSharp.Aws.Glue with AvroSharp.Aws.Glue.Kafka, all without Apache.Avro;
  • a logo, and the packages' icon;
  • the fixes below.

The performance gate passes on an i7-12800H, an EPYC 7543 and a Ryzen 5 3500U with .NET 8, 9 and 10, apart from two near-ties: single-value varint writes on Zen+ CPUs, the one exception in docs/design.md §11 (#168), and one bzip2 container write, where both libraries compress with the same SharpZipLib.

Added

  • A logo: the packages' icon on NuGet, and the logo in the READMEs and on the documentation site (#220).
  • AvroSharp.Confluent, a new package: Confluent Schema Registry serializers and deserializers for Confluent.Kafka, without Apache.Avro (#185). It is released with AvroSharp at the same version.
    • Serializers: AvroSharpSerializer<T> and AvroSharpDeserializer<T> work with:
      • generated and [AvroSerializable] types, found through AvroTypes, or given as an AvroTypeInfo<T>;
      • the primitives string, int, long, float, double, bool and byte[];
      • generic values through AvroSharpGeneric.
    • Confluent's own code handles the registry: the serializers derive from Confluent's AsyncSerializer/AsyncDeserializer, so subject name strategies, registration, use.latest.version, schema ID strategies (prefix or header), references and domain and encoding rules are Confluent's code.
    • Matches Confluent's Avro serializer:
      • the same message bytes, and a top-level plain bytes value is the message body alone; with a logical type, such as a decimal, it keeps Avro's length prefix, as Java's serializer writes it (#211);
      • the same default subject name strategy (Associated, falling back to Topic);
      • the same configuration keys (avro.serializer.*, avro.deserializer.*). Unknown keys are rejected.
    • Schema evolution: the deserializer reads each message in its writer's schema and resolves it to the type's. A generic deserializer with use.latest.version and no reader schema reads each message as the latest version. A registered schema with an invalid field default or name still reads: the registry is the authority.
    • Checks:
      • With use.latest.version, use.latest.with.metadata or use.schema.id, the serializer checks that the type's schema encodes like the schema whose ID the message carries: the same canonical form, and the same logical types (a decimal's precision and scale, a timestamp's unit).
      • Tombstones follow Confluent: a null value is written with no body, and reading one into a value type throws.
      • The None subject name strategy is rejected with an error that says so, with or without use.schema.id, which Confluent's client can't look up without a subject (#211).
    • Synchronous too: the serializer and deserializer also implement Confluent.Kafka's ISerializer<T> and IDeserializer<T>, so producer.Produce works, and a consumer takes the deserializer without AsSyncOverAsync (#189). Once a topic's schema ID, or a schema ID's reader, is known, the synchronous path runs without a task; with use.latest.version, use.latest.with.metadata, use.schema.id or rules, it waits for the asynchronous one.
    • Builder extensions for Confluent.Kafka's producer and consumer builders. They set the synchronous interfaces, and take a RuleRegistry. SetAvroSharpGenericValueSerializer and SetAvroSharpGenericValueDeserializer set generic serdes (#211).
    • Generic records of several schemas: AvroSharpGeneric.CreateSerializer(registry) writes each record with its own schema, as Confluent's generic serializer does (#193).
    • Configuration: the config classes copy any key-value pairs, as Confluent's do, such as a configuration section's (#195). A cached schema ID is read without a lock (#197).
    • Dependencies and targets: net10.0, net9.0, net8.0 and netstandard2.0, on Confluent.SchemaRegistry [2.14.0, 3.0.0). The tests run against both ends of that range (-p:ConfluentVersion=2.14.0 for the lowest) and against Redpanda.
    • CEL rules name the Avro fields, as with Java's and Confluent's serializers, on every type: generated and [AvroSerializable] types, and generic values (#191). The rules get the value as Apache.Avro's generic model; Apache.Avro comes with Confluent.SchemaRegistry.Rules, which has the CEL executor.
    • Not supported yet: field rules (field-level encryption, CEL_FIELD, #186) and migration rules (#187).
  • AvroSharp.KafkaFlow, a new package: KafkaFlow serializer middleware for Confluent Schema Registry on AvroSharp.Confluent, in place of KafkaFlow's KafkaFlow.Serializer.SchemaRegistry.ConfluentAvro, without Apache.Avro (#154). It is released with AvroSharp at the same version.
    • Setup: AddSchemaRegistryAvroSharpSerializer and AddSchemaRegistryAvroSharpDeserializer on producers and consumers. The registry comes from KafkaFlow's WithSchemaRegistry on the cluster.
    • Matches KafkaFlow's Confluent Avro serializer: the same subjects and message bytes, and each reads what the other writes.
    • Several record types on one topic: AvroSharpMessageTypeResolver picks each message's type from its writer's schema. It matches a list of types by their schemas' record names, without reflection, or, like KafkaFlow's, finds the type by its .NET full name in the loaded assemblies.
    • Not supported: schema IDs in headers, because KafkaFlow's middleware gives serializers no headers.
    • More of Confluent's patterns (#212): a top-level union of records under the topic subject, resolved by each message's union branch; a type's record aliases, so a renamed record still matches; and Protobuf or JSON Schema messages reported as such. Without a list of types, a record no loaded type has is looked for again only after more assemblies load.
    • Finding types by name: AddSchemaRegistryAvroSharpDeserializerByTypeName(), in place of KafkaFlow's AddSchemaRegistryAvroDeserializer(), needs reflection; the overload with the message types doesn't, and rejects an empty list (#212).
    • Dependencies and targets: net10.0, net9.0 and net8.0, on KafkaFlow.SchemaRegistry [4.0.0, 5.0.0). There's no netstandard2.0 build: every KafkaFlow 4.x needs a newer System.Threading.Tasks.Extensions assembly than AvroSharp's netstandard2.0 dependencies bring. Not strong-named, because KafkaFlow isn't.
  • AvroSharp.Azure.SchemaRegistry, a new package: an Azure Schema Registry serializer for Event Hubs and Service Bus messages, in place of Microsoft's Microsoft.Azure.Data.SchemaRegistry.ApacheAvro, without Apache.Avro (#155). It is released with AvroSharp at the same version.
    • AvroSharpSchemaRegistrySerializer: Serialize and Deserialize, sync and async, to and from MessageContent and the types derived from it (EventData, ServiceBusMessage). Its methods are virtual, with a protected constructor, so it can be mocked, as Microsoft's can.
    • Matches Microsoft's serializer: the body is the Avro encoding, and the content type is avro/binary+<schema ID>. The tests compare the body byte for byte, and each reads what the other writes.
    • Types: generated and [AvroSerializable] types, resolved from the writer's schema, and generic records (GenericRecord, AvroValue), read in the writer's schema.
    • Registry: schemas are registered or looked up once, by full name in the serializer's group (AutoRegisterSchemas, off by default as in Microsoft's), and writer schemas are fetched once per ID, even by concurrent readers, and parsed without validating defaults or names (#214). Without auto-registration, a schema the group doesn't have is an InvalidOperationException that says how to fix it.
    • Generic values: records, enums and fixed values, which have the names the registry needs (#214).
    • Dependencies and targets: net10.0, net9.0, net8.0 and netstandard2.0, on Azure.Data.SchemaRegistry [1.2.0, 2.0.0).
  • AvroSharp.Aws.Glue, a new package: a managed AWS Glue Schema Registry serializer for Kafka and Kinesis, on every platform, in place of AWS's native, Linux-only AWS.Glue.SchemaRegistry, without Apache.Avro (#156). It is released with AvroSharp at the same version.
    • AvroSharpGlueSerializer: writes and reads AWS's wire format: 0x03, the compression byte (none, or zlib), the schema version UUID with the most significant bits first, then the Avro data. A test vector from AWS's Java encoder pins the format.
    • Kafka, in AvroSharp.Aws.Glue.Kafka: AvroSharpGlueKafkaSerializer<T> and AvroSharpGlueKafkaDeserializer<T> for Confluent.Kafka, both asynchronous and synchronous, with builder extensions. They're a package of their own, so Kinesis and other users don't take Confluent.Kafka. A null value, or a null AvroValue, is a tombstone.
    • Registry calls, as AWS's serializer makes them: the version is looked up by its definition (Java's Schema.toString() text, as AWS registers it). With auto-registration the serializer registers a new version, or creates the schema with the configured compatibility, and waits while a new version is PENDING (a TimeoutException after 10 checks). Each version ID and each schema is cached. A writer schema with an invalid default or name still reads: the registry is the authority. A version another producer is registering is waited for too, concurrent first writes of a schema make one lookup, and each version the serializer registers gets x-amz-meta-transport and the Metadata option's metadata, as with AWS's serializer. An unknown or deleted version ID in a message is an AvroDataException (#213).
    • Settings: AvroSharpGlueOptions has AWS's settings and defaults: default-registry, schemas named after the topic or stream, auto-registration off, no compression, and BACKWARD compatibility. SchemaNameStrategy gets the transport name, the schema, and whether a key is written (AvroSharpGlueSchemaNameContext), as AWS's naming strategy does.
    • Types: generated and [AvroSerializable] types, resolved from the writer's schema, and generic records.
    • Dependencies and targets: net10.0, net9.0, net8.0 and netstandard2.0, on AWSSDK.Glue [4.0.0, 5.0.0), and AWSSDK.Core [4.0.3.3, 5.0.0) (below 4.0.3.3 it has an advisory). AvroSharp.Aws.Glue.Kafka adds Confluent.Kafka [2.0.2, 3.0.0).
    • Tests: an in-memory Glue client, and moto (an AWS emulator in Docker) through the AWS SDK's real client. Not yet against AWS's own package, whose native library didn't start in the Docker environment tried.
  • The add-ons on .NET Framework and Native AOT (#198): the Confluent, Azure and Glue tests also run on .NET Framework 4.8.1, and CI publishes an application using AvroSharp.Confluent, AvroSharp.Azure.SchemaRegistry, AvroSharp.Aws.Glue and AvroSharp.Aws.Glue.Kafka with Native AOT, failing on any trim or AOT warning in an AvroSharp assembly.

Fixed

  • Schema JSON that ends in a backslash (an unfinished escape) is a schema error. The check for escaped unpaired surrogates, which runs before the JSON parser, threw ArgumentOutOfRangeException for it, from AvroSchema.Parse and from a container file whose header held such a schema. Found by the nightly fuzzing (SchemaParse and ContainerFile).
  • AvroTypes.TryGet(Type) and AvroTypes.Get<T>() find a generated type whose assembly no code has run from yet, such as one loaded only for its metadata, by a type name or a message handler's signature. They run the assembly's module initializer, which registers its types, before they report the type unknown.

1.0.0-rc.1 - 2026-10-01

The release candidate for 1.0: the public API is frozen, and 1.0.0 follows when .NET 11 is released. New since 0.2.0:

  • schemas and serializers from your own C# types ([AvroSerializable]), and lookup by type without reflection (AvroTypes);
  • schema compatibility checks with every reason, in code and as avrosharp schema compat;
  • container files whose schema only Java accepted are now read;
  • the fixes from three reviews: an independent one, Apache.Avro's C# test cases, and Apache Avro Java's behavior.

The performance gate passes on an EPYC 7543 with .NET 8, 9 and 10: 330 of 330 comparisons are faster than Apache.Avro and allocate no more.

Breaking

The public API review before 1.0 (#134) renames and moves members, so the API can be frozen without them (#73). Regenerate checked-in generated code with the new avrosharp gen: code from 0.2 calls support members that moved. From 1.0, code generated by a version keeps compiling against later runtimes of the same major version (the policy).

  • Generated-code support: AvroGeneratedCode, AvroRecordPlan, AvroPlanCache, AvroConversion and AvroUninitialized move to AvroSharp.Serialization.Generated.
    • IAvroCodec<T> is now IAvroValueSerializer<T>, the primitive codecs Avro…Serializer, and the generated nested AvroCodec struct ValueSerializer: "codec" means block compression only.
    • Removed, because generated code no longer calls them: GetRecordPlan(writer, reader) without the cache, PutTypeMismatch(object?, string, string), AvroRecordPlan.Target(int) and Conversion(int), and IsSameSchema (use AvroSchema.HasSameCanonicalForm).
  • Exceptions: AvroDataException (was AvroSharp.IO) and AvroSchemaException (was AvroSharp.Schemas) are in AvroSharp.
  • Schema lookup: IAvroSchemaStore is now IAvroSchemaResolver, with GetSchemaAsync, which an implementer has to add. The AvroMessageReader factories name its parameter resolver, and AvroMessageReader.CreateGeneric names its reader options readerOptions, as the other reader factories do.
  • AvroSerializer: Serialize(output, value) and TrySerialize(destination, value, out bytesWritten) take the output first.
  • Renamed:
    • GenericDatumReader.Schema and GenericDatumJsonReader.Schema are now WriterSchema;
    • EnumSchema.Default is now DefaultSymbol;
    • AvroReader.Skip(long) is now SkipRaw;
    • AvroValue.FromByteArray, FromGenericRecord and FromGenericFixed are now FromBytes, FromRecord and FromFixed;
    • AvroFileReader<T>.GetMetadataString is now TryGetMetadataString(key, out value);
    • GenericRecord.TryGetValue(int index, out AvroValue value) names its first parameter position, as the rest of GenericRecord does;
    • CodeGenOptions.DefaultNamespace, NamespaceMapping and PropertyNaming are now Namespace, NamespaceMap and PropertyNames.
  • Changed types:
    • AvroFileReader<T>.Codec is the AvroCodec (its name is Codec.Name);
    • AvroFileWriterOptions.Metadata holds ReadOnlyMemory<byte> values;
    • SchemaFingerprint.Crc64AvroEmpty is a long;
    • ConfluentSchemaIdHeader.Encode takes an AvroSchemaId, and rejects a numeric ID: the header holds a GUID.
  • AvroCodec.CreateDeflate is replaced by new DeflateCodec(level).
  • AvroCodecNames moves to AvroSharp.Containers, next to AvroCodec (#73).
  • uuid strings with spaces around them are rejected (AvroDataException): the reader takes the RFC 4122 form only, as Java's UUID.fromString does. Before, Guid.TryParseExact trimmed them (#135).
  • Validation: the generic and schema parse options reject a MaxDepth below 1 and a negative MaxZeroSizeItems. CodeGenOptions rejects nullable annotations with a LanguageVersion below 8, and any version below 7.

Added

  • The public API is frozen for 1.0 (#73). Every package declares its API as shipped, and package validation compares each with its 0.2.0 on nuget.org. The breaks listed above are the only ones allowed; any other fails the pack.
  • AvroMessageReader<T>.ReadAsync, which fetches an unknown fingerprint through the resolver and keeps its reader, and AvroSchemaStore.GetSchemaAsync (#134).
  • Schemas and serializers from C# types (#31): mark a partial class or record class [AvroSerializable], and the AvroSharp.Generators package writes its schema and serializers from its members.
    • What the type gets: the same members as a type generated from a .avsc file, so AvroSerializer, container files, messages and the registry readers take it alike. The serializers are the same code, with no reflection, and they're AOT-clean.
    • The mapping:
      • primitives;
      • nullable types, as unions with null and a default of null;
      • C# enums, List<T> and Dictionary<string, T>;
      • other [AvroSerializable] classes, as records;
      • Guid, DateOnly, TimeOnly and DateTimeOffset;
      • decimal with [AvroDecimal], and byte[] or Guid as fixed with [AvroFixed];
      • object with [AvroUnion], as a union of records.
    • Field names are the member names as written, as Apache Avro's do. AvroNaming.CamelCase converts them, per type or with [assembly: AvroSerializableDefaults(FieldNames = ...)]. [AvroName], [AvroAlias], [AvroDoc] (or the XML summary), [AvroDefault], [AvroIgnore] and [AvroFieldPosition] cover the rest. [AvroLogicalType] on a raw long, int or string keeps the value raw.
    • Diagnostics AVROGEN101–118 report what the first version doesn't support, at the code: init-only members, primary constructors, nested or generic types, a DateTime without a logical type, and enums whose values aren't 0, 1, 2 and so on.
    • The design is in docs/design.md §6.5, the guide in Code generation, and the code in the SerializableTypes sample.
  • AvroTypes (#31): a type's schema and read and write functions, by type argument or by Type, without reflection, for integrations and generic code on every target.
    • Every generated type has AvroTypeInfo. It registers itself on .NET 5 and later (with a module initializer); elsewhere, call AvroTypes.Register once.
    • The primitives that Confluent's serializers support are registered, and read promoted writer schemas.
  • Schema compatibility checks (#165): AvroSchemaCompatibility.Check(writerSchema, readerSchema) says whether data written with one schema can be read with another.
    • The verdict is Compatible, Partial or Incompatible. Partial means some values can't be read: an enum symbol, or a writer union branch, that the reader lacks. The readers leave these until such a value is read. The verdict is a fact about the schemas, which the options don't change. IsCompatible says whether the check passes: only Compatible passes by default, as in Java's SchemaCompatibility and schema registries; AllowPartial accepts Partial.
    • Every incompatibility is reported, not just the first (Incompatibilities, with the warnings in Warnings). Each has a kind, a message, and a path into the data ($.items[].sku, $.payment[1:Card].number). A named type used in several places lists its other paths.
    • The kinds are Java's six, plus InvalidDefault.
    • The check uses the resolving readers' own rules, and the tests check every row against GenericDatumReader.Create and the generated-code plans.
    • Warnings cover differences that the specification allows but that can change the values read: decimal scale or precision, other logical type changes, lossy promotions, names matched without their namespace, enum defaults used for unknown symbols, and ambiguous field aliases. WarningsAsErrors makes them fail the check.
    • Schema registry levels: CheckVersions(schema, previousVersions, level) checks a new version at Confluent Schema Registry's levels (Backward, Forward, Full and their transitive versions), with a result for each pair; each pair's Index is the earlier version's position in the list.
    • The CLI: avrosharp schema compat <writer> <reader>, or --level backward-transitive v1 v2 v3. It has text or --json output (with a formatVersion), --warnings-as-errors (or --strict) and --allow-partial. Exit code 3 means partially compatible, and 4 incompatible or, with --warnings-as-errors, warnings.
    • Documentation: a section in Schema evolution, and "Replacing Schema.CanRead" in the migration guide, with a table of where Apache.Avro's verdicts differ.
  • AvroSchema.HasSameCanonicalForm: whether two schemas have the same encoding. Equals stays reference equality, as AvroSchema's docs now say (#134).
  • DeflateCodec: the built-in codec is public, with Default and Level, as the Codecs package's codecs have (#134). It rejects a level the runtime doesn't have when it is made, rather than when a writer compresses its first block.
  • New overloads:
    • AvroSerializer.Deserialize<T>(in ReadOnlySequence<byte>, AvroSchema writerSchema);
    • GenericDatumReader.Read(in ReadOnlySequence<byte>);
    • AvroMessage.Write(output, in AvroValue, GenericDatumWriter);
    • AvroRegistryMessage.Write(output, framing, id, in AvroValue, GenericDatumWriter) (#134).
  • Default limits as constants: GenericDatumReaderOptions.DefaultMaxDepth and DefaultMaxZeroSizeItems, GenericDatumWriterOptions.DefaultMaxDepth and AvroSchemaParseOptions.DefaultMaxDepth (#134).
  • Generator settings (#134):
    • The AvroSharpNamespaceMap MSBuild property, avro.ns:CSharp.Ns entries separated by ; or ,, like the CLI's --namespace-map. The package's build/AvroSharp.Generators.targets passes it on with , for ; and line breaks, since the compiler reads it from an .editorconfig file, where ; starts a comment and a line break ends the value. So the entries can also be written one per line, or given with -p:.
    • Generator properties accept their values in any case.
    • A value the generator doesn't recognize is warning AVROGEN006; before, a typo was ignored silently.
    • --language-version 7 implies --no-nullable.
  • Records built from another record's fields: a record given a field that belongs to another record takes a copy, so new RecordSchema(name, other.Fields.Append(field)) works. It threw before (#134).
  • Documentation:
    • Migrating from Apache.Avro: the plan, and the Apache.Avro API next to AvroSharp's. Its code is the Migration sample's, which CI runs against both libraries, and a test keeps the two in step (#21).
    • The GeneratorPackage sample: an application's project with the AvroSharp.Generators package, its MSBuild settings and schemas across files, and how to use the generator from source instead (#126, #74).
    • A Getting started section: installing, a first program, generated types, container files, schema evolution, logical types and JSON. Each page's code is a sample's (the new ContainerFiles, SchemaEvolution, LogicalTypes and Json samples), and a test keeps them in step.
    • Links from the guides and READMEs to the API reference, an API map, namespace overviews, and cross-references between the API pages.
  • A dev container (.devcontainer/) and build/ci-local.sh, which build and test as the Linux CI job does. The container has Ubuntu 24.04 with the .NET 8, 9 and 10 SDKs, the Native AOT prerequisites, CI's variables and a cached NuGet volume; the script runs the job's steps in order. DocFX and ReportGenerator are pinned local tools (.config/dotnet-tools.json), used by CI, the docs workflow and the container alike (#139).

Fixed

  • Container files whose schema Java reads are read. A file's schema is parsed as Java parses it, without validating names or defaults, and with Java's leniencies: a "doc" or an enum's "default" of null, and a namespace that is not a string. Before, a field named user-id or größe, or a "default": null on a string field, made the file unreadable (The file's schema is invalid). These are common in files written by Java before 1.9, and by other tools. AvroSchema.Parse still validates.
  • Out-of-range logical values in generated types: a timestamp-millis of Long.MaxValue, a common "never" sentinel that Java reads, can't be a DateTimeOffset.
    • "avrosharp.logicalType": "raw" on a schema now keeps that one logical type's underlying type, so the property is a long and reads it. The other fields keep their .NET types. "native" maps one schema to its .NET type under AvroSharpLogicalTypes=raw.
    • The range errors say how to read such values.
    • Not available with AvroSharpApacheCompatible, which reports it.
  • Documented: logical types don't take part in schema resolution, as in Java, so a decimal of another scale reads as another number, although the specification says such decimals don't match. AvroSchemaCompatibility warns about it.
  • Message readers (#159, #160):
    • Concurrent ReadAsync calls that met the same new fingerprint or registry ID each asked the resolver. Now one fetches and the others wait for it. A fetch that fails or is cancelled caches nothing.
    • Bad data made ReadAsync and ReadPayloadAsync throw from the call itself. It now faults the returned task, so reads started together, then awaited, all complete.
  • Code generation (#161): a type whose C# name would hide a namespace or type the generated code uses is an error. For example, a namespace map onto AvroSharp with a record Serialization used to generate code that didn't compile.
  • Generated types (#162): a stored default of more than 65,536 zero-size items (an array of nulls, say) failed every read of older data. Defaults now decode without that limit, through the new AvroRecordPlan.DefaultReader.
  • Enum and fixed values written as another schema: the generic binary and JSON writers wrote an enum value's ordinal in its own schema, so a symbol at another position in the target schema became another symbol, silently, and one the target lacks was written anyway; and a fixed value of another size wrote all its bytes, shifting every later field. The symbol is now found in the target schema, and the size is checked. Values of the target schema itself, or of a copy parsed separately, write faster than before.
  • Escaped unpaired surrogates (\ud800) in a schema or in JSON data threw InvalidOperationException. A schema with one, anywhere, is now an AvroSchemaException (container file headers included), and JSON data an AvroDataException.
  • .NET Framework: the JSON reader read -0.0 as +0.0 and rejected a number beyond double's range; it reads -0.0 and an infinity, as on the other targets and in Java.
  • A bad namespace is reported at $.namespace, not $.name.
  • A fixed type named Equals or GetHashCode generated code that did not compile (CS0542); it is renamed, with a note, like the other generated member names (#141).
  • A decimal default that the generated C# decimal cannot hold (beyond 96 bits, or with more digits than the precision) generated code whose constructor threw; it is now a generation error (#141).
  • Union defaults in generated code are for the first branch they are a value of, as Avro 1.12 says and the readers do, not always the first branch. ["int","string"] with a string default failed to generate, a [enum,"string"] default that isn't a symbol generated code that didn't compile, and a decimal default in a later branch escaped the check above.
  • C# 7.0 and 7.1 projects: the generator read C# 7.0 as version 0 and accepted 7.1, whose generated code doesn't compile (readonly structs need C# 7.2). Both are now error AVROGEN003, which says to set <LangVersion> to 7.3 or later. CodeGenOptions.LanguageVersion 7 means C# 7.2 or later.
  • A default that isn't a value of its field's schema, on a field built in code (the parser checks parsed ones), made GenericDatumReader.Create and the generated-code plan throw AvroDataException. It is now the documented AvroSchemaException, naming the field.

Changed

  • A coverage badge in the README.
    • CI's merged line coverage is written by build/coverage-summary.sh in shields.io's endpoint format.
    • The docs workflow publishes it with the site. It now deploys after CI passes on main, not on every push.
  • JSON numbers as Java reads them: the JSON reader accepts a whole number written as 1.0 or 1e2 for an int or long, and "INF" and "-INF" for a float or double, as Java's JsonDecoder does. The writer is unchanged.
  • Faster logical values, parsing, resolution and strings (#135). Measured on the EPYC 7543 with main and the branch in parallel lanes on neighbouring CCDs, on both sockets:
    • Decimal writes: 2.9× faster on bytes (178–188 to 62 µs per 1,024 decimal(18,4) values; Apache.Avro takes 94 µs), and 2.6× on fixed. Pow10's multiplication loop became a table, decimal.GetBits writes into the stack on .NET 5+, and the rounding check runs only for a value with more fractional digits than the scale (.NET 7+).
    • uuid strings: writes −40%, reads −50%. They're parsed from and formatted to UTF-8 (Utf8Parser/Utf8Formatter), without a string, and are faster than Apache.Avro both ways.
    • Schema parsing: 9–13% faster, with 33–37% less allocation, for small and large schemas, with or without the fingerprint. There's no closure per named type, field, union and alias, a logical type is read without a dictionary, and the canonical form's string is made only when asked for.
    • The resolving generic reader: 9–10% faster; with record, array, map and bytes defaults, 30–37% faster. Each default is converted from JSON once: immutable ones are shared, mutable ones decoded from their encoding per read.
    • String writes: 16% faster (the Strings workload). The length prefix is reserved for 1 byte per char, so ASCII strings of 22–63 chars aren't moved.
    • A LogicalTypeBenchmarks and a resolution benchmark with defaults were added to measure these.
  • Tests for the remaining gaps of the test review (#142): enum and fixed aliases, the resolving reader's options, the JSON writer's widening, depth limit and shape errors, maps of any IReadOnlyDictionary, hostile container files (overlong varints, truncation on the async path, trailing bytes when pipelined), AvroRegistryMessageReader.ReadAsync with a missing schema or a cancelled fetch, logical-value range errors, and control characters in schemas.
  • The nightly fuzzing workflow also runs a random-schema code-generation test (#141): random schemas, with hostile names and every kind of default, are generated, compiled for C# 7.3, 12 and the latest version, and round-tripped. PR CI runs it on 100 schemas.
  • The performance gate checks every benchmark (#157). LogicalTypeBenchmarks was not gated: its methods weren't named AvroSharp_*, and only one of its categories had an Apache.Avro baseline.
    • Each logical type and operation is now a group with an Apache.Avro baseline, adding decimal reads, decimal on fixed and timestamp reads. uuid on fixed has no Apache.Avro equivalent.
    • ResolutionBenchmarks gains an Apache.Avro baseline for reading with defaults.
    • The gate fails a benchmark it can't compare, instead of skipping it. Benchmarks without an Apache.Avro equivalent are marked Ungated.
  • The API reference on the documentation site is built from the net10.0 build, so it shows the .NET 8+ API (AvroSerializer, IAvroSerializable<T> and the overloads that take no delegates), and lists those members with their equivalents on other targets (#136).

0.2.0 - 2026-09-29

The avrosharp command-line tool; generated types that read older versions of their schema, write and read memory without allocating, and start with their schema defaults; and the fixes from an independent review: bounded memory and nesting for hostile input, schema resolution that follows the specification and Java, and code generation fixes. The source generator needs the .NET 10 SDK or Visual Studio 2026 and later; the generated code and the tool run on .NET 8 and later (and the code on .NET Standard 2.0).

Added

  • The avrosharp command-line tool, package AvroSharp.Tool: a dotnet tool like Apache.Avro's avrogen, for .NET 8 and later (#33).
    • avrosharp gen writes the source generator's code for .avsc files, or folders of them, with the generator's options: one file per type, in folders for its namespace, or with --flat in one folder.
    • avrosharp schema canonical and schema fingerprint print the Parsing Canonical Form, and the CRC-64-AVRO, MD5 or SHA-256 fingerprint as hex, base64 or (CRC-64) the decimal Java prints.
    • Files may refer to each other's types in any order. Errors are in the compiler's format, with the generator's diagnostic IDs. The exit code is 0 on success, 1 when the command fails and 2 for an invalid command line.
  • CodeGenOptions.NamespaceMapping (--namespace-map in the tool): C# namespaces for Avro namespaces and the namespaces under them, as avrogen's --namespace a:b. Only the generated C# namespace changes, not the schema. GeneratedSource.Namespace gives each file's C# namespace.
  • SchemaFileSet in AvroSharp.CodeGen: parses schema files that refer to each other's named types, in any order, as the source generator and the tool do.
  • Bulk booleans (#26): AvroReader.ReadBooleans checks and copies a block of booleans as one span, and AvroWriter.WriteBooleans writes one as a single copy. Generated code uses them for boolean arrays, and the generic reader and writer use them for boolean arrays.
  • Generated records implement new interfaces (#115): IAvroWritable (WriteTo(ref AvroWriter)) and IAvroReadable (ReadFrom(ref AvroReader), which fills an existing instance and reuses its lists, dictionaries and records), and on .NET 8 and later IAvroSerializable<T> with the static Schema, Write and Read members. With it, AvroSerializer.Serialize/TrySerialize/Deserialize, AvroFileWriter.Create<T>(stream), AvroFileReader.Open<T>(stream)/OpenAsync<T>, AvroStreamWriter.Create<T>, AvroStreamReader.Open<T>, AvroMessage.ToArray<T>/Write<T>, AvroMessageReader.Create<T>, AvroRegistryMessage.ToArray<T>/Write<T> and AvroRegistryMessageReader.Create<T> need no delegates. Readers of files and messages resolve other versions of the type's schema.
  • Generated records write and read memory without allocating (#115): TryWriteAvroBytes(Span<byte>, out int) (no exception when the value does not fit), WriteAvroBytes(Span<byte>), WriteAvroBytes(IBufferWriter<byte>), FromAvroBytes(data, out int bytesConsumed) and FromAvroBytes(in ReadOnlySequence<byte>).
  • AvroWriter.WriteInts and WriteLongs write array items in bulk.
  • CodeGenOptions.LanguageVersion: the C# version generated code may use. The source generator sets it from the project; with C# 11 or later, generated code adds IAvroSerializable<T> for .NET 8 and later.
  • Generated code is easier to debug and read (#117): records have a [DebuggerDisplay] with their first three simple fields, the serializers are [DebuggerNonUserCode], and a property whose logical type keeps its raw type (a decimal wider than 28 digits, duration, timestamp-nanos, or any logical type with AvroSharpLogicalTypes=raw) says in its documentation what the value means.
  • Fixed types have ==, != and AsSpan() (#117).
  • The generator reports a property renamed to avoid a clash (for example user_id and userId in one record) as informational diagnostic AVROGEN005, naming both fields; GeneratedSource.Notes carries these notes for other callers of CSharpCodeGenerator (#117).
  • Generated files suppress CS8981, so all-lower-case Avro type names compile without warnings (#117).
  • AvroFileReaderOptions.MaxSchemaLength and MaxZeroSizeValuesPerBlock (#129, #130).
  • GenericDatumJsonReader.ReadDefault: converts a schema default to an AvroValue, the value a reader gives a field that the data lacks (#131).

Changed

  • Packages list only the dependencies they use (#20). AvroSharp.Generators and AvroSharp.CodeGen depend only on AvroSharp, and AvroSharp.Codecs only on AvroSharp and its compression libraries; before, they also listed System.Memory, System.Text.Json and other packages that AvroSharp brings. AvroSharp's netstandard2.1 build no longer depends on System.Memory or Microsoft.Bcl.AsyncInterfaces, which that target has built in. The generator still needs the .NET 10 SDK or Visual Studio 2026 and later; docs/design.md records what fails on older SDKs.
  • Faster varint encode (#102). Where PDEP is fast (Intel from Haswell, AMD from Zen 3), 3- to 8-byte values are one inline 8-byte store; elsewhere one out-of-line call writes 3 to 5 bytes with one 4-byte store. On the i7-12800H, 3- to 8-byte encodes are 28–38% faster and 3-byte encode is 1.52× faster than Apache.Avro (1.09× before); on the EPYC 7543 they are 25–35% faster, and on an i5-3570K and a Ryzen 5 3500U without fast PDEP, 3 to 5 bytes are 7–33% faster. 1-byte encode is 6–8% slower on the fast-PDEP machines. Record writes are unchanged or faster; generated writes are 7–8% faster on the i7 and the EPYC.
  • AvroWriter.WriteLongs/WriteInts keep the write position in a local while the buffer has room, instead of storing and reloading it after every value (#27, the per-value round trip through memory that limited one-at-a-time writes on the EPYC). They are at least as fast as single writes for every length on the machines measured, and 1-byte bulk writes are 19–48% faster than single ones.
  • Bulk ReadLongs/ReadInts decode runs of one-byte values 16 at a time with vectors, without a loop over the run (#24). Mixed data (90% small values) is read 1.48× faster than the plain loop on the i7 and the EPYC, and small values 2.7–3.6× faster; timestamps stay within 3% of the plain loop, as #29's rule requires.
  • Schema parsing uses a pooled JsonDocument and copies out only what a schema keeps (defaults and custom properties) (#103). Parsing a small schema is 2.07× faster than Apache.Avro on the i7 (1.08× before) and allocates 7.2 KB instead of 8.4 KB. Content after the schema JSON is now reported by the JSON parser, with its line and column.
  • The generic reader stores arrays of boolean, int, long, float and double items as the primitives themselves, not one AvroValue each (#23). AsArray() still returns the items as values. The new TryGetInt64Array and the other typed accessors return the memory without a copy, and FromInt64Array and the other typed factories wrap existing memory. The writer takes these arrays from their memory: booleans, floats and doubles as one copy, ints and longs with no kind check per item. On the i7-12800H (.NET 10), reading an array of 1,000 items with the generic model is 4.4× faster for ints (3,241 to 734 ns), 3.5× for longs and 6.2× for doubles, and allocates a quarter to a half as much. A record with 64 long counters reads in 591 ns instead of 748 ns. Against Apache.Avro, array reads are now 6.6–16.6× faster, up from 1.7–2.8×.
  • The array returned by AsArray() for these item types is no longer a List<AvroValue>. Code that cast it to List<AvroValue> has to copy it instead (AsArray() has always been documented as IReadOnlyList<AvroValue>).
  • Generated code is smaller (#114). Nullable unions of a primitive or a record, and arrays and maps of primitives, strings, bytes or records, are one call to an AvroGeneratedCode helper instead of an inlined switch or loop. The union helpers are inlined; the collection helpers take struct codecs (IAvroCodec<T>) that the JIT specializes. The schema-evolution reader is one ReadField switch instead of a case and a helper method per field. For a production-style schema of 12 records, generated code went from 18,653 to 9,193 lines and its assembly from 377 KB to 165 KB; the 31-field test Order record went from 2,200 to 1,038 lines.
  • Generated records with more than 32 fields are serialized by several methods of up to 32 fields each, which keeps each within the JIT's limit on tracked locals (#113). On a 140-field record (WideRecordBenchmarks, i7-12800H, .NET 10), a generated read takes 525 ns, against 661 ns with one method per record and 729 ns before this change; a write takes 374 ns, against 419 and 492 ns.
  • Generated readers no longer allocate each collection twice: they construct records with a constructor that skips the property initializers (#113). Block counts are checked against the input without a 64-bit division, the plan for another writer schema is cached per generated type, int and long arrays are written with the new AvroWriter.WriteInts/WriteLongs, and date converts through DateOnly.DayNumber. A generated read of the benchmark Order record allocates 2,048 B instead of 2,192 B, with times unchanged within noise; reading another version of a schema takes 179 ns (181 to 193 ns before).
  • new T() gives every field with a schema default its default (#112): primitives, strings, bytes, enums, nullable unions whose default is for their first branch, and arrays and maps of those. Before, only reading data that lacked the field applied the default. Code that relied on these fields starting at zero or empty sees the defaults instead.
  • The Apache compatibility mode makes code written for avrogen classes compile unchanged (#116): property names default to the Avro field names, as avrogen's are (AvroSharpPropertyNames=pascal keeps PascalCase), and types have avrogen's static _SCHEMA and an instance Schema (Apache's Avro.Schema). AvroSharp's schema is AvroSharpSchema in that mode, on records as on fixed types. CodeGenOptions.PropertyNaming is now nullable: null means the mode's default.
  • Generated property names title-case segments in capitals (#117): USER_ID becomes UserId (it was USERID) and HTTP2_PORT becomes Http2Port; segments with lower-case letters keep their capitals (txId is TxId). This renames such properties in existing generated code.
  • [GeneratedCode] on generated types carries the generator's package version (for example 0.1.2) instead of 0.0.0.0 (#107): MinVer sets AssemblyVersion to major.0.0.0.
  • Generated code no longer checks a union branch's value for null after a type pattern proved it is not, and null checks have one pair of parentheses (#107).

Fixed

  • Generated code: a line terminator other than \n in a schema's doc text (\r, U+0085, U+2028 or U+2029) ended the generated /// comment, so the rest of the doc was compiled as C#. For a field's doc, that could add members to the generated type. Doc text is now split on every C# line terminator, and other control characters are dropped (#110).
  • Generated readers limited zero-size array items (nulls, empty records) per array, not per value. Arrays nested in arrays multiplied the limit: an 8 KB input could declare about 131 million items and allocate about 2 GB. The limit of 65,536 now covers the whole value, as in the generic reader. Each generated Read starts a new budget, so values read one after another from the same reader, as in a container block, each get the full limit (#109).
  • The generator reports a union whose branches map to the same C# type, for example a uuid string and a uuid fixed (both Guid), as error AVROGEN003. Before, it generated code that didn't compile (CS8120), and whose writer couldn't have chosen the branch anyway. The message suggests AvroSharpLogicalTypes=raw (#108).
  • Generated writers check enum values (#111): a C# enum can hold any number, which was written as an out-of-range ordinal that readers reject. Writing it now throws AvroException naming the field.
  • A bytes decimal with no bytes is rejected with AvroDataException instead of being read as 0, as in Java (#111). The Apache compatibility mode's decimals check the same.
  • Put on a generated record's null-typed field rejects values other than null, which were silently dropped on write (#111).
  • A generated type's Schema is one instance even when first read on several threads at once. Before, a thread that lost the race to parse it could keep its own instance, which missed the cache of resolution plans and the same-schema fast path. The new AvroGeneratedCode.PublishSchema stores the first one.
  • Hostile input (#129):
    • A writer schema that nested arrays, maps or unions between recursive records overflowed the stack, which ends the process: MaxDepth counted only records. Values may now be nested 8 × MaxDepth levels deep (1,024 by default), counting arrays, maps and records, on every path that reads by the writer's schema: the generic and resolving readers, skipped fields, and the transcoder generated types use. The thread's stack is also checked every 16 levels of new depth.
    • A record of many null fields takes no bytes but creates a value per field, so 3 bytes could allocate 160 MB. MaxZeroSizeItems now counts each zero-size item as the values reading it creates: one, plus one per field of each record in it.
    • The bzip2 codec let IndexOutOfRangeException escape on corrupt blocks; bzip2, xz, zstandard and snappy now report every corrupt block as InvalidDataException. A snappy length of 2^31 or more is rejected instead of ending in ArgumentOutOfRangeException.
    • The resolving-reader cache kept every writer schema alive for as long as its reader schema lived, so each file opened with OpenGeneric(stream, readerSchema) leaked its schema.
    • Parsing allocated in proportion to the nesting depth for each field default, because it copied the JSON path; a deeply nested schema allocated 55 times its JSON, and now about 20 times at any depth.
    • A container header may hold at most 1,024 metadata entries, and its avro.schema entry at most AvroFileReaderOptions.MaxSchemaLength bytes (4 MiB by default).
    • Pipelined container reading reserved up to 4 × MaxBlockLength per block; it now reserves at most MaxBlockLength.
    • AvroStreamReader decoded an object again after every read, so a stream that returned one byte per read made a 100 KB string cost 100,000 decodes. A decode cut off inside a string, bytes or fixed value now waits for that value's bytes.
    • createReader of AvroMessageReader and AvroRegistryMessageReader could run more than once per schema when several threads met the schema at once; it runs once, as documented.
  • Schema resolution (#130):
    • A writer union branch was read as the first reader branch of the same unqualified name, although another branch had its full name: com.y.Event in a union with com.x.Event failed to read, or was read as com.x.Event and changed branch when written again. Branches now match by full name (or a reader alias) across the whole union first, then by unqualified name, then by promotion, as in Java.
    • A reader field's alias takes the writer field before a reader field of the writer field's name does, as in Java: the specification defines aliases as rewriting the writer's schema.
    • Container blocks of more than 65,536 zero-size objects (nulls, empty records), which Java and AvroFileWriter write, were rejected. They are now limited by AvroFileReaderOptions.MaxZeroSizeValuesPerBlock (16,777,216 values by default, counted as MaxZeroSizeItems is), and AvroFileWriter starts a new block every 65,536 objects. A block of objects that take at least a byte each can no longer declare more objects than it has bytes.
  • Code generation (#131):
    • new T() gives every field its schema default, as reading data that lacks the field does. Before, defaults with no C# literal were dropped: records, fixed values, logical types (date of 0 became 0001-01-01, uuid became Guid.Empty, a decimal of 0.01 became 0), and collections and unions of them. A fixed default left the field null, so the new value could not be written. Such defaults are now stored as their Avro encoding and decoded by the field's reader.

    • Arrays of enums, fixed values, unions, nullable values, logical types and nested collections are written with a loop over the list on C# 12 and earlier. The loop over the list's span, which needs C# 13 on .NET 9 and 10, made projects pinned to C# 12 fail to compile (CS9202).

    • A float or double default beyond the type's range is float.PositiveInfinity (and the like), not the invalid Infinityf.

    • A type named like a member the generator adds to it (Schema, Read, ToAvroBytes, a fixed type's Size or Value), an enum named like one of its symbols, and a type named var are renamed with a trailing _ and reported as AVROGEN005. Before, they did not compile. In the Apache.Avro compatibility mode, which finds types by name, they are errors (AVROGEN003).

    • Contextual keywords as type or namespace names (record, file, scoped, required, partial, nameof, _ and others) are escaped with @, and generated code no longer uses nameof, which a namespace of that name captured.

    • The source generator failed on two types whose names differ only by case (cs.Order and cs.order) and dropped all its output; they now generate as two types. avrosharp gen rejects them with exit code 1 instead of writing one file over the other on Windows and macOS.

    • AVROGEN003 now also reports:

      • two types that map to the same C# name (-m a:M -m b:M with a.X and b.X);
      • a type whose C# name is also a namespace (app.events next to app.events.Click);
      • an AvroSharpNamespace that is not a C# namespace.

      avrosharp gen --namespace rejects such a namespace as a usage error (exit code 2).

    • Two schema files that need each other's types are reported as a circular reference, naming the other file. Before, both errors only said that a type was not defined.

  • AvroSharp.CodeGen could not be loaded by .NET 8 and 9 applications: its only build, for netstandard2.0, was compiled against the System.Text.Json 10 package, which those runtimes do not have. It now also targets net8.0, which uses the framework's System.Text.Json. The source generator keeps the netstandard2.0 build.

0.1.1 - 2026-09-28

The first complete release of all four packages. 0.1.0's publish stopped partway, so AvroSharp.CodeGen 0.1.0 was never published; use 0.1.1. The code is the same as 0.1.0.

Fixed

  • Releasing: AvroSharp.Generators no longer produces a symbol package. It had no .pdb in it, since the generator ships under analyzers/, and nuget.org's rejection stopped the 0.1.0 publish before AvroSharp.CodeGen. The generator's PDB is embedded in its DLL instead. The release workflow now pushes each package separately and checks, while packing, that every symbol package contains a PDB.

0.1.0 - 2026-09-28

The first preview. It covers:

  • Schemas: parsing, writing, canonical form and fingerprints.
  • The generic data model: binary and JSON encoding, and schema evolution.
  • Code generation from .avsc files.
  • Container files with every codec in the specification.
  • Messages: single-object messages and schema-registry framing.
  • Streams of objects.

Packages: AvroSharp, AvroSharp.Codecs, AvroSharp.CodeGen and AvroSharp.Generators. As a 0.x release, the API may still change before 1.0.

Added

  • Nightly fuzzing (.github/workflows/fuzz.yml): every libFuzzer target runs for 30 minutes a night with SharpFuzz, keeping its corpus between runs and uploading crash inputs. The libFuzzer steps in fuzz/README.md are now verified; a first 20-minute run of all seven targets found no crashes.
  • Generated readers resolve through a plan built once per writer schema (#69): Read(ref reader, writerSchema) reads the writer's fields straight into the type's properties, in the writer's order. Fields of the same schema are read directly, numbers (also arrays of them) are promoted in place, enum ordinals are remapped, writer-only fields are skipped and missing fields take their pre-encoded defaults; only other differences (nested records of another version, unions, logical types) are transcoded, one field at a time. On the evolution benchmark this reads in 387 ns, against 506 ns for the generic resolving reader and 1,033 ns for Apache.Avro (i5-3570K); it was 616 ns.
  • AvroValueTransformer.Transform(schema, value, transform): walks a generic value with its schema, through records, arrays, maps and union branches (resolved as the generic writer resolves them), and calls the transform for each non-null primitive, enum or fixed value inside a record field. The transform gets an AvroFieldContext: the record, the field with its properties (for example confluent:tags), the value's own schema, and the field's FullName as Confluent's data rules name it. Records, arrays and maps are copied only where a value changes, so a transform that changes nothing returns the same instance. A depth limit (128 by default) stops cyclic records. This is the basis for field-level data rules, encryption and redaction (#82).
  • Schema-registry wire framing (AvroSharp.Messages), next to single-object encoding and with no registry client dependency: AvroRegistryFraming for Confluent (0x00 + 4-byte ID, byte-identical to Confluent's serializer; and the version 1 0x01 + GUID framing), Apicurio (4- and 8-byte IDs) and AWS Glue (0x03, compression byte, UUID; zlib-compressed payloads read and written), AvroSchemaId, AvroRegistryMessage for writing, and AvroRegistryMessageReader for reading through an IAvroSchemaIdResolver (a synchronous lookup with an asynchronous fill; AvroSchemaIdStore in memory). Read functions are cached per ID with a last-hit check. ConfluentSchemaIdHeader encodes and decodes Confluent's __key_schema_id/__value_schema_id header values, read with ReadPayload. Hostile input (short or unknown headers, unknown IDs, trailing bytes, corrupt zlib, payloads expanding past MaxPayloadLength) raises AvroDataException; a RegistryMessage fuzz target covers it.
  • Schema references for registries: AvroSchema.ToJson(referencedSchemas) writes the named types of referenced subjects by name, as Java's Schema.toString(referencedSchemas, false) does, for a schema parsed against them (AvroSchemaParser.AddNamedSchemas). Checked against Apache Avro Java 1.12 output and against Confluent Schema Registry 7.7, which stores the text unchanged and finds the registered version when it is registered again.
  • Pipelined container reading: AvroFileReader<T>.ReadAllPipelinedAsync(blocksAhead) reads and decompresses blocks on a background task, up to blocksAhead (2 by default) ahead of the caller, who decodes them. Blocks pass through a bounded System.Threading.Channels channel, so memory stays bounded, and each block's buffers return to the pool once it is decoded. An error in a block is raised when the caller reaches that block, and stopping the enumeration stops the background task. The netstandard targets reference System.Threading.Channels (#32).
  • Streams of objects without a container (AvroSharp.Streams): AvroStreamWriter writes objects one after another in the binary encoding, buffered (AvroStreamOptions.BufferSize, 64 KiB by default), with Write/WriteAsync/Flush/FlushAsync. AvroStreamReader reads them back to the end of the stream with TryRead, ReadAll and ReadAllAsync (IAsyncEnumerable<T>), for generated types or as generic values, optionally resolved to a reader schema. Nothing in the encoding delimits objects, so each is decoded to find its end. An object cut off by the buffer is decoded again once more data is read, and a failure is reported only once the stream has ended or MaxDatumLength (64 MiB by default) bytes are buffered, with the object's stream offset. Objects that encode to no bytes cannot be delimited and are rejected (#32).
  • GenericRecord.TryGetValue(int, out AvroValue): field access by position without an exception for a position the record lacks, alongside the name-based overload.
  • AvroSharp.Codecs: the snappy, zstandard, bzip2 and xz codecs in one package, on fully managed libraries: Snappier (plus System.IO.Hashing for snappy's CRC-32), ZstdSharp.Port, SharpZipLib and Lzma.Net. AvroCodecs.All gives a reader every codec, since a file's codec is not known until it is opened. The writing settings and defaults match Apache Avro Java's: zstandard level 3 with an optional content checksum, bzip2 block size 9, and xz level 6. Snappy blocks end with the big-endian CRC-32 of their data, which is checked on read, and zstandard frames are read with or without a content size or checksum. Damaged blocks, and blocks that decompress past MaxBlockLength, are AvroDataException. Checked against files written by Apache Avro Java 1.12.2 in every codec (tests/TestData/java-avro, with the script that writes them), and Java reads the files these codecs write. Snappy and bzip2 are also checked against Apache.Avro C#'s codec packages in both directions. When a file uses a standard codec that is not available, the reader's error names this package (#32).
  • Seeking and splitting container files, as in the Java implementation: AvroFileReader<T>.PreviousSync (the current block's start), Seek (to a block start), Sync (to the first block after a position, found by scanning for the sync marker) and PastSync (whether the current block belongs to the next split). A split [start, end) is read with Sync(start) and while (TryRead(out var v) && !PastSync(end)). Positions before the header's marker go to the first block, so a marker in the metadata (Apache's syncInMeta.avro) is not mistaken for a block. Needs a seekable stream.
  • Asynchronous container files: AvroFileReader.OpenAsync/OpenGenericAsync and ReadAllAsync (IAsyncEnumerable<T>), and AvroFileWriter<T>.WriteAsync/FlushAsync/DisposeAsync; both types implement IAsyncDisposable. The asynchronous paths do no synchronous I/O (tested with a stream that throws on it); objects are encoded into memory and blocks decoded from memory synchronously. Cancellation is checked before every block. The writer now writes the header with its first block, or when flushed or disposed, instead of when created. Microsoft.Bcl.AsyncInterfaces is referenced for netstandard2.0 (it already came in through System.Text.Json).
  • Single-object encoding (AvroSharp.Messages): AvroMessage writes and reads the C3 01 marker and the writer schema's CRC-64-AVRO fingerprint, and writes whole messages for generic values or any type with a write delegate. AvroMessageReader looks each message's fingerprint up in an IAvroSchemaStore (AvroSchemaStore is a thread-safe in-memory one), creates the read function once per writer schema, and can resolve every message to one reader schema. Headers that are missing, unknown fingerprints and bytes left after the object are AvroDataException. Checked byte for byte against Java's messageV1 test message. Apache.Avro C# has no single-object API, so there is no benchmark baseline.
  • Object container files (AvroSharp.Containers): AvroFileWriter and AvroFileReader for generic values or any type with a write or read delegate (generated types pass their static Write/Read methods). The built-in codecs are null and deflate (BCL, raw DEFLATE); others plug in through AvroCodec. Blocks are written at a configurable sync interval, the header carries application metadata, and a write that throws leaves nothing behind. The reader resolves to an optional reader schema, verifies each block's sync marker, and bounds hostile input: block and metadata sizes are limited (MaxBlockLength, 64 MiB by default, also applied to decompressed data), object counts are checked against the block's bytes, and bytes left after a block's last object are an error. Checked against Apache's weather.avro, weather-sorted.avro (deflate) and syncInMeta.avro, and against Apache.Avro in both directions with random schemas and data.
  • Schema resolution for generated types: Read(ref reader, writerSchema) and FromAvroBytes(data, writerSchema) read data written with another version of the type's schema. Data of the same canonical schema takes the direct path; other data is transcoded, following the resolution rules, straight into the type's own encoding in a reused per-thread buffer and read from there, without creating generic values. On the evolution benchmark this reads 1.65x faster than Apache.Avro's resolving reader with 29% of its allocations (i5-3570K).
  • Schema resolution for the generic model: GenericDatumReader.Create(writerSchema, readerSchema) reads data written with one schema version as another, following the specification. Record fields are matched by name or alias, writer-only fields are skipped (sized blocks in one step), and missing reader fields take their defaults. Named types match by full name, unqualified name or alias. Numbers are promoted, string and bytes convert, and enum symbols are matched with the reader's default for unknown ones. Unions resolve per branch. Incompatible schemas are rejected when the reader is created; a mismatch that only some data would hit (a union branch, an unknown enum symbol without a default) is reported when such a value is read.
  • AvroSchemaParseOptions.AllowIdenticalRedefinitions: a parser may accept a named type that an earlier Parse call defined, when both definitions have the same canonical form; the first definition stays in use. The source generator turns it on, so schema sets that inline shared types in every file (as Apache's one-file-at-a-time tooling requires) generate each type once. Different definitions are still an error that names both files.
  • Generator option AvroSharpPropertyNames=avro (CodeGenOptions.PropertyNaming): keep the Avro field names as property names, as Apache's avrogen does, so code written against avrogen classes compiles unchanged. C# keywords are escaped.
  • Apache.Avro compatibility mode for generated code (AvroSharpApacheCompatible=true, requires a reference to Apache.Avro): records also implement Avro.Specific.ISpecificRecord, fixed types derive from Avro.Specific.SpecificFixed, and logical types use Apache's .NET types, so Apache's SpecificDatumWriter<T>/SpecificDatumReader<T> and AvroSharp's serializers work on the same classes and produce the same bytes. Put also accepts what Apache's reader passes: an enum's ordinal, and an AvroDecimal for a decimal on fixed. The generator reports AVROGEN004 when the property is set without the reference.
  • Logical types in generated code:
    • date becomes DateOnly and time-millis/time-micros become TimeOnly (DateTime/TimeSpan where those types don't exist);
    • timestamp-millis/timestamp-micros become DateTimeOffset, and the local-timestamp variants become DateTime;
    • uuid (on string or fixed(16)) becomes Guid;
    • decimal with a precision up to 28 becomes decimal.
    • Set the MSBuild property AvroSharpLogicalTypes=raw to keep the underlying types.
    • The conversions are public in AvroLogicalValues:
      • decimals are exact, raising an error instead of rounding;
      • times and timestamps are truncated towards negative infinity to the logical type's precision;
      • out-of-range data raises AvroDataException.
  • Generated records implement IAvroSpecificRecord: Schema, Get(int) and Put(int, object?), field access by position following the contract of Apache.Avro's ISpecificRecord without depending on it. Put checks the value's type (no implicit widening, as with Apache's casts) and names the field in errors.
  • Code generation from schema files: the AvroSharp.Generators source generator (an incremental generator for .avsc files passed as AdditionalFiles) and the AvroSharp.CodeGen engine it uses. Records become partial classes with static Write/Read methods (plus ToAvroBytes/FromAvroBytes) that call AvroWriter/AvroReader directly in schema order; enums become C# enums; fixed types become size-checked wrappers. Schema files may refer to each other's named types; errors are reported at the file, line and column. Generated readers enforce the same hostile-input limits as the generic reader. The generator needs the .NET 10 SDK or Visual Studio 2026.
  • JSON encoding for the generic model: GenericDatumJsonWriter and GenericDatumJsonReader, following the specification (wrapped union values, byte strings for bytes and fixed, enum symbols). Record fields may appear in any order and missing fields take their defaults; NaN and infinities are written as strings, as Apache.Avro C# does. Checked both ways against Apache.Avro's JsonEncoder/JsonDecoder.
  • Fuzz targets (fuzz/AvroSharp.Fuzz, SharpFuzz/libFuzzer) for schema parsing and generic binary and JSON data, container files, single-object messages and schema resolution (the transcoder checked against the resolving reader), each also checking round trips. They run on every build as a seeded mutation smoke test.
  • Generic data model (AvroSharp.Generic): AvroValue, a 16-byte struct holding any Avro value without boxing (primitives inline, enums as schema plus ordinal), GenericRecord and GenericFixed. GenericDatumWriter and GenericDatumReader compile a schema once into a cached, thread-safe plan of typed nodes; union branches are selected from the value's kind or schema name. Arrays of int/long/float/double are read in bulk. Hostile input is bounded: block counts are checked against the remaining input (using each record's minimum encoded size), pre-allocation is capped, zero-size items draw from a per-read budget, and record nesting is limited when reading and writing (GenericDatumReaderOptions, GenericDatumWriterOptions).
  • AvroReader.ReadLongs/ReadInts: bulk varint reads; on net8+ a Vector128 check decodes runs of one-byte values 16 at a time.
  • Binary encoding (AvroSharp.IO): AvroWriter writes directly into an IBufferWriter<byte> or a Span<byte>; AvroReader reads from a ReadOnlySpan<byte> or a multi-segment ReadOnlySequence<byte>, returning slices of the input for bytes, string and fixed when contiguous. Bulk double/float array items are a single copy on little-endian hardware. Length prefixes are checked against the remaining input before any allocation; malformed data raises AvroDataException.
  • Schema model (AvroSharp.Schemas): immutable primitive, record, enum, array, map, union and fixed schemas; names, namespaces and aliases; record fields with defaults, order and aliases; custom properties.
  • All Avro 1.12 logical types. Unknown or invalid logical types are ignored and kept as properties, as the specification requires.
  • AvroSchemaParser and AvroSchema.Parse/ParseAsync: System.Text.Json parser over UTF-8 with default-value validation, optional comments, and errors that report the JSON path, line and column. A parser keeps named types across calls, so schemas split over several files can refer to each other.
  • Full schema JSON writer, Parsing Canonical Form, and CRC-64-AVRO, MD5 and SHA-256 fingerprints.
  • Tests against Apache Avro's schema-tests.txt vectors, property-based interop tests against Apache.Avro (C#), a Native AOT smoke test, and schema-parse benchmarks gated against Apache.Avro.
  • AvroCodecNames: the codec names defined by the specification.
  • Repository skeleton: build settings, analyzers, public API tracking, strong naming, TUnit tests on .NET 8/9/10 and .NET Framework 4.8.1, and CI on Linux and Windows (x64 and Arm64), including .NET Framework 4.8.1 on Windows (#72).

Changed

  • AvroSchema.ToJson() now gives the text of Apache Avro Java's Schema.toString(): a named type's aliases come after its custom properties, numbers with a fraction or exponent in defaults and properties are printed as Java prints them (1e10 becomes 1.0E10, computed exactly on every runtime), and only ", \ and control characters are escaped. The parsed schema and its canonical form are unchanged.
  • AvroWriter.WriteString encodes in one pass when the buffer has room for the longest possible encoding. It reserves the prefix for 3 bytes per char, encodes, then writes the real length, moving the bytes down when the prefix is shorter. Otherwise it counts first, as before, so a fixed-size span destination that holds exactly the encoding still works. On the string encode benchmark (i7, ShortRun), it went from 0.88x to 0.83x Apache.Avro's time (#25).
  • Container files use their streams less:
    • The writer writes each block, with its count, size and sync marker, in one stream write instead of three. That is one call or one await per block. Its block buffer is sized from the sync interval (#62).
    • The reader keeps its block buffer between blocks while it is large enough, reuses the buffer that limits decompressed size, and decodes objects from an array segment (#62).
    • PastSync checks the stream's length once per block, not for every object of the documented split loop (#62).
    • A header that needs several fills is still re-scanned, but an attempt that runs out of data no longer allocates. Metadata is recorded as positions and turned into strings and copies once, when the header is complete (#70).
  • The container benchmarks and the Native AOT smoke test cover every codec. The benchmark baselines are Apache.Avro's codec packages (#32).
  • Reading generated types with a writer schema (Read(ref reader, writerSchema), as container files and single-object messages do for every object) no longer compares the two schemas' canonical forms for every record. A writer schema remembers the last schema found to have its canonical form, so repeated checks compare references (#60).
  • Resolving records whose reader field order differs from the writer's no longer allocates an array per record; the slots are rented from the shared array pool (#61).
  • AvroMessageReader checks the last schema used before its fingerprint dictionary, which saves the dictionary lookup when messages repeat one schema (#63).
  • Each package ships its own README (AvroSharp.CodeGen and AvroSharp.Generators no longer show the repository README), and all three include THIRD-PARTY-NOTICES.md (#66).

Fixed

  • THIRD-PARTY-NOTICES.md said the packages contain no third-party code, but they compile in source from Polyfill (MIT). The notice now includes Polyfill's copyright and license, and is packed into every package (#66).
  • Generated code needed C# 9 (new() initializers, ??=, is { } and is not patterns), so it failed to compile in netstandard2.0 and .NET Framework projects, which default to C# 7.3. It now uses constructs every version accepts, and emits nullable annotations only for C# 8 and later.
  • Invalid UTF-8 inside a JSON string (schema JSON or JSON data) raised InvalidOperationException instead of AvroSchemaException/AvroDataException. Found by the fuzz smoke test.