Changelog
All notable changes to this project are documented here. The format follows Keep a Changelog and the project uses Semantic Versioning.
Unreleased
1.0.0 - 2026-10-02
The first stable release. The core packages' public API is the one frozen in 1.0.0-rc.1, and from now on it follows semantic versioning: no breaking changes before 2.0. New since 1.0.0-rc.1:
- five add-on packages, released with AvroSharp at the same version: AvroSharp.Confluent for Confluent.Kafka with Confluent Schema Registry, AvroSharp.KafkaFlow, AvroSharp.Azure.SchemaRegistry, and AvroSharp.Aws.Glue with AvroSharp.Aws.Glue.Kafka, all without Apache.Avro;
- a logo, and the packages' icon;
- the fixes below.
The performance gate passes on an i7-12800H, an EPYC 7543 and a Ryzen 5 3500U with .NET 8, 9 and 10, apart from two near-ties: single-value varint writes on Zen+ CPUs, the one exception in docs/design.md §11 (#168), and one bzip2 container write, where both libraries compress with the same SharpZipLib.
Added
- A logo: the packages' icon on NuGet, and the logo in the READMEs and on the documentation site (#220).
- AvroSharp.Confluent, a new package: Confluent Schema Registry serializers and deserializers for Confluent.Kafka, without Apache.Avro (#185). It is released with AvroSharp at the same version.
- Serializers:
AvroSharpSerializer<T>andAvroSharpDeserializer<T>work with:- generated and
[AvroSerializable]types, found throughAvroTypes, or given as anAvroTypeInfo<T>; - the primitives
string,int,long,float,double,boolandbyte[]; - generic values through
AvroSharpGeneric.
- generated and
- Confluent's own code handles the registry: the serializers derive from Confluent's
AsyncSerializer/AsyncDeserializer, so subject name strategies, registration,use.latest.version, schema ID strategies (prefix or header), references and domain and encoding rules are Confluent's code. - Matches Confluent's Avro serializer:
- the same message bytes, and a top-level plain
bytesvalue is the message body alone; with a logical type, such as a decimal, it keeps Avro's length prefix, as Java's serializer writes it (#211); - the same default subject name strategy (
Associated, falling back toTopic); - the same configuration keys (
avro.serializer.*,avro.deserializer.*). Unknown keys are rejected.
- the same message bytes, and a top-level plain
- Schema evolution: the deserializer reads each message in its writer's schema and resolves it to the type's. A generic deserializer with
use.latest.versionand no reader schema reads each message as the latest version. A registered schema with an invalid field default or name still reads: the registry is the authority. - Checks:
- With
use.latest.version,use.latest.with.metadataoruse.schema.id, the serializer checks that the type's schema encodes like the schema whose ID the message carries: the same canonical form, and the same logical types (a decimal's precision and scale, a timestamp's unit). - Tombstones follow Confluent: a null value is written with no body, and reading one into a value type throws.
- The
Nonesubject name strategy is rejected with an error that says so, with or withoutuse.schema.id, which Confluent's client can't look up without a subject (#211).
- With
- Synchronous too: the serializer and deserializer also implement Confluent.Kafka's
ISerializer<T>andIDeserializer<T>, soproducer.Produceworks, and a consumer takes the deserializer withoutAsSyncOverAsync(#189). Once a topic's schema ID, or a schema ID's reader, is known, the synchronous path runs without a task; withuse.latest.version,use.latest.with.metadata,use.schema.idor rules, it waits for the asynchronous one. - Builder extensions for Confluent.Kafka's producer and consumer builders. They set the synchronous interfaces, and take a
RuleRegistry.SetAvroSharpGenericValueSerializerandSetAvroSharpGenericValueDeserializerset generic serdes (#211). - Generic records of several schemas:
AvroSharpGeneric.CreateSerializer(registry)writes each record with its own schema, as Confluent's generic serializer does (#193). - Configuration: the config classes copy any key-value pairs, as Confluent's do, such as a configuration section's (#195). A cached schema ID is read without a lock (#197).
- Dependencies and targets: net10.0, net9.0, net8.0 and netstandard2.0, on Confluent.SchemaRegistry [2.14.0, 3.0.0). The tests run against both ends of that range (
-p:ConfluentVersion=2.14.0for the lowest) and against Redpanda. - CEL rules name the Avro fields, as with Java's and Confluent's serializers, on every type: generated and
[AvroSerializable]types, and generic values (#191). The rules get the value as Apache.Avro's generic model; Apache.Avro comes with Confluent.SchemaRegistry.Rules, which has the CEL executor. - Not supported yet: field rules (field-level encryption,
CEL_FIELD, #186) and migration rules (#187).
- Serializers:
- AvroSharp.KafkaFlow, a new package: KafkaFlow serializer middleware for Confluent Schema Registry on AvroSharp.Confluent, in place of KafkaFlow's
KafkaFlow.Serializer.SchemaRegistry.ConfluentAvro, without Apache.Avro (#154). It is released with AvroSharp at the same version.- Setup:
AddSchemaRegistryAvroSharpSerializerandAddSchemaRegistryAvroSharpDeserializeron producers and consumers. The registry comes from KafkaFlow'sWithSchemaRegistryon the cluster. - Matches KafkaFlow's Confluent Avro serializer: the same subjects and message bytes, and each reads what the other writes.
- Several record types on one topic:
AvroSharpMessageTypeResolverpicks each message's type from its writer's schema. It matches a list of types by their schemas' record names, without reflection, or, like KafkaFlow's, finds the type by its .NET full name in the loaded assemblies. - Not supported: schema IDs in headers, because KafkaFlow's middleware gives serializers no headers.
- More of Confluent's patterns (#212): a top-level union of records under the topic subject, resolved by each message's union branch; a type's record aliases, so a renamed record still matches; and Protobuf or JSON Schema messages reported as such. Without a list of types, a record no loaded type has is looked for again only after more assemblies load.
- Finding types by name:
AddSchemaRegistryAvroSharpDeserializerByTypeName(), in place of KafkaFlow'sAddSchemaRegistryAvroDeserializer(), needs reflection; the overload with the message types doesn't, and rejects an empty list (#212). - Dependencies and targets: net10.0, net9.0 and net8.0, on KafkaFlow.SchemaRegistry [4.0.0, 5.0.0). There's no netstandard2.0 build: every KafkaFlow 4.x needs a newer System.Threading.Tasks.Extensions assembly than AvroSharp's netstandard2.0 dependencies bring. Not strong-named, because KafkaFlow isn't.
- Setup:
- AvroSharp.Azure.SchemaRegistry, a new package: an Azure Schema Registry serializer for Event Hubs and Service Bus messages, in place of Microsoft's
Microsoft.Azure.Data.SchemaRegistry.ApacheAvro, without Apache.Avro (#155). It is released with AvroSharp at the same version.AvroSharpSchemaRegistrySerializer:SerializeandDeserialize, sync and async, to and fromMessageContentand the types derived from it (EventData,ServiceBusMessage). Its methods are virtual, with a protected constructor, so it can be mocked, as Microsoft's can.- Matches Microsoft's serializer: the body is the Avro encoding, and the content type is
avro/binary+<schema ID>. The tests compare the body byte for byte, and each reads what the other writes. - Types: generated and
[AvroSerializable]types, resolved from the writer's schema, and generic records (GenericRecord,AvroValue), read in the writer's schema. - Registry: schemas are registered or looked up once, by full name in the serializer's group (
AutoRegisterSchemas, off by default as in Microsoft's), and writer schemas are fetched once per ID, even by concurrent readers, and parsed without validating defaults or names (#214). Without auto-registration, a schema the group doesn't have is anInvalidOperationExceptionthat says how to fix it. - Generic values: records, enums and fixed values, which have the names the registry needs (#214).
- Dependencies and targets: net10.0, net9.0, net8.0 and netstandard2.0, on Azure.Data.SchemaRegistry [1.2.0, 2.0.0).
- AvroSharp.Aws.Glue, a new package: a managed AWS Glue Schema Registry serializer for Kafka and Kinesis, on every platform, in place of AWS's native, Linux-only
AWS.Glue.SchemaRegistry, without Apache.Avro (#156). It is released with AvroSharp at the same version.AvroSharpGlueSerializer: writes and reads AWS's wire format:0x03, the compression byte (none, or zlib), the schema version UUID with the most significant bits first, then the Avro data. A test vector from AWS's Java encoder pins the format.- Kafka, in AvroSharp.Aws.Glue.Kafka:
AvroSharpGlueKafkaSerializer<T>andAvroSharpGlueKafkaDeserializer<T>for Confluent.Kafka, both asynchronous and synchronous, with builder extensions. They're a package of their own, so Kinesis and other users don't take Confluent.Kafka. A null value, or a nullAvroValue, is a tombstone. - Registry calls, as AWS's serializer makes them: the version is looked up by its definition (Java's
Schema.toString()text, as AWS registers it). With auto-registration the serializer registers a new version, or creates the schema with the configured compatibility, and waits while a new version isPENDING(aTimeoutExceptionafter 10 checks). Each version ID and each schema is cached. A writer schema with an invalid default or name still reads: the registry is the authority. A version another producer is registering is waited for too, concurrent first writes of a schema make one lookup, and each version the serializer registers getsx-amz-meta-transportand theMetadataoption's metadata, as with AWS's serializer. An unknown or deleted version ID in a message is anAvroDataException(#213). - Settings:
AvroSharpGlueOptionshas AWS's settings and defaults:default-registry, schemas named after the topic or stream, auto-registration off, no compression, andBACKWARDcompatibility.SchemaNameStrategygets the transport name, the schema, and whether a key is written (AvroSharpGlueSchemaNameContext), as AWS's naming strategy does. - Types: generated and
[AvroSerializable]types, resolved from the writer's schema, and generic records. - Dependencies and targets: net10.0, net9.0, net8.0 and netstandard2.0, on AWSSDK.Glue [4.0.0, 5.0.0), and AWSSDK.Core [4.0.3.3, 5.0.0) (below 4.0.3.3 it has an advisory). AvroSharp.Aws.Glue.Kafka adds Confluent.Kafka [2.0.2, 3.0.0).
- Tests: an in-memory Glue client, and moto (an AWS emulator in Docker) through the AWS SDK's real client. Not yet against AWS's own package, whose native library didn't start in the Docker environment tried.
- The add-ons on .NET Framework and Native AOT (#198): the Confluent, Azure and Glue tests also run on .NET Framework 4.8.1, and CI publishes an application using AvroSharp.Confluent, AvroSharp.Azure.SchemaRegistry, AvroSharp.Aws.Glue and AvroSharp.Aws.Glue.Kafka with Native AOT, failing on any trim or AOT warning in an AvroSharp assembly.
Fixed
- Schema JSON that ends in a backslash (an unfinished escape) is a schema error. The check for escaped unpaired surrogates, which runs before the JSON parser, threw
ArgumentOutOfRangeExceptionfor it, fromAvroSchema.Parseand from a container file whose header held such a schema. Found by the nightly fuzzing (SchemaParse and ContainerFile). AvroTypes.TryGet(Type)andAvroTypes.Get<T>()find a generated type whose assembly no code has run from yet, such as one loaded only for its metadata, by a type name or a message handler's signature. They run the assembly's module initializer, which registers its types, before they report the type unknown.
1.0.0-rc.1 - 2026-10-01
The release candidate for 1.0: the public API is frozen, and 1.0.0 follows when .NET 11 is released. New since 0.2.0:
- schemas and serializers from your own C# types (
[AvroSerializable]), and lookup by type without reflection (AvroTypes); - schema compatibility checks with every reason, in code and as
avrosharp schema compat; - container files whose schema only Java accepted are now read;
- the fixes from three reviews: an independent one, Apache.Avro's C# test cases, and Apache Avro Java's behavior.
The performance gate passes on an EPYC 7543 with .NET 8, 9 and 10: 330 of 330 comparisons are faster than Apache.Avro and allocate no more.
Breaking
The public API review before 1.0 (#134) renames and moves members, so the API can be frozen without them (#73). Regenerate checked-in generated code with the new avrosharp gen: code from 0.2 calls support members that moved. From 1.0, code generated by a version keeps compiling against later runtimes of the same major version (the policy).
- Generated-code support:
AvroGeneratedCode,AvroRecordPlan,AvroPlanCache,AvroConversionandAvroUninitializedmove toAvroSharp.Serialization.Generated.IAvroCodec<T>is nowIAvroValueSerializer<T>, the primitive codecsAvro…Serializer, and the generated nestedAvroCodecstructValueSerializer: "codec" means block compression only.- Removed, because generated code no longer calls them:
GetRecordPlan(writer, reader)without the cache,PutTypeMismatch(object?, string, string),AvroRecordPlan.Target(int)andConversion(int), andIsSameSchema(useAvroSchema.HasSameCanonicalForm).
- Exceptions:
AvroDataException(wasAvroSharp.IO) andAvroSchemaException(wasAvroSharp.Schemas) are inAvroSharp. - Schema lookup:
IAvroSchemaStoreis nowIAvroSchemaResolver, withGetSchemaAsync, which an implementer has to add. TheAvroMessageReaderfactories name its parameterresolver, andAvroMessageReader.CreateGenericnames its reader optionsreaderOptions, as the other reader factories do. AvroSerializer:Serialize(output, value)andTrySerialize(destination, value, out bytesWritten)take the output first.- Renamed:
GenericDatumReader.SchemaandGenericDatumJsonReader.Schemaare nowWriterSchema;EnumSchema.Defaultis nowDefaultSymbol;AvroReader.Skip(long)is nowSkipRaw;AvroValue.FromByteArray,FromGenericRecordandFromGenericFixedare nowFromBytes,FromRecordandFromFixed;AvroFileReader<T>.GetMetadataStringis nowTryGetMetadataString(key, out value);GenericRecord.TryGetValue(int index, out AvroValue value)names its first parameterposition, as the rest ofGenericRecorddoes;CodeGenOptions.DefaultNamespace,NamespaceMappingandPropertyNamingare nowNamespace,NamespaceMapandPropertyNames.
- Changed types:
AvroFileReader<T>.Codecis theAvroCodec(its name isCodec.Name);AvroFileWriterOptions.MetadataholdsReadOnlyMemory<byte>values;SchemaFingerprint.Crc64AvroEmptyis along;ConfluentSchemaIdHeader.Encodetakes anAvroSchemaId, and rejects a numeric ID: the header holds a GUID.
AvroCodec.CreateDeflateis replaced bynew DeflateCodec(level).AvroCodecNamesmoves toAvroSharp.Containers, next toAvroCodec(#73).uuidstrings with spaces around them are rejected (AvroDataException): the reader takes the RFC 4122 form only, as Java'sUUID.fromStringdoes. Before,Guid.TryParseExacttrimmed them (#135).- Validation: the generic and schema parse options reject a
MaxDepthbelow 1 and a negativeMaxZeroSizeItems.CodeGenOptionsrejects nullable annotations with aLanguageVersionbelow 8, and any version below 7.
Added
- The public API is frozen for 1.0 (#73). Every package declares its API as shipped, and package validation compares each with its 0.2.0 on nuget.org. The breaks listed above are the only ones allowed; any other fails the pack.
AvroMessageReader<T>.ReadAsync, which fetches an unknown fingerprint through the resolver and keeps its reader, andAvroSchemaStore.GetSchemaAsync(#134).- Schemas and serializers from C# types (#31): mark a
partialclass or record class[AvroSerializable], and theAvroSharp.Generatorspackage writes its schema and serializers from its members.- What the type gets: the same members as a type generated from a
.avscfile, soAvroSerializer, container files, messages and the registry readers take it alike. The serializers are the same code, with no reflection, and they're AOT-clean. - The mapping:
- primitives;
- nullable types, as unions with
nulland a default ofnull; - C# enums,
List<T>andDictionary<string, T>; - other
[AvroSerializable]classes, as records; Guid,DateOnly,TimeOnlyandDateTimeOffset;decimalwith[AvroDecimal], andbyte[]orGuidasfixedwith[AvroFixed];objectwith[AvroUnion], as a union of records.
- Field names are the member names as written, as Apache Avro's do.
AvroNaming.CamelCaseconverts them, per type or with[assembly: AvroSerializableDefaults(FieldNames = ...)].[AvroName],[AvroAlias],[AvroDoc](or the XML summary),[AvroDefault],[AvroIgnore]and[AvroFieldPosition]cover the rest.[AvroLogicalType]on a rawlong,intorstringkeeps the value raw. - Diagnostics AVROGEN101–118 report what the first version doesn't support, at the code:
init-only members, primary constructors, nested or generic types, aDateTimewithout a logical type, and enums whose values aren't 0, 1, 2 and so on. - The design is in docs/design.md §6.5, the guide in Code generation, and the code in the SerializableTypes sample.
- What the type gets: the same members as a type generated from a
AvroTypes(#31): a type's schema and read and write functions, by type argument or byType, without reflection, for integrations and generic code on every target.- Every generated type has
AvroTypeInfo. It registers itself on .NET 5 and later (with a module initializer); elsewhere, callAvroTypes.Registeronce. - The primitives that Confluent's serializers support are registered, and read promoted writer schemas.
- Every generated type has
- Schema compatibility checks (#165):
AvroSchemaCompatibility.Check(writerSchema, readerSchema)says whether data written with one schema can be read with another.- The verdict is
Compatible,PartialorIncompatible.Partialmeans some values can't be read: an enum symbol, or a writer union branch, that the reader lacks. The readers leave these until such a value is read. The verdict is a fact about the schemas, which the options don't change.IsCompatiblesays whether the check passes: onlyCompatiblepasses by default, as in Java'sSchemaCompatibilityand schema registries;AllowPartialacceptsPartial. - Every incompatibility is reported, not just the first (
Incompatibilities, with the warnings inWarnings). Each has a kind, a message, and a path into the data ($.items[].sku,$.payment[1:Card].number). A named type used in several places lists its other paths. - The kinds are Java's six, plus
InvalidDefault. - The check uses the resolving readers' own rules, and the tests check every row against
GenericDatumReader.Createand the generated-code plans. - Warnings cover differences that the specification allows but that can change the values read: decimal scale or precision, other logical type changes, lossy promotions, names matched without their namespace, enum defaults used for unknown symbols, and ambiguous field aliases.
WarningsAsErrorsmakes them fail the check. - Schema registry levels:
CheckVersions(schema, previousVersions, level)checks a new version at Confluent Schema Registry's levels (Backward, Forward, Full and their transitive versions), with a result for each pair; each pair'sIndexis the earlier version's position in the list. - The CLI:
avrosharp schema compat <writer> <reader>, or--level backward-transitive v1 v2 v3. It has text or--jsonoutput (with aformatVersion),--warnings-as-errors(or--strict) and--allow-partial. Exit code 3 means partially compatible, and 4 incompatible or, with--warnings-as-errors, warnings. - Documentation: a section in Schema evolution, and "Replacing
Schema.CanRead" in the migration guide, with a table of where Apache.Avro's verdicts differ.
- The verdict is
AvroSchema.HasSameCanonicalForm: whether two schemas have the same encoding.Equalsstays reference equality, asAvroSchema's docs now say (#134).DeflateCodec: the built-in codec is public, withDefaultandLevel, as the Codecs package's codecs have (#134). It rejects a level the runtime doesn't have when it is made, rather than when a writer compresses its first block.- New overloads:
AvroSerializer.Deserialize<T>(in ReadOnlySequence<byte>, AvroSchema writerSchema);GenericDatumReader.Read(in ReadOnlySequence<byte>);AvroMessage.Write(output, in AvroValue, GenericDatumWriter);AvroRegistryMessage.Write(output, framing, id, in AvroValue, GenericDatumWriter)(#134).
- Default limits as constants:
GenericDatumReaderOptions.DefaultMaxDepthandDefaultMaxZeroSizeItems,GenericDatumWriterOptions.DefaultMaxDepthandAvroSchemaParseOptions.DefaultMaxDepth(#134). - Generator settings (#134):
- The
AvroSharpNamespaceMapMSBuild property,avro.ns:CSharp.Nsentries separated by;or,, like the CLI's--namespace-map. The package'sbuild/AvroSharp.Generators.targetspasses it on with,for;and line breaks, since the compiler reads it from an.editorconfigfile, where;starts a comment and a line break ends the value. So the entries can also be written one per line, or given with-p:. - Generator properties accept their values in any case.
- A value the generator doesn't recognize is warning AVROGEN006; before, a typo was ignored silently.
--language-version 7implies--no-nullable.
- The
- Records built from another record's fields: a record given a field that belongs to another record takes a copy, so
new RecordSchema(name, other.Fields.Append(field))works. It threw before (#134). - Documentation:
- Migrating from Apache.Avro: the plan, and the Apache.Avro API next to AvroSharp's. Its code is the Migration sample's, which CI runs against both libraries, and a test keeps the two in step (#21).
- The GeneratorPackage sample: an application's project with the
AvroSharp.Generatorspackage, its MSBuild settings and schemas across files, and how to use the generator from source instead (#126, #74). - A Getting started section: installing, a first program, generated types, container files, schema evolution, logical types and JSON. Each page's code is a sample's (the new ContainerFiles, SchemaEvolution, LogicalTypes and Json samples), and a test keeps them in step.
- Links from the guides and READMEs to the API reference, an API map, namespace overviews, and cross-references between the API pages.
- A dev container (
.devcontainer/) andbuild/ci-local.sh, which build and test as the Linux CI job does. The container has Ubuntu 24.04 with the .NET 8, 9 and 10 SDKs, the Native AOT prerequisites, CI's variables and a cached NuGet volume; the script runs the job's steps in order. DocFX and ReportGenerator are pinned local tools (.config/dotnet-tools.json), used by CI, the docs workflow and the container alike (#139).
Fixed
- Container files whose schema Java reads are read. A file's schema is parsed as Java parses it, without validating names or defaults, and with Java's leniencies: a
"doc"or an enum's"default"ofnull, and anamespacethat is not a string. Before, a field nameduser-idorgröße, or a"default": nullon astringfield, made the file unreadable (The file's schema is invalid). These are common in files written by Java before 1.9, and by other tools.AvroSchema.Parsestill validates. - Out-of-range logical values in generated types: a
timestamp-millisofLong.MaxValue, a common "never" sentinel that Java reads, can't be aDateTimeOffset."avrosharp.logicalType": "raw"on a schema now keeps that one logical type's underlying type, so the property is alongand reads it. The other fields keep their .NET types."native"maps one schema to its .NET type underAvroSharpLogicalTypes=raw.- The range errors say how to read such values.
- Not available with
AvroSharpApacheCompatible, which reports it.
- Documented: logical types don't take part in schema resolution, as in Java, so a decimal of another scale reads as another number, although the specification says such decimals don't match.
AvroSchemaCompatibilitywarns about it. - Message readers (#159, #160):
- Concurrent
ReadAsynccalls that met the same new fingerprint or registry ID each asked the resolver. Now one fetches and the others wait for it. A fetch that fails or is cancelled caches nothing. - Bad data made
ReadAsyncandReadPayloadAsyncthrow from the call itself. It now faults the returned task, so reads started together, then awaited, all complete.
- Concurrent
- Code generation (#161): a type whose C# name would hide a namespace or type the generated code uses is an error. For example, a namespace map onto
AvroSharpwith a recordSerializationused to generate code that didn't compile. - Generated types (#162): a stored default of more than 65,536 zero-size items (an array of nulls, say) failed every read of older data. Defaults now decode without that limit, through the new
AvroRecordPlan.DefaultReader. - Enum and fixed values written as another schema: the generic binary and JSON writers wrote an enum value's ordinal in its own schema, so a symbol at another position in the target schema became another symbol, silently, and one the target lacks was written anyway; and a fixed value of another size wrote all its bytes, shifting every later field. The symbol is now found in the target schema, and the size is checked. Values of the target schema itself, or of a copy parsed separately, write faster than before.
- Escaped unpaired surrogates (
\ud800) in a schema or in JSON data threwInvalidOperationException. A schema with one, anywhere, is now anAvroSchemaException(container file headers included), and JSON data anAvroDataException. - .NET Framework: the JSON reader read
-0.0as+0.0and rejected a number beyonddouble's range; it reads-0.0and an infinity, as on the other targets and in Java. - A bad namespace is reported at
$.namespace, not$.name. - A fixed type named
EqualsorGetHashCodegenerated code that did not compile (CS0542); it is renamed, with a note, like the other generated member names (#141). - A decimal default that the generated C#
decimalcannot hold (beyond 96 bits, or with more digits than the precision) generated code whose constructor threw; it is now a generation error (#141). - Union defaults in generated code are for the first branch they are a value of, as Avro 1.12 says and the readers do, not always the first branch.
["int","string"]with a string default failed to generate, a[enum,"string"]default that isn't a symbol generated code that didn't compile, and a decimal default in a later branch escaped the check above. - C# 7.0 and 7.1 projects: the generator read C# 7.0 as version 0 and accepted 7.1, whose generated code doesn't compile (readonly structs need C# 7.2). Both are now error AVROGEN003, which says to set
<LangVersion>to 7.3 or later.CodeGenOptions.LanguageVersion7 means C# 7.2 or later. - A default that isn't a value of its field's schema, on a field built in code (the parser checks parsed ones), made
GenericDatumReader.Createand the generated-code plan throwAvroDataException. It is now the documentedAvroSchemaException, naming the field.
Changed
- A coverage badge in the README.
- CI's merged line coverage is written by
build/coverage-summary.shin shields.io's endpoint format. - The docs workflow publishes it with the site. It now deploys after CI passes on main, not on every push.
- CI's merged line coverage is written by
- JSON numbers as Java reads them: the JSON reader accepts a whole number written as
1.0or1e2for anintorlong, and"INF"and"-INF"for afloatordouble, as Java'sJsonDecoderdoes. The writer is unchanged. - Faster logical values, parsing, resolution and strings (#135). Measured on the EPYC 7543 with main and the branch in parallel lanes on neighbouring CCDs, on both sockets:
- Decimal writes: 2.9× faster on bytes (178–188 to 62 µs per 1,024
decimal(18,4)values; Apache.Avro takes 94 µs), and 2.6× on fixed.Pow10's multiplication loop became a table,decimal.GetBitswrites into the stack on .NET 5+, and the rounding check runs only for a value with more fractional digits than the scale (.NET 7+). uuidstrings: writes −40%, reads −50%. They're parsed from and formatted to UTF-8 (Utf8Parser/Utf8Formatter), without a string, and are faster than Apache.Avro both ways.- Schema parsing: 9–13% faster, with 33–37% less allocation, for small and large schemas, with or without the fingerprint. There's no closure per named type, field, union and alias, a logical type is read without a dictionary, and the canonical form's string is made only when asked for.
- The resolving generic reader: 9–10% faster; with record, array, map and bytes defaults, 30–37% faster. Each default is converted from JSON once: immutable ones are shared, mutable ones decoded from their encoding per read.
- String writes: 16% faster (the Strings workload). The length prefix is reserved for 1 byte per char, so ASCII strings of 22–63 chars aren't moved.
- A
LogicalTypeBenchmarksand a resolution benchmark with defaults were added to measure these.
- Decimal writes: 2.9× faster on bytes (178–188 to 62 µs per 1,024
- Tests for the remaining gaps of the test review (#142): enum and fixed aliases, the resolving reader's options, the JSON writer's widening, depth limit and shape errors, maps of any
IReadOnlyDictionary, hostile container files (overlong varints, truncation on the async path, trailing bytes when pipelined),AvroRegistryMessageReader.ReadAsyncwith a missing schema or a cancelled fetch, logical-value range errors, and control characters in schemas. - The nightly fuzzing workflow also runs a random-schema code-generation test (#141): random schemas, with hostile names and every kind of default, are generated, compiled for C# 7.3, 12 and the latest version, and round-tripped. PR CI runs it on 100 schemas.
- The performance gate checks every benchmark (#157).
LogicalTypeBenchmarkswas not gated: its methods weren't namedAvroSharp_*, and only one of its categories had an Apache.Avro baseline.- Each logical type and operation is now a group with an Apache.Avro baseline, adding decimal reads, decimal on fixed and timestamp reads.
uuidon fixed has no Apache.Avro equivalent. ResolutionBenchmarksgains an Apache.Avro baseline for reading with defaults.- The gate fails a benchmark it can't compare, instead of skipping it. Benchmarks without an Apache.Avro equivalent are marked
Ungated.
- Each logical type and operation is now a group with an Apache.Avro baseline, adding decimal reads, decimal on fixed and timestamp reads.
- The API reference on the documentation site is built from the net10.0 build, so it shows the .NET 8+ API (
AvroSerializer,IAvroSerializable<T>and the overloads that take no delegates), and lists those members with their equivalents on other targets (#136).
0.2.0 - 2026-09-29
The avrosharp command-line tool; generated types that read older versions of their schema, write and read memory without allocating, and start with their schema defaults; and the fixes from an independent review: bounded memory and nesting for hostile input, schema resolution that follows the specification and Java, and code generation fixes. The source generator needs the .NET 10 SDK or Visual Studio 2026 and later; the generated code and the tool run on .NET 8 and later (and the code on .NET Standard 2.0).
Added
- The
avrosharpcommand-line tool, packageAvroSharp.Tool: adotnet toollike Apache.Avro'savrogen, for .NET 8 and later (#33).avrosharp genwrites the source generator's code for.avscfiles, or folders of them, with the generator's options: one file per type, in folders for its namespace, or with--flatin one folder.avrosharp schema canonicalandschema fingerprintprint the Parsing Canonical Form, and the CRC-64-AVRO, MD5 or SHA-256 fingerprint as hex, base64 or (CRC-64) the decimal Java prints.- Files may refer to each other's types in any order. Errors are in the compiler's format, with the generator's diagnostic IDs. The exit code is 0 on success, 1 when the command fails and 2 for an invalid command line.
CodeGenOptions.NamespaceMapping(--namespace-mapin the tool): C# namespaces for Avro namespaces and the namespaces under them, as avrogen's--namespace a:b. Only the generated C# namespace changes, not the schema.GeneratedSource.Namespacegives each file's C# namespace.SchemaFileSetinAvroSharp.CodeGen: parses schema files that refer to each other's named types, in any order, as the source generator and the tool do.- Bulk booleans (#26):
AvroReader.ReadBooleanschecks and copies a block of booleans as one span, andAvroWriter.WriteBooleanswrites one as a single copy. Generated code uses them forbooleanarrays, and the generic reader and writer use them for boolean arrays. - Generated records implement new interfaces (#115):
IAvroWritable(WriteTo(ref AvroWriter)) andIAvroReadable(ReadFrom(ref AvroReader), which fills an existing instance and reuses its lists, dictionaries and records), and on .NET 8 and laterIAvroSerializable<T>with the staticSchema,WriteandReadmembers. With it,AvroSerializer.Serialize/TrySerialize/Deserialize,AvroFileWriter.Create<T>(stream),AvroFileReader.Open<T>(stream)/OpenAsync<T>,AvroStreamWriter.Create<T>,AvroStreamReader.Open<T>,AvroMessage.ToArray<T>/Write<T>,AvroMessageReader.Create<T>,AvroRegistryMessage.ToArray<T>/Write<T>andAvroRegistryMessageReader.Create<T>need no delegates. Readers of files and messages resolve other versions of the type's schema. - Generated records write and read memory without allocating (#115):
TryWriteAvroBytes(Span<byte>, out int)(no exception when the value does not fit),WriteAvroBytes(Span<byte>),WriteAvroBytes(IBufferWriter<byte>),FromAvroBytes(data, out int bytesConsumed)andFromAvroBytes(in ReadOnlySequence<byte>). AvroWriter.WriteIntsandWriteLongswrite array items in bulk.CodeGenOptions.LanguageVersion: the C# version generated code may use. The source generator sets it from the project; with C# 11 or later, generated code addsIAvroSerializable<T>for .NET 8 and later.- Generated code is easier to debug and read (#117): records have a
[DebuggerDisplay]with their first three simple fields, the serializers are[DebuggerNonUserCode], and a property whose logical type keeps its raw type (a decimal wider than 28 digits,duration,timestamp-nanos, or any logical type withAvroSharpLogicalTypes=raw) says in its documentation what the value means. - Fixed types have
==,!=andAsSpan()(#117). - The generator reports a property renamed to avoid a clash (for example
user_idanduserIdin one record) as informational diagnostic AVROGEN005, naming both fields;GeneratedSource.Notescarries these notes for other callers ofCSharpCodeGenerator(#117). - Generated files suppress CS8981, so all-lower-case Avro type names compile without warnings (#117).
AvroFileReaderOptions.MaxSchemaLengthandMaxZeroSizeValuesPerBlock(#129, #130).GenericDatumJsonReader.ReadDefault: converts a schema default to anAvroValue, the value a reader gives a field that the data lacks (#131).
Changed
- Packages list only the dependencies they use (#20).
AvroSharp.GeneratorsandAvroSharp.CodeGendepend only on AvroSharp, andAvroSharp.Codecsonly on AvroSharp and its compression libraries; before, they also listed System.Memory, System.Text.Json and other packages that AvroSharp brings. AvroSharp's netstandard2.1 build no longer depends on System.Memory or Microsoft.Bcl.AsyncInterfaces, which that target has built in. The generator still needs the .NET 10 SDK or Visual Studio 2026 and later;docs/design.mdrecords what fails on older SDKs. - Faster varint encode (#102). Where PDEP is fast (Intel from Haswell, AMD from Zen 3), 3- to 8-byte values are one inline 8-byte store; elsewhere one out-of-line call writes 3 to 5 bytes with one 4-byte store. On the i7-12800H, 3- to 8-byte encodes are 28–38% faster and 3-byte encode is 1.52× faster than Apache.Avro (1.09× before); on the EPYC 7543 they are 25–35% faster, and on an i5-3570K and a Ryzen 5 3500U without fast PDEP, 3 to 5 bytes are 7–33% faster. 1-byte encode is 6–8% slower on the fast-PDEP machines. Record writes are unchanged or faster; generated writes are 7–8% faster on the i7 and the EPYC.
AvroWriter.WriteLongs/WriteIntskeep the write position in a local while the buffer has room, instead of storing and reloading it after every value (#27, the per-value round trip through memory that limited one-at-a-time writes on the EPYC). They are at least as fast as single writes for every length on the machines measured, and 1-byte bulk writes are 19–48% faster than single ones.- Bulk
ReadLongs/ReadIntsdecode runs of one-byte values 16 at a time with vectors, without a loop over the run (#24). Mixed data (90% small values) is read 1.48× faster than the plain loop on the i7 and the EPYC, and small values 2.7–3.6× faster; timestamps stay within 3% of the plain loop, as #29's rule requires. - Schema parsing uses a pooled
JsonDocumentand copies out only what a schema keeps (defaults and custom properties) (#103). Parsing a small schema is 2.07× faster than Apache.Avro on the i7 (1.08× before) and allocates 7.2 KB instead of 8.4 KB. Content after the schema JSON is now reported by the JSON parser, with its line and column. - The generic reader stores arrays of
boolean,int,long,floatanddoubleitems as the primitives themselves, not oneAvroValueeach (#23).AsArray()still returns the items as values. The newTryGetInt64Arrayand the other typed accessors return the memory without a copy, andFromInt64Arrayand the other typed factories wrap existing memory. The writer takes these arrays from their memory: booleans, floats and doubles as one copy, ints and longs with no kind check per item. On the i7-12800H (.NET 10), reading an array of 1,000 items with the generic model is 4.4× faster for ints (3,241 to 734 ns), 3.5× for longs and 6.2× for doubles, and allocates a quarter to a half as much. A record with 64 long counters reads in 591 ns instead of 748 ns. Against Apache.Avro, array reads are now 6.6–16.6× faster, up from 1.7–2.8×. - The array returned by
AsArray()for these item types is no longer aList<AvroValue>. Code that cast it toList<AvroValue>has to copy it instead (AsArray()has always been documented asIReadOnlyList<AvroValue>). - Generated code is smaller (#114). Nullable unions of a primitive or a record, and arrays and maps of primitives, strings, bytes or records, are one call to an
AvroGeneratedCodehelper instead of an inlined switch or loop. The union helpers are inlined; the collection helpers take struct codecs (IAvroCodec<T>) that the JIT specializes. The schema-evolution reader is oneReadFieldswitch instead of a case and a helper method per field. For a production-style schema of 12 records, generated code went from 18,653 to 9,193 lines and its assembly from 377 KB to 165 KB; the 31-field testOrderrecord went from 2,200 to 1,038 lines. - Generated records with more than 32 fields are serialized by several methods of up to 32 fields each, which keeps each within the JIT's limit on tracked locals (#113). On a 140-field record (
WideRecordBenchmarks, i7-12800H, .NET 10), a generated read takes 525 ns, against 661 ns with one method per record and 729 ns before this change; a write takes 374 ns, against 419 and 492 ns. - Generated readers no longer allocate each collection twice: they construct records with a constructor that skips the property initializers (#113). Block counts are checked against the input without a 64-bit division, the plan for another writer schema is cached per generated type,
intandlongarrays are written with the newAvroWriter.WriteInts/WriteLongs, anddateconverts throughDateOnly.DayNumber. A generated read of the benchmarkOrderrecord allocates 2,048 B instead of 2,192 B, with times unchanged within noise; reading another version of a schema takes 179 ns (181 to 193 ns before). new T()gives every field with a schema default its default (#112): primitives, strings, bytes, enums, nullable unions whose default is for their first branch, and arrays and maps of those. Before, only reading data that lacked the field applied the default. Code that relied on these fields starting at zero or empty sees the defaults instead.- The Apache compatibility mode makes code written for
avrogenclasses compile unchanged (#116): property names default to the Avro field names, as avrogen's are (AvroSharpPropertyNames=pascalkeeps PascalCase), and types have avrogen's static_SCHEMAand an instanceSchema(Apache'sAvro.Schema). AvroSharp's schema isAvroSharpSchemain that mode, on records as on fixed types.CodeGenOptions.PropertyNamingis now nullable:nullmeans the mode's default. - Generated property names title-case segments in capitals (#117):
USER_IDbecomesUserId(it wasUSERID) andHTTP2_PORTbecomesHttp2Port; segments with lower-case letters keep their capitals (txIdisTxId). This renames such properties in existing generated code. [GeneratedCode]on generated types carries the generator's package version (for example0.1.2) instead of0.0.0.0(#107): MinVer setsAssemblyVersiontomajor.0.0.0.- Generated code no longer checks a union branch's value for null after a type pattern proved it is not, and null checks have one pair of parentheses (#107).
Fixed
- Generated code: a line terminator other than
\nin a schema'sdoctext (\r, U+0085, U+2028 or U+2029) ended the generated///comment, so the rest of the doc was compiled as C#. For a field's doc, that could add members to the generated type. Doc text is now split on every C# line terminator, and other control characters are dropped (#110). - Generated readers limited zero-size array items (
nulls, empty records) per array, not per value. Arrays nested in arrays multiplied the limit: an 8 KB input could declare about 131 million items and allocate about 2 GB. The limit of 65,536 now covers the whole value, as in the generic reader. Each generatedReadstarts a new budget, so values read one after another from the same reader, as in a container block, each get the full limit (#109). - The generator reports a union whose branches map to the same C# type, for example a
uuidstring and auuidfixed (bothGuid), as error AVROGEN003. Before, it generated code that didn't compile (CS8120), and whose writer couldn't have chosen the branch anyway. The message suggestsAvroSharpLogicalTypes=raw(#108). - Generated writers check enum values (#111): a C# enum can hold any number, which was written as an out-of-range ordinal that readers reject. Writing it now throws
AvroExceptionnaming the field. - A
bytesdecimal with no bytes is rejected withAvroDataExceptioninstead of being read as 0, as in Java (#111). The Apache compatibility mode's decimals check the same. Puton a generated record'snull-typed field rejects values other thannull, which were silently dropped on write (#111).- A generated type's
Schemais one instance even when first read on several threads at once. Before, a thread that lost the race to parse it could keep its own instance, which missed the cache of resolution plans and the same-schema fast path. The newAvroGeneratedCode.PublishSchemastores the first one. - Hostile input (#129):
- A writer schema that nested arrays, maps or unions between recursive records overflowed the stack, which ends the process:
MaxDepthcounted only records. Values may now be nested 8 ×MaxDepthlevels deep (1,024 by default), counting arrays, maps and records, on every path that reads by the writer's schema: the generic and resolving readers, skipped fields, and the transcoder generated types use. The thread's stack is also checked every 16 levels of new depth. - A record of many
nullfields takes no bytes but creates a value per field, so 3 bytes could allocate 160 MB.MaxZeroSizeItemsnow counts each zero-size item as the values reading it creates: one, plus one per field of each record in it. - The bzip2 codec let
IndexOutOfRangeExceptionescape on corrupt blocks; bzip2, xz, zstandard and snappy now report every corrupt block asInvalidDataException. A snappy length of 2^31 or more is rejected instead of ending inArgumentOutOfRangeException. - The resolving-reader cache kept every writer schema alive for as long as its reader schema lived, so each file opened with
OpenGeneric(stream, readerSchema)leaked its schema. - Parsing allocated in proportion to the nesting depth for each field default, because it copied the JSON path; a deeply nested schema allocated 55 times its JSON, and now about 20 times at any depth.
- A container header may hold at most 1,024 metadata entries, and its
avro.schemaentry at mostAvroFileReaderOptions.MaxSchemaLengthbytes (4 MiB by default). - Pipelined container reading reserved up to 4 ×
MaxBlockLengthper block; it now reserves at mostMaxBlockLength. AvroStreamReaderdecoded an object again after every read, so a stream that returned one byte per read made a 100 KB string cost 100,000 decodes. A decode cut off inside a string, bytes or fixed value now waits for that value's bytes.createReaderofAvroMessageReaderandAvroRegistryMessageReadercould run more than once per schema when several threads met the schema at once; it runs once, as documented.
- A writer schema that nested arrays, maps or unions between recursive records overflowed the stack, which ends the process:
- Schema resolution (#130):
- A writer union branch was read as the first reader branch of the same unqualified name, although another branch had its full name:
com.y.Eventin a union withcom.x.Eventfailed to read, or was read ascom.x.Eventand changed branch when written again. Branches now match by full name (or a reader alias) across the whole union first, then by unqualified name, then by promotion, as in Java. - A reader field's alias takes the writer field before a reader field of the writer field's name does, as in Java: the specification defines aliases as rewriting the writer's schema.
- Container blocks of more than 65,536 zero-size objects (
nulls, empty records), which Java andAvroFileWriterwrite, were rejected. They are now limited byAvroFileReaderOptions.MaxZeroSizeValuesPerBlock(16,777,216 values by default, counted asMaxZeroSizeItemsis), andAvroFileWriterstarts a new block every 65,536 objects. A block of objects that take at least a byte each can no longer declare more objects than it has bytes.
- A writer union branch was read as the first reader branch of the same unqualified name, although another branch had its full name:
- Code generation (#131):
new T()gives every field its schema default, as reading data that lacks the field does. Before, defaults with no C# literal were dropped: records, fixed values, logical types (dateof 0 became 0001-01-01,uuidbecameGuid.Empty, a decimal of 0.01 became 0), and collections and unions of them. A fixed default left the fieldnull, so the new value could not be written. Such defaults are now stored as their Avro encoding and decoded by the field's reader.Arrays of enums, fixed values, unions, nullable values, logical types and nested collections are written with a loop over the list on C# 12 and earlier. The loop over the list's span, which needs C# 13 on .NET 9 and 10, made projects pinned to C# 12 fail to compile (CS9202).
A float or double default beyond the type's range is
float.PositiveInfinity(and the like), not the invalidInfinityf.A type named like a member the generator adds to it (
Schema,Read,ToAvroBytes, a fixed type'sSizeorValue), an enum named like one of its symbols, and a type namedvarare renamed with a trailing_and reported as AVROGEN005. Before, they did not compile. In the Apache.Avro compatibility mode, which finds types by name, they are errors (AVROGEN003).Contextual keywords as type or namespace names (
record,file,scoped,required,partial,nameof,_and others) are escaped with@, and generated code no longer usesnameof, which a namespace of that name captured.The source generator failed on two types whose names differ only by case (
cs.Orderandcs.order) and dropped all its output; they now generate as two types.avrosharp genrejects them with exit code 1 instead of writing one file over the other on Windows and macOS.AVROGEN003 now also reports:
- two types that map to the same C# name (
-m a:M -m b:Mwitha.Xandb.X); - a type whose C# name is also a namespace (
app.eventsnext toapp.events.Click); - an
AvroSharpNamespacethat is not a C# namespace.
avrosharp gen --namespacerejects such a namespace as a usage error (exit code 2).- two types that map to the same C# name (
Two schema files that need each other's types are reported as a circular reference, naming the other file. Before, both errors only said that a type was not defined.
AvroSharp.CodeGencould not be loaded by .NET 8 and 9 applications: its only build, for netstandard2.0, was compiled against the System.Text.Json 10 package, which those runtimes do not have. It now also targets net8.0, which uses the framework's System.Text.Json. The source generator keeps the netstandard2.0 build.
0.1.1 - 2026-09-28
The first complete release of all four packages. 0.1.0's publish stopped partway, so AvroSharp.CodeGen 0.1.0 was never published; use 0.1.1. The code is the same as 0.1.0.
Fixed
- Releasing:
AvroSharp.Generatorsno longer produces a symbol package. It had no.pdbin it, since the generator ships underanalyzers/, and nuget.org's rejection stopped the 0.1.0 publish beforeAvroSharp.CodeGen. The generator's PDB is embedded in its DLL instead. The release workflow now pushes each package separately and checks, while packing, that every symbol package contains a PDB.
0.1.0 - 2026-09-28
The first preview. It covers:
- Schemas: parsing, writing, canonical form and fingerprints.
- The generic data model: binary and JSON encoding, and schema evolution.
- Code generation from
.avscfiles. - Container files with every codec in the specification.
- Messages: single-object messages and schema-registry framing.
- Streams of objects.
Packages: AvroSharp, AvroSharp.Codecs, AvroSharp.CodeGen and AvroSharp.Generators. As a 0.x release, the API may still change before 1.0.
Added
- Nightly fuzzing (
.github/workflows/fuzz.yml): every libFuzzer target runs for 30 minutes a night with SharpFuzz, keeping its corpus between runs and uploading crash inputs. The libFuzzer steps infuzz/README.mdare now verified; a first 20-minute run of all seven targets found no crashes. - Generated readers resolve through a plan built once per writer schema (#69):
Read(ref reader, writerSchema)reads the writer's fields straight into the type's properties, in the writer's order. Fields of the same schema are read directly, numbers (also arrays of them) are promoted in place, enum ordinals are remapped, writer-only fields are skipped and missing fields take their pre-encoded defaults; only other differences (nested records of another version, unions, logical types) are transcoded, one field at a time. On the evolution benchmark this reads in 387 ns, against 506 ns for the generic resolving reader and 1,033 ns for Apache.Avro (i5-3570K); it was 616 ns. AvroValueTransformer.Transform(schema, value, transform): walks a generic value with its schema, through records, arrays, maps and union branches (resolved as the generic writer resolves them), and calls the transform for each non-null primitive, enum or fixed value inside a record field. The transform gets anAvroFieldContext: the record, the field with its properties (for exampleconfluent:tags), the value's own schema, and the field'sFullNameas Confluent's data rules name it. Records, arrays and maps are copied only where a value changes, so a transform that changes nothing returns the same instance. A depth limit (128 by default) stops cyclic records. This is the basis for field-level data rules, encryption and redaction (#82).- Schema-registry wire framing (
AvroSharp.Messages), next to single-object encoding and with no registry client dependency:AvroRegistryFramingfor Confluent (0x00+ 4-byte ID, byte-identical to Confluent's serializer; and the version 10x01+ GUID framing), Apicurio (4- and 8-byte IDs) and AWS Glue (0x03, compression byte, UUID; zlib-compressed payloads read and written),AvroSchemaId,AvroRegistryMessagefor writing, andAvroRegistryMessageReaderfor reading through anIAvroSchemaIdResolver(a synchronous lookup with an asynchronous fill;AvroSchemaIdStorein memory). Read functions are cached per ID with a last-hit check.ConfluentSchemaIdHeaderencodes and decodes Confluent's__key_schema_id/__value_schema_idheader values, read withReadPayload. Hostile input (short or unknown headers, unknown IDs, trailing bytes, corrupt zlib, payloads expanding pastMaxPayloadLength) raisesAvroDataException; aRegistryMessagefuzz target covers it. - Schema references for registries:
AvroSchema.ToJson(referencedSchemas)writes the named types of referenced subjects by name, as Java'sSchema.toString(referencedSchemas, false)does, for a schema parsed against them (AvroSchemaParser.AddNamedSchemas). Checked against Apache Avro Java 1.12 output and against Confluent Schema Registry 7.7, which stores the text unchanged and finds the registered version when it is registered again. - Pipelined container reading:
AvroFileReader<T>.ReadAllPipelinedAsync(blocksAhead)reads and decompresses blocks on a background task, up toblocksAhead(2 by default) ahead of the caller, who decodes them. Blocks pass through a boundedSystem.Threading.Channelschannel, so memory stays bounded, and each block's buffers return to the pool once it is decoded. An error in a block is raised when the caller reaches that block, and stopping the enumeration stops the background task. The netstandard targets referenceSystem.Threading.Channels(#32). - Streams of objects without a container (
AvroSharp.Streams):AvroStreamWriterwrites objects one after another in the binary encoding, buffered (AvroStreamOptions.BufferSize, 64 KiB by default), withWrite/WriteAsync/Flush/FlushAsync.AvroStreamReaderreads them back to the end of the stream withTryRead,ReadAllandReadAllAsync(IAsyncEnumerable<T>), for generated types or as generic values, optionally resolved to a reader schema. Nothing in the encoding delimits objects, so each is decoded to find its end. An object cut off by the buffer is decoded again once more data is read, and a failure is reported only once the stream has ended orMaxDatumLength(64 MiB by default) bytes are buffered, with the object's stream offset. Objects that encode to no bytes cannot be delimited and are rejected (#32). GenericRecord.TryGetValue(int, out AvroValue): field access by position without an exception for a position the record lacks, alongside the name-based overload.AvroSharp.Codecs: the snappy, zstandard, bzip2 and xz codecs in one package, on fully managed libraries: Snappier (plusSystem.IO.Hashingfor snappy's CRC-32), ZstdSharp.Port, SharpZipLib and Lzma.Net.AvroCodecs.Allgives a reader every codec, since a file's codec is not known until it is opened. The writing settings and defaults match Apache Avro Java's: zstandard level 3 with an optional content checksum, bzip2 block size 9, and xz level 6. Snappy blocks end with the big-endian CRC-32 of their data, which is checked on read, and zstandard frames are read with or without a content size or checksum. Damaged blocks, and blocks that decompress pastMaxBlockLength, areAvroDataException. Checked against files written by Apache Avro Java 1.12.2 in every codec (tests/TestData/java-avro, with the script that writes them), and Java reads the files these codecs write. Snappy and bzip2 are also checked against Apache.Avro C#'s codec packages in both directions. When a file uses a standard codec that is not available, the reader's error names this package (#32).- Seeking and splitting container files, as in the Java implementation:
AvroFileReader<T>.PreviousSync(the current block's start),Seek(to a block start),Sync(to the first block after a position, found by scanning for the sync marker) andPastSync(whether the current block belongs to the next split). A split[start, end)is read withSync(start)andwhile (TryRead(out var v) && !PastSync(end)). Positions before the header's marker go to the first block, so a marker in the metadata (Apache'ssyncInMeta.avro) is not mistaken for a block. Needs a seekable stream. - Asynchronous container files:
AvroFileReader.OpenAsync/OpenGenericAsyncandReadAllAsync(IAsyncEnumerable<T>), andAvroFileWriter<T>.WriteAsync/FlushAsync/DisposeAsync; both types implementIAsyncDisposable. The asynchronous paths do no synchronous I/O (tested with a stream that throws on it); objects are encoded into memory and blocks decoded from memory synchronously. Cancellation is checked before every block. The writer now writes the header with its first block, or when flushed or disposed, instead of when created.Microsoft.Bcl.AsyncInterfacesis referenced for netstandard2.0 (it already came in through System.Text.Json). - Single-object encoding (
AvroSharp.Messages):AvroMessagewrites and reads theC3 01marker and the writer schema's CRC-64-AVRO fingerprint, and writes whole messages for generic values or any type with a write delegate.AvroMessageReaderlooks each message's fingerprint up in anIAvroSchemaStore(AvroSchemaStoreis a thread-safe in-memory one), creates the read function once per writer schema, and can resolve every message to one reader schema. Headers that are missing, unknown fingerprints and bytes left after the object areAvroDataException. Checked byte for byte against Java'smessageV1test message. Apache.Avro C# has no single-object API, so there is no benchmark baseline. - Object container files (
AvroSharp.Containers):AvroFileWriterandAvroFileReaderfor generic values or any type with a write or read delegate (generated types pass their staticWrite/Readmethods). The built-in codecs arenullanddeflate(BCL, raw DEFLATE); others plug in throughAvroCodec. Blocks are written at a configurable sync interval, the header carries application metadata, and a write that throws leaves nothing behind. The reader resolves to an optional reader schema, verifies each block's sync marker, and bounds hostile input: block and metadata sizes are limited (MaxBlockLength, 64 MiB by default, also applied to decompressed data), object counts are checked against the block's bytes, and bytes left after a block's last object are an error. Checked against Apache'sweather.avro,weather-sorted.avro(deflate) andsyncInMeta.avro, and against Apache.Avro in both directions with random schemas and data. - Schema resolution for generated types:
Read(ref reader, writerSchema)andFromAvroBytes(data, writerSchema)read data written with another version of the type's schema. Data of the same canonical schema takes the direct path; other data is transcoded, following the resolution rules, straight into the type's own encoding in a reused per-thread buffer and read from there, without creating generic values. On the evolution benchmark this reads 1.65x faster than Apache.Avro's resolving reader with 29% of its allocations (i5-3570K). - Schema resolution for the generic model:
GenericDatumReader.Create(writerSchema, readerSchema)reads data written with one schema version as another, following the specification. Record fields are matched by name or alias, writer-only fields are skipped (sized blocks in one step), and missing reader fields take their defaults. Named types match by full name, unqualified name or alias. Numbers are promoted,stringandbytesconvert, and enum symbols are matched with the reader's default for unknown ones. Unions resolve per branch. Incompatible schemas are rejected when the reader is created; a mismatch that only some data would hit (a union branch, an unknown enum symbol without a default) is reported when such a value is read. AvroSchemaParseOptions.AllowIdenticalRedefinitions: a parser may accept a named type that an earlierParsecall defined, when both definitions have the same canonical form; the first definition stays in use. The source generator turns it on, so schema sets that inline shared types in every file (as Apache's one-file-at-a-time tooling requires) generate each type once. Different definitions are still an error that names both files.- Generator option
AvroSharpPropertyNames=avro(CodeGenOptions.PropertyNaming): keep the Avro field names as property names, as Apache'savrogendoes, so code written against avrogen classes compiles unchanged. C# keywords are escaped. - Apache.Avro compatibility mode for generated code (
AvroSharpApacheCompatible=true, requires a reference to Apache.Avro): records also implementAvro.Specific.ISpecificRecord, fixed types derive fromAvro.Specific.SpecificFixed, and logical types use Apache's .NET types, so Apache'sSpecificDatumWriter<T>/SpecificDatumReader<T>and AvroSharp's serializers work on the same classes and produce the same bytes.Putalso accepts what Apache's reader passes: an enum's ordinal, and anAvroDecimalfor a decimal on fixed. The generator reportsAVROGEN004when the property is set without the reference. - Logical types in generated code:
datebecomesDateOnlyandtime-millis/time-microsbecomeTimeOnly(DateTime/TimeSpanwhere those types don't exist);timestamp-millis/timestamp-microsbecomeDateTimeOffset, and thelocal-timestampvariants becomeDateTime;uuid(onstringorfixed(16)) becomesGuid;decimalwith a precision up to 28 becomesdecimal.- Set the MSBuild property
AvroSharpLogicalTypes=rawto keep the underlying types. - The conversions are public in
AvroLogicalValues:- decimals are exact, raising an error instead of rounding;
- times and timestamps are truncated towards negative infinity to the logical type's precision;
- out-of-range data raises
AvroDataException.
- Generated records implement
IAvroSpecificRecord:Schema,Get(int)andPut(int, object?), field access by position following the contract of Apache.Avro'sISpecificRecordwithout depending on it.Putchecks the value's type (no implicit widening, as with Apache's casts) and names the field in errors. - Code generation from schema files: the
AvroSharp.Generatorssource generator (an incremental generator for.avscfiles passed asAdditionalFiles) and theAvroSharp.CodeGenengine it uses. Records become partial classes with staticWrite/Readmethods (plusToAvroBytes/FromAvroBytes) that callAvroWriter/AvroReaderdirectly in schema order; enums become C# enums; fixed types become size-checked wrappers. Schema files may refer to each other's named types; errors are reported at the file, line and column. Generated readers enforce the same hostile-input limits as the generic reader. The generator needs the .NET 10 SDK or Visual Studio 2026. - JSON encoding for the generic model:
GenericDatumJsonWriterandGenericDatumJsonReader, following the specification (wrapped union values, byte strings forbytesandfixed, enum symbols). Record fields may appear in any order and missing fields take their defaults; NaN and infinities are written as strings, as Apache.Avro C# does. Checked both ways against Apache.Avro'sJsonEncoder/JsonDecoder. - Fuzz targets (
fuzz/AvroSharp.Fuzz, SharpFuzz/libFuzzer) for schema parsing and generic binary and JSON data, container files, single-object messages and schema resolution (the transcoder checked against the resolving reader), each also checking round trips. They run on every build as a seeded mutation smoke test. - Generic data model (
AvroSharp.Generic):AvroValue, a 16-byte struct holding any Avro value without boxing (primitives inline, enums as schema plus ordinal),GenericRecordandGenericFixed.GenericDatumWriterandGenericDatumReadercompile a schema once into a cached, thread-safe plan of typed nodes; union branches are selected from the value's kind or schema name. Arrays ofint/long/float/doubleare read in bulk. Hostile input is bounded: block counts are checked against the remaining input (using each record's minimum encoded size), pre-allocation is capped, zero-size items draw from a per-read budget, and record nesting is limited when reading and writing (GenericDatumReaderOptions,GenericDatumWriterOptions). AvroReader.ReadLongs/ReadInts: bulk varint reads; on net8+ aVector128check decodes runs of one-byte values 16 at a time.- Binary encoding (
AvroSharp.IO):AvroWriterwrites directly into anIBufferWriter<byte>or aSpan<byte>;AvroReaderreads from aReadOnlySpan<byte>or a multi-segmentReadOnlySequence<byte>, returning slices of the input forbytes,stringandfixedwhen contiguous. Bulkdouble/floatarray items are a single copy on little-endian hardware. Length prefixes are checked against the remaining input before any allocation; malformed data raisesAvroDataException. - Schema model (
AvroSharp.Schemas): immutable primitive, record, enum, array, map, union and fixed schemas; names, namespaces and aliases; record fields with defaults, order and aliases; custom properties. - All Avro 1.12 logical types. Unknown or invalid logical types are ignored and kept as properties, as the specification requires.
AvroSchemaParserandAvroSchema.Parse/ParseAsync:System.Text.Jsonparser over UTF-8 with default-value validation, optional comments, and errors that report the JSON path, line and column. A parser keeps named types across calls, so schemas split over several files can refer to each other.- Full schema JSON writer, Parsing Canonical Form, and CRC-64-AVRO, MD5 and SHA-256 fingerprints.
- Tests against Apache Avro's
schema-tests.txtvectors, property-based interop tests against Apache.Avro (C#), a Native AOT smoke test, and schema-parse benchmarks gated against Apache.Avro. AvroCodecNames: the codec names defined by the specification.- Repository skeleton: build settings, analyzers, public API tracking, strong naming, TUnit tests on .NET 8/9/10 and .NET Framework 4.8.1, and CI on Linux and Windows (x64 and Arm64), including .NET Framework 4.8.1 on Windows (#72).
Changed
AvroSchema.ToJson()now gives the text of Apache Avro Java'sSchema.toString(): a named type'saliasescome after its custom properties, numbers with a fraction or exponent in defaults and properties are printed as Java prints them (1e10becomes1.0E10, computed exactly on every runtime), and only",\and control characters are escaped. The parsed schema and its canonical form are unchanged.AvroWriter.WriteStringencodes in one pass when the buffer has room for the longest possible encoding. It reserves the prefix for 3 bytes per char, encodes, then writes the real length, moving the bytes down when the prefix is shorter. Otherwise it counts first, as before, so a fixed-size span destination that holds exactly the encoding still works. On the string encode benchmark (i7, ShortRun), it went from 0.88x to 0.83x Apache.Avro's time (#25).- Container files use their streams less:
- The writer writes each block, with its count, size and sync marker, in one stream write instead of three. That is one call or one await per block. Its block buffer is sized from the sync interval (#62).
- The reader keeps its block buffer between blocks while it is large enough, reuses the buffer that limits decompressed size, and decodes objects from an array segment (#62).
PastSyncchecks the stream's length once per block, not for every object of the documented split loop (#62).- A header that needs several fills is still re-scanned, but an attempt that runs out of data no longer allocates. Metadata is recorded as positions and turned into strings and copies once, when the header is complete (#70).
- The container benchmarks and the Native AOT smoke test cover every codec. The benchmark baselines are Apache.Avro's codec packages (#32).
- Reading generated types with a writer schema (
Read(ref reader, writerSchema), as container files and single-object messages do for every object) no longer compares the two schemas' canonical forms for every record. A writer schema remembers the last schema found to have its canonical form, so repeated checks compare references (#60). - Resolving records whose reader field order differs from the writer's no longer allocates an array per record; the slots are rented from the shared array pool (#61).
AvroMessageReaderchecks the last schema used before its fingerprint dictionary, which saves the dictionary lookup when messages repeat one schema (#63).- Each package ships its own README (
AvroSharp.CodeGenandAvroSharp.Generatorsno longer show the repository README), and all three includeTHIRD-PARTY-NOTICES.md(#66).
Fixed
THIRD-PARTY-NOTICES.mdsaid the packages contain no third-party code, but they compile in source from Polyfill (MIT). The notice now includes Polyfill's copyright and license, and is packed into every package (#66).- Generated code needed C# 9 (
new()initializers,??=,is { }andis notpatterns), so it failed to compile in netstandard2.0 and .NET Framework projects, which default to C# 7.3. It now uses constructs every version accepts, and emits nullable annotations only for C# 8 and later. - Invalid UTF-8 inside a JSON string (schema JSON or JSON data) raised
InvalidOperationExceptioninstead ofAvroSchemaException/AvroDataException. Found by the fuzz smoke test.