Run these two statements against any Neo4j database:
CREATE (:Product {price: 9.99});
CREATE (:Product {price: '9.99'});Both succeed. You now have two Product nodes that disagree about what price is - one stores a FLOAT, the other a STRING. Nothing complains until you try to use the value: avg(p.price) fails on the string, and a filter like WHERE p.price > 10 silently drops the mismatched node, because comparing a string to a number in Cypher evaluates to null.
Why you need to care about property types
Neo4j is schema-optional by design, and most of the time that flexibility is a feature. The labels, relationship types, and property types you rely on form your graph's ontology - but by default that ontology lives in convention and application code, not in the database. A property's type is whatever the last write said it was, per node - Neo4j does not enforce types by default, drivers faithfully serialise whatever they are given, and compile-time types in your application do not exist at runtime.
As an example, the LOAD CSV command loads every value into memory as a string before processing. If you do not consistently cast the values to their original type, you may end up with a mix of types in a single property.
What Cypher gives you
Cypher has three tools that enforce consistency, all part of the same type-system work that landed across the Neo4j 5 releases:
- Type predicate expressions (
value IS :: STRING), introduced in Neo4j 5.9, test the type of any value in a query. - The
valueType()function, added in Neo4j 5.13, returns the type of a value as a string, so you can audit what is actually stored. - Property type constraints, also introduced in Neo4j 5.9 (Enterprise Edition, including Aura), reject any write that stores the wrong type.
The first two can be used to identify problems. Adding a constraint prevents them happening in the first place.
Checking values
To find every product where price is a string, use a type predicate in the WHERE clause:
MATCH (p:Product)
WHERE p.price IS :: STRING
RETURN p.sku, p.priceThe predicate works on any expression, supports the full Cypher type system - INTEGER, FLOAT, STRING, BOOLEAN, temporal types, POINT, and parameterised list types like LIST<STRING> - and has a negated form, IS NOT ::.
One behaviour catches almost everyone the first time: type predicates treat types as nullable by default, so null IS :: STRING evaluates to true. A query for wrong-typed properties therefore also matches every node where the property is missing - on a large dataset, that turns "find the handful of bad values" into "match nearly everything". Exclude nulls with an explicit guard, or - since Neo4j 5.10 - by marking the type as non-nullable:
// equivalent predicates
WHERE p.price IS NOT NULL AND p.price IS :: STRING
WHERE p.price IS :: STRING NOT NULLRather than testing types one by one, valueType() tells you what is stored:
UNWIND [
1, // (1) an integer
'true', // (2) a string
true, // (3) a boolean
null, // (4) a null value
[true, false] // (5) a list of booleans
] AS value
RETURN value, valueType(value) // (6)- 1
1is an integer - 2
'true'is a string - notice the single quotes - 3
truewithout the single quotes is a boolean - 4
nullis a null value - 5
[true, false]is a list of booleans - 6
valueType()returns the most specific type of each value as a string
The Cypher statement returns the following values:
| value | valueType(value) |
|---|---|
1 | INTEGER NOT NULL |
'true' | STRING NOT NULL |
true | BOOLEAN NOT NULL |
null | NULL |
[true, false] | LIST<BOOLEAN NOT NULL> NOT NULL |
The NOT NULL suffix means the value itself is not null - a concrete value never is, so only null reports the plain NULL type. Combined with count(), the function makes a one-query audit of any property:
MATCH (p:Product)
WHERE p.price IS NOT NULL
RETURN valueType(p.price) AS type, count(*) AS countIf the query returns more than one row, you have a drift in your property types, and the fix is usually a small, targeted SET:
MATCH (p:Product)
WHERE p.price IS :: STRING NOT NULL
SET p.price = toFloat(p.price)toFloat() returns null for values it cannot parse, and setting a property to null removes it - so check for unparseable values first if losing them matters. Re-running the statement matches nothing and changes nothing, which is the property you want in any data migration.
Setting up a constraint to check automatically
Checking values fixes the past. A property type constraint fixes the future:
CREATE CONSTRAINT product_price_type IF NOT EXISTS
FOR (p:Product)
REQUIRE p.price IS :: FLOATFrom this point on, any write that tries to store a non-float price fails at the database, whatever the writer - the application, a backfill script, a colleague in the Query tool. That part of your ontology is now understood and enforced by the database itself, not left to every writer to remember.
Nodes without the property still pass, because type constraints do not imply existence. The NOT NULL modifier from type predicates is not valid in a constraint - to require the property as well, add a separate property existence constraint:
CREATE CONSTRAINT product_price_exists IF NOT EXISTS
FOR (p:Product)
REQUIRE p.price IS :: FLOAT NOT NULLCreating a constraint will fail while violating data exists. Run any cleanup scripts before creating the constrant.
Learn more
If the type of a property is part of your graph's ontology, say so in the schema. Constraints, indexes, and the rest of Neo4j's schema toolkit are covered in the Cypher Indexes and Constraints course on GraphAcademy:
Cypher Indexes and Constraints
Comments (0)
Loading comments...