https://youtrack.jetbrains.com/issue/KT-60003/K2-Disappeared-INVALIDCHARACTERSNATIVEERROR ^KT-60003 Fixed Merge-request: KT-MR-12686 Merged-by: Vladimir Sukharev <Vladimir.Sukharev@jetbrains.com>
13 KiB
FIR Checkers
Checkers structure
There are six kinds of checkers:
- DeclarationChecker
- ExpressionChecker
- FirTypeChecker
- FirLanguageVersionSettingsChecker
- FirControlFlowChecker
The first three kinds are typed and may be restricted to checking only a specific type of declaration/expression/type ref. To simplify working with checkers for different FIR elements, there is a number of typed typealiases:
- Declarations: FirDeclarationCheckerAliases.kt
- Expressions: FirExpressionCheckerAliases.kt
- Type refs: FirTypeCheckerAliases.kt
The next kind, FirLanguageVersionSettingsChecker, is to check language version settings independently of particular code pieces.
The last kind of checker, FirControlFlowChecker, is for checkers which perform Control Flow Analysis (CFA) and is supposed to work with every declaration that has its own Control Flow Graph (CFG)
Checkers contracts
All checkers are supposed to satisfy the following contracts:
- checkers are stateless
- checkers are independent
- checkers are as specific as possible
- checkers should try to avoid traversing the subtree of the element it checks
- checkers should not rely on the syntax
Those contracts imply the following:
- Usually a checker is an
objectwithout any state - Each checker should work correctly even if all other checkers are disabled
- If a checker is meant to check only simple functions, there is no need to parameterize it with
FirDeclarationand check if the declaration is aFirSimpleFunction. Just parameterize the checker itself withFirSimpleFunction- this is needed not only for simplification of code, but also for the sake of performance. Typed checkers are run only on elements with a suitable type. So if you declared a
FirRegularClassCheckerit will never be run for aFirAnonymousObject
- this is needed not only for simplification of code, but also for the sake of performance. Typed checkers are run only on elements with a suitable type. So if you declared a
- If a checker is supposed to check anonymous initializers, it's better to create a
FirAnonymousInitializerCheckerwhich will be separately run for eachinitblock in the class rather than creating aFirClassCheckerwhich will manually iterate over eachinitblock in this class. There are several reasons for that:- the diagnostic suppression mechanism is implemented in the checkers dispatcher, so reporting something on a sub-element may cause false-positive diagnostics, if there was a
@Suppressannotation between the root element (passed to the checker) and the sub-element. There is a mechanism to fix it, but it's not recommended to use - checkers with smaller scope increase IDE performance because they require fewer elements to be resolved in order to check something
- the diagnostic suppression mechanism is implemented in the checkers dispatcher, so reporting something on a sub-element may cause false-positive diagnostics, if there was a
- FIR compiler is made syntax-agnostic and can work with different parsers and syntax tree (at this moment it already supports PSI and LightTree syntax trees), so checkers should not rely on any syntax implementation details. Instead of that, checkers should use positioning strategies to more precise positioning of diagnostics for specific elements (e.g. it allows to render diagnostic on class name using the source of the whole class). The only exception from this rule are inheritors of FirSyntaxChecker, which work directly with a syntax tree (and must support several implementations for different ASTs)
Checkers pipeline
All checkers are collected in special containers, named DeclarationCheckers, ExpressionCheckers and TypeCheckers. Those containers have fields with sets of checkers for each possible type of checker of corresponding kind
There is a number of different container groups:
- Common checkers, which always run on any platform
- Checkers for each specific platform (lay in the corresponding
:compiler:fir:checkers:checkers.platformmodules) - Extended checkers. Those checkers are disabled by default and can be enabled with the
-Xuse-fir-extended-checkerscompiler flag. This group includes experimental and not very performant checkers, which are not crucial for regular compilation
At the beginning of the compilation, in the initialization phase, all required checker containers are collected inside a session component named CheckersComponent. When the time of checker phase comes, the compiler creates an instance of AbstractDiagnosticCollector, which is responsible to run all checkers. DiagnosticCollector traverses the whole given FIR tree, collects CheckerContext during this traversal, and runs all checkers that suite the element type on each element.
Checker Context
CheckerContext contains all information which can be used by checkers, including
sessionandscopeSession- the list of
containingDeclarations - various information about the body which is analyzed
- the stack of implicit receivers
- information about suppressed diagnostics
CheckerContext is meant to be read-only for checkers
Diagnostic reporting
All diagnostics which can be reported by the compiler are stored within the FirErrors, FirJvmErrors, FirJsErrors and FirNativeErrors objects. Those diagnostics are auto-generated based on the diagnostic description in one of a diagnostic list in checkers-component-generator.
The generation is needed, because Analysis API (AA), which is used in IDE, generates a separate class for each compiler diagnostic with proper conversions of arguments for parametrized diagnostics. And the goal of the code generator is to automatically generate those classes and conversions. To run the diagnostics generation use the Generators -> Generate FIR Checker Components and FIR/IDE Diagnostics run configuration.
Diagnostic messages must be added manually to FirErrorsDefaultMessages, FirJvmErrorsDefaultMessages, FirJsErrorsDefaultMessages and FirNativeErrorsDefaultMessages respectively. Guidelines for diagnostic messages are described in the header of FirErrorsDefaultMessages
To report diagnostics, each checker takes an instance of DiagnosticReporter as a parameter. To reduce the boilerplate needed to instantiate a diagnostic from the given factory and ensure it's not missed due to reporting on the null source, a one should use the utilities from KtDiagnosticReportHelpers
FIR contracts at checker stage
In CLI mode the compiler runs checkers only after it has analyzed the whole world up to the final FIR phase (BODY_RESOLVE). But the IDE uses lazy resolve, so there can be a situation when some files have been analyzed to BODY_RESOLVE and other files have not been analyzed at all. This means that in a checker one can not rely on the fact that some FIR elements should have been resolved to some specific phase. The only exception is the following: If some element was passed directly to the checker then it is guaranteed that this element is already resolved to the BODY_RESOLVE phase. If some declaration is received somewhere from outside (from a type, a symbol provider or a scope), then it could have been resolved up to an arbitrary phase.
So, to avoid possible problems with accessing some information from FIR elements which was not yet calculated in the AA mode, there are the following restrictions and recommendations:
- Access to
FirBasedSymbol<*>.firis prohibited. One can not extract any FIR element from the corresponding symbol - Instead of that, if some information about the declaration is needed (e.g., the list of supertypes for some class symbol), special accessors from that symbol should be used (they are declared as members of symbols). Those accessors call lazy resolution to the least required phase and after that extract the required information from FIR
Resolution diagnostics
While all checkers are run after resolution of the code is finished, some diagnostics can be actually detected only during resolution, such as
- inference errors (type mismatch, no information for type parameter)
- call resolution errors (overload resolution ambiguity)
- type resolution errors (cycle in supertypes)
- visibility errors (invisible reference)
- etc
And at the same time, there is a contract that FIR resolution is side effect free (not very formal but still) and produces only a resolved FIR tree. So diagnostics can not be reported from resolution directly.
To support such diagnostics, there is the following mechanism:
- some FIR nodes (mostly with word
Errorin name, like FirResolvedErrorReference) have a property which contain aConeDiagnostic - ConeDiagnostic is an indicator that something went wrong during resolution
- there are a lot of different kinds of
ConeDiagnosticfor any possible problems, see ConeDiagnostics.kt
- there are a lot of different kinds of
ConeDiagnosticis saved in the FIR tree, and then the special checker component (ErrorNodeDiagnosticCollectorComponent) checks all FIR nodes and report proper diagnostics based on the foundConeDiagnostic