A method and system for automatically generating a report of a water quality chemical analysis device

CN122452505APending Publication Date: 2026-07-24GUANGDONG SPECIAL EQUIP TESTING INST DONGGUAN TESTING INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG SPECIAL EQUIP TESTING INST DONGGUAN TESTING INST
Filing Date
2026-06-24
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as insufficient accuracy of semantic mapping of natural language requests, lack of compliance analysis logic, and inability of report generation templates to dynamically adapt to query confidence levels, resulting in low report generation efficiency and insufficient professionalism and credibility.

Method used

By extracting multi-dimensional conversation features from natural language analysis requests to generate conversation analysis focus vectors, using domain knowledge graphs for entity ambiguity resolution and pattern linking, and combining compliance judgment logic, a multi-dimensional probability scoring model is constructed to evaluate query quality. Furthermore, a linkage mechanism between query confidence and report generation strategy is established to dynamically adjust the level of detail and tone of the report content.

Benefits of technology

It enables the intelligent, professional, and standardized generation of water quality chemical analysis reports, improving the efficiency, professionalism, and credibility of report generation, accurately understanding user query intent, reducing entity mapping bias, and ensuring that query logic conforms to water quality analysis standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122452505A_ABST
    Figure CN122452505A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of report automatic generation, and particularly relates to a report automatic generation method and system of water quality chemical analysis equipment, which comprises the following steps: receiving a natural language analysis request, extracting multi-dimensional conversation features to generate a conversation analysis focus vector; extracting a request entity, and searching a candidate database mode guided by the vector; performing entity ambiguity resolution and mode linking based on correlation degree features to generate a linking confidence; deconstructing the request into an execution sequence containing compliance judgment logic, evaluating logical rigor to generate a structure confidence; generating a candidate query statement based on the execution sequence and the linked database mode, and calculating a comprehensive confidence by combining the linking confidence and the structure confidence; executing a query according to the query statement selected according to the comprehensive confidence, dynamically regulating a report generation strategy, and generating a report by combining the query result and the conversation analysis focus vector. The present application can improve the efficiency, professionalism and reliability of water quality chemical analysis report generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automatic report generation technology. More specifically, this invention relates to a method and system for automatically generating reports for a water quality chemical analysis device. Background Technology

[0002] With increasing awareness of ecological and environmental protection and higher requirements for water quality from industrial production, water chemical analysis has become an important tool for environmental monitoring, drinking water safety, and watershed pollution control. my country has established a comprehensive monitoring network covering major river basins, key lakes and reservoirs, and drinking water sources, generating massive amounts of multi-dimensional chemical indicator data daily. The standardized analytical reports generated are a crucial foundation for water management and environmental decision-making.

[0003] Traditional report generation heavily relies on technical professionals, requiring manual extraction of monitoring data for specific time periods, compliance assessments, and conclusion writing. This process is time-consuming, labor-intensive, and prone to bias due to subjective experience. Furthermore, a significant semantic gap exists between the natural language query needs of non-professional users and the underlying structured database. Conventional retrieval systems cannot accurately understand varied, colloquial requests, resulting in consistently low utilization rates of water quality data.

[0004] To address these issues, existing technologies have begun exploring intelligent report generation solutions. For example, Chinese patent document CN113673970B discloses a water quality report generation method based on distributed nodes. This method collects video data and extracts water quality parameters through distributed nodes in a watershed, then calculates the water quality level using weights to generate a report. This solution automates the report generation process but only supports limited visual indicator analysis, lacks natural language interaction and built-in compliance judgment capabilities, and the report content and format rely entirely on fixed templates. Another example is Chinese patent application CN116108198A, which discloses a water quality diagnosis method based on big data AI to construct a knowledge graph. This method integrates multi-source data through knowledge graph construction to achieve automatic inspection and global analysis of water quality anomalies. However, this application lacks a natural language interaction interface, has insufficient ability to resolve ambiguities related to professional entities, cannot establish an accurate mapping between natural language expressions and database schemas, and cannot deconstruct analysis requirements into a standardized execution sequence containing compliance logic, making it difficult to meet the professional and standardized requirements of water quality chemical analysis reports.

[0005] Furthermore, most current automatic report generation methods employ static template mechanisms, failing to establish a linkage between query accuracy and report quality. When entity links or query structures present uncertainty, the system still outputs the same level of certainty, unable to dynamically adjust the level of detail and tone of the report content. This can easily mislead decision-making and hinders the practical application of intelligent report generation technology.

[0006] To address the aforementioned technical deficiencies, this invention provides a method and system for automatically generating reports for water quality chemical analysis equipment. The aim is to solve problems such as low natural language semantic matching, lack of compliant analysis logic, rigid report templates, and inability to adapt to query confidence levels, thereby achieving intelligent, professional, and standardized report generation. Summary of the Invention

[0007] To address the issues of insufficient semantic mapping accuracy of natural language requests, lack of compliance analysis logic, and inability of report generation templates to dynamically adapt to query confidence levels in the automatic generation of water quality chemical analysis reports in existing technologies, this invention provides solutions in the following aspects.

[0008] In a first aspect, the present invention provides a method for automatically generating reports from a water quality chemical analysis device, comprising: S1, receiving a natural language analysis request, extracting multi-dimensional conversation features of the analysis request to generate a conversation analysis focus vector; extracting request entities from the analysis request, and, guided by the conversation analysis focus vector, retrieving candidate database patterns associated with the request entities in a domain knowledge graph; S2, calculating the correlation features between the candidate database patterns and the conversation analysis focus vector, performing entity ambiguity resolution and pattern linking based on the correlation features, and generating link confidence; and then... The analysis request is deconstructed into a logical execution sequence containing compliance judgment logic, and the logical rigor of the logical execution sequence is evaluated to generate structural confidence; S3, based on the logical execution sequence and the candidate database pattern that completes the pattern link, candidate query statements are generated, and the comprehensive confidence of the candidate query statements is calculated by combining the link confidence and the structural confidence; S4, a target query statement is selected to perform data query based on the comprehensive confidence, and the report generation strategy is dynamically adjusted based on the comprehensive confidence, and a water quality chemical analysis report is generated by combining the query results and the session analysis focus vector.

[0009] The core principle of this invention lies in the following: First, the spatiotemporal features, historical preferences, and analytical intent of the natural language analysis request are extracted and fused to generate a conversation analysis focus vector. This vector guides the retrieval of candidate database patterns in the knowledge graph. Then, semantic consistency and connection path weights are used to resolve entity ambiguities and link patterns. Simultaneously, the request is deconstructed into a sequence of operators with built-in compliance judgment logic, and its logical rigor is evaluated, generating link confidence and structural confidence respectively. Finally, the comprehensive confidence of the candidate query statements is calculated by combining the two confidence levels. Based on this confidence level, the target query statement is selected for execution, and the report generation strategy is dynamically adjusted. This solution accurately understands the user's query intent, reduces entity mapping bias, ensures that the query logic conforms to professional water quality analysis standards, and adjusts the level of detail and tone of the report content according to the reliability of the query, thereby generating a water quality chemical analysis report with standardized content and reliable conclusions.

[0010] Preferably, extracting multidimensional session features of the analysis request to generate a session analysis focus vector includes: extracting spatiotemporal context features, user historical preference features, and analysis intent features of the analysis request; performing feature concatenation and dimensionality reduction fusion on the spatiotemporal context features, the user historical preference features, and the analysis intent features to generate the session analysis focus vector.

[0011] By combining the explicit analytical intent of a user's request with implicit spatiotemporal context and historical operation preference features, a unified session analysis focus vector is generated. This vector can comprehensively capture the user's true needs, avoid the comprehension bias caused by relying on a single intent feature, and provide clear guidance for subsequent entity retrieval, pattern linking, and report framework construction, thereby improving the targeting and accuracy of the entire process.

[0012] Preferably, extracting the spatiotemporal context features and user historical preference features of the analysis request includes: parsing the analysis request to extract the timestamp and geographic coordinates, converting the timestamp into a standard time value and the geographic coordinates into a latitude-longitude floating-point pair, normalizing the standard time value and the latitude-longitude floating-point pair, and then combining them to construct the spatiotemporal context features; querying the user's historical operation logs, extracting the water quality indicator parameters with the highest search frequency, and vectorizing them as the user historical preference features.

[0013] By transforming unstructured timestamps and geographic coordinates into standardized numerical features, a unified representation of spatiotemporal information is achieved. Furthermore, by statistically analyzing high-frequency retrieval indicators in users' historical operation logs, water quality parameters of long-term user interest are identified. This method accurately extracts key contextual information from user requests, providing a reliable feature foundation for generating focus vectors in session analysis.

[0014] Preferably, the process involves extracting request entities from the analysis request and retrieving candidate database patterns associated with the request entities in the domain knowledge graph, guided by the session analysis focus vector. This includes: segmenting the natural language analysis request using a bidirectional maximum matching algorithm; extracting water quality indicators, monitoring locations, and time ranges from the segmentation results using a pre-trained named entity recognition model as the request entities; calculating the similarity between the request entities and standard entity names in the domain knowledge graph; retaining standard entity names with similarity higher than a preset threshold to establish a candidate entity set; and mapping this set to the underlying database to form the candidate database patterns.

[0015] Preferably, entity ambiguity resolution and pattern linking based on the correlation features between the candidate database pattern and the session analysis focus vector include: extracting the node vector corresponding to the candidate database pattern in the domain knowledge graph, calculating the cosine similarity between the node vector and the session analysis focus vector to obtain a semantic consistency score; calculating the shortest hop count from the request entity to the candidate database pattern node in the domain knowledge graph, and converting the shortest hop count into a connection path weight score according to a preset distance decay function, wherein the semantic consistency score and the connection path weight score are used as the correlation features; performing a weighted summation of the correlation features using a weighted formula to obtain a pattern link comprehensive score, and selecting the node with the highest pattern link comprehensive score to establish a unique mapping relationship to complete the pattern linking.

[0016] The relevance of candidate patterns is comprehensively evaluated by simultaneously calculating the semantic consistency score between candidate pattern nodes and focus vectors, as well as the connection path weight score between entities in the knowledge graph. This method considers both semantic matching and the topological relationships of the knowledge graph, effectively resolving ambiguity issues related to specialized entities and establishing an accurate mapping between natural language representations and underlying database schemas.

[0017] Preferably, the generation of link confidence includes: counting the total number of candidate database patterns in the domain knowledge graph that have a connection relationship with the requesting entity and whose hop count is within a preset range; calculating a basic quotient by dividing the pattern link comprehensive score of the node that establishes the unique mapping relationship by the sum of the pattern link comprehensive scores of all the candidate database patterns; and mapping the basic quotient through a nonlinear activation function.

[0018] By statistically analyzing the total number of candidate pattern nodes for the requested entity within a preset range, the comprehensive score of the target node is transformed into a relative advantage value, which is then mapped to a standardized probability value using a nonlinear activation function. This method objectively characterizes the deterministic level of the entity linking process, providing important feature parameters for the comprehensive quality evaluation of subsequent query statements.

[0019] Preferably, the comprehensive confidence of the candidate query statement is calculated by combining the link confidence and the structure confidence, including: extracting the feature vector of the candidate database pattern actually accessed by the candidate query statement as the query database pattern feature vector; calculating the vector inner product between the query database pattern feature vector and the session analysis focus vector as the pattern matching feature value; and inputting the link confidence, the structure confidence, and the pattern matching feature value as feature parameters into a pre-trained probability scoring model, and having the probability scoring model output the comprehensive confidence of the candidate query statement.

[0020] Preferably, the probability scoring model adopts a gradient boosting decision tree model, and the probability scoring model outputs the comprehensive confidence of the candidate query statement, including: inputting the feature parameters sequentially into the gradient boosting decision tree model, accumulating the weights of the leaf nodes output by multiple decision trees in the model, and obtaining the comprehensive confidence with the value compressed within a preset interval through the mapping transformation of the regression function.

[0021] Preferably, the water quality chemical analysis report is generated based on the comprehensive confidence level dynamic adjustment report generation strategy, combined with the query results and the session analysis focus vector, including: when the comprehensive confidence level is greater than a preset confidence level threshold, a detailed analysis paragraph containing the data deduction process is generated in the water quality chemical analysis report, and an explanatory text is generated by calling a deterministic tone vocabulary; when the comprehensive confidence level is less than or equal to the confidence level threshold, only the core indicator data results are extracted in the water quality chemical analysis report, and an explanatory text is generated by calling a conservative tone vocabulary containing probabilistic or advisory statements.

[0022] By establishing a linkage mechanism between comprehensive confidence level and report generation strategy, the level of detail and tone of the report content are dynamically adjusted according to the reliability of the query. When the confidence level is high, in-depth analysis and definitive conclusions are provided, while when the confidence level is low, only core data is presented and conservative statements are used. This effectively reduces the risk of being misled by low-quality queries and improves the practicality and credibility of the report.

[0023] Secondly, the present invention provides an automatic report generation system for a water quality chemical analysis device, comprising a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned automatic report generation method for a water quality chemical analysis device is implemented.

[0024] By adopting the above technical solution, the above-mentioned method for automatically generating reports of a water quality chemical analysis device is used to generate a computer program and store it in a memory so that it can be loaded and executed by a processor. This allows for the creation of a terminal device based on the memory and processor, making it convenient to use.

[0025] The technical solution of the present invention has the following beneficial technical effects: This invention employs a technique that integrates spatiotemporal context, historical preferences, and analytical intent to generate session analysis focus vectors. It combines this with a water quality analysis knowledge graph for entity ambiguity resolution and database schema linking. The analysis request is deconstructed into a standardized execution sequence containing compliance judgment logic. A multi-dimensional probability scoring model is constructed to evaluate query quality, and a linkage mechanism between query confidence and report generation strategy is established. This allows for accurate understanding of users' natural language analysis needs, automatic data retrieval and compliance judgment, and dynamic adjustment of report content detail and tone based on query reliability. This invention effectively solves the problems of low semantic matching degree of natural language requests, lack of built-in compliance analysis logic, static and rigid report templates, and inability to adapt to query confidence in existing technologies, significantly improving the efficiency, professionalism, and credibility of water quality chemical analysis report generation. Attached Figure Description

[0026] Figure 1 A flowchart of a method for automatically generating reports from a water quality chemical analysis device; Figure 2 This is a schematic diagram of path distance attenuation; Figure 3 This is a schematic diagram comparing ablation experiments. Detailed Implementation

[0027] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0028] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0029] This invention discloses a method for automatically generating reports from a water quality chemical analysis device, referring to... Figure 1 This includes the following steps: S1. Generate session analysis focus vectors and retrieve candidate database patterns.

[0030] The purpose of this step is to parse the natural language analysis request input by the user and fuse multi-dimensional features. The goal is to improve the semantic matching accuracy between the natural language request and the underlying structured water quality database, and to obtain guiding features that can represent user intent and business context.

[0031] In the specific execution process, the multi-dimensional session features of the analysis request are first extracted to generate a session analysis focus vector. Optionally, the system parses the time representation and geographic location information in the analysis request, converting the time representation into standard time values, such as Unix millisecond-level timestamps, and converting the geographic location information into latitude and longitude floating-point pairs, for example, by performing coordinate system transformation based on the WGS-84 standard. To eliminate differences between different physical dimensions, the above values ​​are normalized, and the time field is decomposed into year, month, day, and hour dimensions and combined with latitude and longitude to construct a 6-dimensional numerical spatiotemporal context feature.

[0032] Simultaneously, the system queries the user's historical operation logs. For example, it uses a word frequency-inverse document frequency algorithm to statistically extract water quality indicator parameters with high search frequency in the past three months, such as chemical oxygen demand (COD), ammonia nitrogen, total phosphorus, total nitrogen, pH, or dissolved oxygen. These parameters are then mapped into 256-dimensional vectorized data using a pre-trained Word2Vec or BERT word embedding model, and after average pooling, are used as 256-dimensional user historical preference features. Furthermore, the system uses a natural language intent classification model, such as a fine-tuned BERT model, to classify the analysis requests and extract the corresponding intent feature vectors as analysis intent features.

[0033] Furthermore, the spatiotemporal context features, user historical preference features, and analysis intent features are concatenated, and then dimensionality reduction and feature fusion are performed through a fully connected layer network to output a session analysis focus vector. The preferred dimension of the session analysis focus vector is 256. The session analysis focus vector is used to represent the monitoring objects, time range, spatial range, indicator preferences, and analysis objectives that the user is interested in during the current session, and is used for targeted guidance in subsequent knowledge graph retrieval.

[0034] Subsequently, request entities are extracted from the analysis request, and guided by the session analysis focus vector, candidate database patterns associated with the request entities are retrieved from the domain knowledge graph, such as a water quality analysis graph database constructed using the Neo4j graph database. Specifically, a bidirectional maximum matching algorithm combined with a water quality chemistry-specific dictionary containing at least 50,000 technical terms is used to segment the natural language analysis request. When encountering segmentation ambiguities, the optimal segmentation sequence is selected based on the highest transition probability calculated by the contextual Hidden Markov Model. Using a pre-trained named entity recognition model, such as a BiLSTM-CRF model, entities such as water quality indicators, monitoring points, administrative regions, watershed names, and time ranges are extracted from the segmentation results and used as request entities.

[0035] To address non-standard colloquial expressions, the system calculates the similarity between the requested entity and standard entity names in the domain knowledge graph. This similarity is determined by combining edit distance similarity, thesaurus matching results, and sentence vector cosine similarity. The system retains standard entity names with similarity scores above a preset threshold, establishing a candidate entity set. This candidate entity set is then further mapped to corresponding table names, field names, field types, and primary / foreign key relationships in the underlying relational database, thus forming a candidate database schema. The candidate database schema may include one or more candidate tables, candidate fields, and corresponding table join relationships.

[0036] S2. Calculate the link confidence of the model and the structural confidence of the sequence.

[0037] The purpose of this step is to perform a dual reliability assessment on the retrieved graph mapping relationship and query action sequence, in order to characterize the determinism of entity links and the compliance and rigor of the query logic.

[0038] In the specific execution process, entity ambiguity resolution and pattern linking are first performed based on the relevance features between candidate database patterns and session analysis focus vectors. When obtaining relevance features, graph neural networks, such as GraphSAGE, can be used to extract pre-trained node features of candidate database patterns in the domain knowledge graph, and the cosine similarity between the node vectors of candidate database patterns and session analysis focus vectors is calculated to obtain a semantic consistency score.

[0039] Simultaneously, the shortest hop count from the requesting entity to the candidate database pattern node in the domain knowledge graph is calculated using a bidirectional breadth-first search graph traversal algorithm, and the shortest hop count is converted into a connection path weight score according to a preset distance decay function; the formula for constructing the distance decay function can be: ; in, Assign weight scores to the connection paths; The shortest number of hops; The set decay constant is preferably 0.5; this relationship indicates that in a domain knowledge graph, the business association tightness between entity nodes decays exponentially with the increase of topological distance. For example... Figure 2 As shown in the figure, this diagram illustrates the trend of the business association density between entity nodes in a knowledge graph as the topological distance hop count increases. With the increase in the topological distance hop count, the calculated attenuation weight value decreases non-linearly and exponentially, reflecting the characteristic that the greater the distance between nodes, the lower the degree of business association. This provides an intuitive basis for calculating the connection path weight score during pattern linking.

[0040] Based on this, semantic consistency score and connection path weight score are used as relevance features, and a weighted summation formula is used to calculate the relevance features. In this embodiment, the weight of semantic consistency score is 0.6, and the weight of connection path weight score is 0.4, resulting in a comprehensive pattern linking score for each candidate database pattern. The system sorts the candidate database patterns according to the comprehensive pattern linking score, selects the candidate database patterns that meet the preset score threshold or are ranked higher as target database patterns, and establishes one or more target mapping relationships between the request entity and the underlying database tables, fields, and connection relationships. Pattern linking is completed and polysemous ambiguity is resolved by writing UUID binding records in the mapping registry.

[0041] Furthermore, to generate link confidence, the system statistically analyzes a set of candidate database pattern nodes in the domain knowledge graph that have connections to the requesting entity and whose hop count is within a preset range. In this embodiment, the preset range is preferably a hop count of no more than 3. Let the overall pattern link score of the target database pattern be... The first node in the candidate database pattern node set The overall score of the pattern links of the candidate nodes is: The number of candidate nodes is The normalized base value can then be calculated using the following formula. : ; in, To prevent extremely small constants with a denominator of zero, this basic value is used to represent the relative credibility of the target database schema within the candidate domain structure. Subsequently, this basic value is mapped using a non-linear activation function, as shown in the following example: ; in, Link confidence; These are the normalized base values; The scaling parameter is preferably 10; The translation parameter is preferably 0.1. Using a Sigmoid function with translation and scaling factors, the output is compressed to the 0-1 range, improving the discriminative power of values ​​within the critical range. This yields a value that can be interpreted as probabilistic confidence, representing the link confidence level of entity link determinism.

[0042] In terms of structural confidence assessment, the system analyzes requests by deconstructing them into a logical execution sequence containing query actions and compliance judgment logic through a semantic parser. This logical execution sequence is composed of predefined Python atomic operators, which may include time filtering operators, location filtering operators, indicator selection operators, aggregation statistics operators, threshold comparison operators, and compliance judgment operators. Among these, the threshold comparison operator internally hard-coded the verification logic for national surface water environmental quality standards and water pollutant discharge standards, automatically comparing query values ​​with standard thresholds.

[0043] The system further constructs a directed acyclic graph to represent the sequential dependencies of operators in the logical execution sequence, and performs consistency verification on the input data type, output data type, time range, spatial range, indicator unit, and compliance rule reference relationship between adjacent operators. For example, when the output of the aggregation operator is a numerical statistical result, its subsequent threshold comparison operator should receive numerical input; when the analysis request involves a specific cross-section, the subsequent compliance determination operator should reference a standard threshold that matches the water body category to which that cross-section belongs.

[0044] In one optional implementation, the system uses a preset weighted scoring function to sum the operator compatibility score, execution order legality score, and standard rule matching score, and subtracts the abnormal path penalty term. After normalization, a structural confidence score for the logical execution sequence is obtained, ranging from 0 to 1. If the operator types match, the rule references are correct, and there are no circular dependencies or execution deadlocks, a high structural confidence score is assigned to the logical execution sequence. If there are inconsistencies in units, missing threshold standards, or incompatible operator inputs and outputs, the structural confidence score is reduced accordingly based on the severity of the problem.

[0045] S3. Generate candidate query statements and calculate the overall confidence level.

[0046] The purpose of this step is to convert natural language intent into executable query instructions and perform a global score. Its aim is to comprehensively evaluate the overall reliability of the query requests that are about to be dispatched to the database for execution, and to ensure the accuracy of the data extraction results.

[0047] During the specific execution process, the system extracts the query operation logic from the aforementioned logical execution sequence, and combines it with the target database schema after linking the abstract syntax tree mapping dictionary to generate one or more candidate query statements that conform to the underlying database dialect. Relational databases correspond to SQL statements, while time-series databases or graph databases correspond to their native query statements.

[0048] Subsequently, the system combines link confidence and structural confidence to calculate the overall confidence of the candidate query statement. Specifically, it extracts pre-trained node feature vectors of the database tables, fields, and connection relationships actually accessed by the candidate query statement from the domain knowledge graph, using them as query database pattern feature vectors. If the relevant feature vectors have been L2 normalized in the preceding steps, the dot product between the query database pattern feature vector and the session analysis focus vector can be directly calculated, and the resulting scalar value can be used as the pattern matching degree feature value to measure the alignment between the query scope and the user's analysis intent. If normalization has not been performed, cosine similarity can be used to calculate the pattern matching degree feature value.

[0049] Next, the link confidence, structural confidence, and pattern matching degree feature values ​​are used as feature parameter vectors and input into a pre-trained probability scoring model. The probability scoring model can be a random forest regression model, preferably a gradient boosting decision tree model. The specific process of outputting the comprehensive confidence of the candidate query statement from the probability scoring model can be as follows: The feature parameter vector is input into the gradient boosting decision tree model. The model reaches the corresponding leaf node based on the feature splitting condition and accumulates the weights of the leaf nodes of multiple decision trees. The accumulation calculation formula is as follows: ; in, To accumulate the output value; The total number of decision trees is set; For the first The leaf node weight mapping function output by the tree. This is the input feature parameter vector. After the model accumulates the weights of the leaf nodes output by multiple decision trees, it is then subjected to a probabilistic mapping using the Logistic function to obtain a comprehensive confidence score with values ​​between 0 and 1. The comprehensive confidence score is used to characterize the overall credibility of the candidate query statement in terms of entity linking, logical structure, and user intent matching, and is used for dynamic adjustment of subsequent data retrieval behavior and report generation strategies.

[0050] S4. Control report generation strategy to output water quality chemical analysis report.

[0051] The purpose of this step is to control the content granularity, reasoning depth, and tone of the natural language report based on the final overall confidence level. Its aim is to establish a linkage mechanism between query reliability and report generation quality, and to avoid over-inference or data illusion under low-confidence query conditions.

[0052] In the specific execution process, the system selects the target query statement with the highest overall confidence level and meeting the preset execution conditions from one or more candidate query statements based on the overall confidence level, and dispatches the target query statement to the water quality IoT database or water quality monitoring database for data query. Simultaneously, the system dynamically adjusts the report generation strategy of the large language model based on the overall confidence level, and generates a water quality chemical analysis report by combining the objective data results returned by the query, the session analysis focus vector, and the preset report template.

[0053] In one implementation, the system sets a preset confidence threshold, which is preferably 0.7 in this embodiment. When the overall confidence level is greater than the confidence threshold, the system confirms that the current query link meets the requirements for high-confidence analysis. At this time, a more complete data interpretation logic is loaded into the report generation framework, generating an analysis paragraph containing the data statistical process and indicator comparison results in the water quality chemical analysis report. The report may include the mean, maximum, minimum, year-on-year and month-on-month change rates of the monitoring indicators, spatial distribution difference analysis of different monitoring points, and comparison results with the corresponding water quality standard limits. The system can also call a relatively deterministic business terminology library, such as "the test results show that," "the monitoring data reflects that the indicator exceeds the limit," and "the indicator is within the compliance range," to generate explanatory text and convey the analysis conclusions based on the test data and standard thresholds to the user.

[0054] When the overall confidence level is less than or equal to the confidence threshold, the system restricts deep inference chains to reduce compliance risks and analytical misrepresentation. The water quality chemical analysis report only extracts core indicator data results for display, presenting only the average pollutant concentration, sampling time, monitoring location, and data source for the core monitoring sections, avoiding strong causal judgments or extended inferences. Simultaneously, the system uses a conservative terminology library containing probabilistic or suggestive statements, such as "preliminary results," "possible abnormal fluctuations," "suggestion for manual review," and "suggestion for supplementing on-site sampling data," to generate cautiously worded explanatory text. The generated report text can be integrated into standard format files such as PDF, HTML, or Word using a report renderer and output to the user.

[0055] To verify the technical effectiveness of the method proposed in this embodiment, a natural language question-answering dataset of 5,000 entries containing multiple intents, high-frequency ambiguous water quality entities, and spatiotemporally ambiguous words was selected as the test sample, and a five-fold cross-validation method was used for partitioning. Four comparative schemes were set up in the experiment: the baseline group used the traditional linking method of extracting natural language entities and performing single-string matching; ablation experiment group 1 removed the spatiotemporal context and historical preference features from the conversation analysis focus vector based on the complete scheme, using only natural language intent features; ablation experiment group 2 removed the shortest hop connection path weights calculated by the graph traversal algorithm based on the complete scheme, using only the semantic consistency score calculated by the graph neural network for database pattern linking; the experimental group used the complete technical scheme of this embodiment.

[0056] The experimental results show that: the database entity linking accuracy rate of the baseline group was 62.4%, the query execution success rate was 55.3%, and the manual verification accuracy score of the generated report was 61.2; the database entity linking accuracy rate of ablation experimental group 1 was 78.6%, the query execution success rate was 71.5%, and the manual verification accuracy score of the generated report was 73.8; the database entity linking accuracy rate of ablation experimental group 2 was 84.3%, the query execution success rate was 82.1%, and the manual verification accuracy score of the generated report was 80.5; the database entity linking accuracy rate of the experimental group reached 95.8%, the query execution success rate improved to 94.2%, and the manual verification accuracy score of the generated report reached 93.6.

[0057] The above experimental results are as follows Figure 3 As shown in the figure, the performance differences of the baseline model, ablation experiment group 1, ablation experiment group 2 and the method proposed in this paper on three evaluation indicators: link accuracy, query success rate and report manual verification score are compared, and the performance gradient changes of each group of methods are presented intuitively.

[0058] The comparative results show that the performance of ablation experiment group 1 is significantly improved compared to the baseline group, indicating that the fusion of spatiotemporal context and historical operation preference features can compensate for implicit user intent and reduce database query bias caused by the lack of time dimension and spatial coordinates. The performance of ablation experiment group 2 is further improved compared to ablation experiment group 1, indicating that simply relying on graph neural networks to extract semantic features can lead to misjudgment of database field nodes with similar structures but no close business relationship. Introducing topological path weights can effectively improve the accuracy of pattern linking. This embodiment comprehensively utilizes multi-dimensional feature joint encoding, connection path hop count decay evaluation, and large language model tone adjustment control, achieving optimal levels in all three indicators. This verifies the effectiveness of the complete technical solution in optimizing the accuracy of entity ambiguity resolution and pattern mapping, and enhancing the rigor and business relevance of water quality chemical analysis reports.

[0059] This invention also discloses an automatic report generation system for a water quality chemical analysis device, including a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement an automatic report generation method for a water quality chemical analysis device according to the present invention.

[0060] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for automatically generating reports from a water quality chemical analysis device, characterized in that, include: S1. Receive a natural language analysis request and extract the multi-dimensional conversation features of the analysis request to generate a conversation analysis focus vector; Extract the request entity from the analysis request, and retrieve the candidate database pattern associated with the request entity in the domain knowledge graph, guided by the session analysis focus vector; S2. Calculate the correlation feature between the candidate database pattern and the session analysis focus vector, perform entity ambiguity resolution and pattern linking based on the correlation feature, and generate link confidence; deconstruct the analysis request into a logical execution sequence containing compliance judgment logic, and evaluate the logical rigor of the logical execution sequence to generate structural confidence. S3. Based on the logical execution sequence and the candidate database pattern that completes the pattern link, generate candidate query statements, and calculate the comprehensive confidence of the candidate query statements by combining the link confidence and the structure confidence. S4. Based on the comprehensive confidence level, select the target query statement to perform data query, and dynamically adjust the report generation strategy based on the comprehensive confidence level. Combine the query results with the session analysis focus vector to generate a water quality chemical analysis report.

2. The method for automatically generating reports from a water quality chemical analysis device according to claim 1, characterized in that, Extracting multidimensional session features of the analysis request to generate a session analysis focus vector includes: extracting spatiotemporal context features, user historical preference features, and analysis intent features of the analysis request; The spatiotemporal context features, user historical preference features, and analysis intent features are concatenated and dimensionality reduced to generate the session analysis focus vector.

3. The method for automatically generating reports from a water quality chemical analysis device according to claim 2, characterized in that, Extracting the spatiotemporal context features and user historical preference features of the analysis request includes: parsing the analysis request to extract the timestamp and geographic coordinates, converting the timestamp into a standard time value and the geographic coordinates into a latitude-longitude floating-point pair, normalizing the standard time value and the latitude-longitude floating-point pair, and then combining them to construct the spatiotemporal context features. The user's historical operation logs are queried, and the water quality index parameters with the highest search frequency are extracted and vectorized as the user's historical preference features.

4. The method for automatically generating reports from a water quality chemical analysis device according to claim 1, characterized in that, Extracting request entities from the analysis request and retrieving candidate database patterns associated with the request entities in the domain knowledge graph, guided by the session analysis focus vector, includes: segmenting the natural language analysis request using a bidirectional maximum matching algorithm, and extracting water quality indicators, monitoring points, and time ranges from the segmentation results using a pre-trained named entity recognition model as the request entities; The similarity between the request entity and the standard entity name in the domain knowledge graph is calculated. Standard entity names with similarity higher than a preset threshold are retained to establish a candidate entity set, which is then mapped to the underlying database to form the candidate database pattern.

5. The method for automatically generating reports from a water quality chemical analysis device according to claim 1, characterized in that, Entity ambiguity resolution and pattern linking based on the correlation features between the candidate database patterns and the session analysis focus vector include: extracting the node vectors corresponding to the candidate database patterns in the domain knowledge graph, calculating the cosine similarity between the node vectors and the session analysis focus vectors to obtain a semantic consistency score; calculating the shortest hop count from the request entity to the candidate database pattern node in the domain knowledge graph, and converting the shortest hop count into a connection path weight score according to a preset distance decay function, with the semantic consistency score and the connection path weight score serving as the correlation features; using a weighted formula to perform a weighted summation of the correlation features to obtain a pattern link comprehensive score, and selecting the node with the highest pattern link comprehensive score to establish a unique mapping relationship to complete the pattern linking.

6. The method for automatically generating reports from a water quality chemical analysis device according to claim 5, characterized in that, The generated link confidence includes: counting the total number of nodes in the domain knowledge graph that have a connection relationship with the request entity and whose hop count is within a preset range; The basic quotient is calculated by dividing the pattern link comprehensive score of the node that establishes the unique mapping relationship by the sum of the pattern link comprehensive scores of all the candidate database patterns, and then mapping the basic quotient through a non-linear activation function.

7. The method for automatically generating reports from a water quality chemical analysis device according to claim 1, characterized in that, The comprehensive confidence of the candidate query statement is calculated by combining the link confidence and the structure confidence, including: extracting the feature vector of the candidate database pattern actually accessed by the candidate query statement as the query database pattern feature vector; Calculate the dot product between the query database pattern feature vector and the session analysis focus vector, and use it as the pattern matching degree feature value; The link confidence, the structure confidence, and the pattern matching degree feature values ​​are used as feature parameters and input into a pre-trained probability scoring model, which then outputs the comprehensive confidence of the candidate query statement.

8. The method for automatically generating reports from a water quality chemical analysis device according to claim 7, characterized in that, The probability scoring model employs a gradient boosting decision tree model, and outputs the comprehensive confidence score of the candidate query statement, including: The feature parameters are sequentially input into the gradient boosting decision tree model, the weights of the leaf nodes output by multiple decision trees in the model are accumulated, and after mapping transformation by the regression function, the comprehensive confidence score with numerical compression within a preset interval is obtained.

9. The method for automatically generating reports from a water quality chemical analysis device according to claim 1, characterized in that, Based on the comprehensive confidence level dynamic adjustment report generation strategy, a water quality chemical analysis report is generated by combining the query results and the session analysis focus vector, including: when the comprehensive confidence level is greater than the preset confidence level threshold, a detailed analysis paragraph containing the data deduction process is generated in the water quality chemical analysis report, and an explanatory text is generated by calling the deterministic tone lexicon. When the overall confidence level is less than or equal to the confidence level threshold, only the core indicator data results are extracted from the water quality chemical analysis report, and an explanatory text is generated by calling a conservative thesaurus containing probabilistic or advisory statements.

10. An automatic report generation system for a water quality chemical analysis device, characterized in that, include: A processor and a memory, the memory storing computer program instructions that, when executed by the processor, implement the automatic report generation method of the water quality chemical analysis equipment according to any one of claims 1-9.