Data query tool method and system based on AI language model

Through the method of layered processing and audit verification, the problem of inefficient existing data query tools in professional queries is solved, and high accuracy and efficient data query results are achieved, which are suitable for data retrieval in professional fields.

CN120256448APending Publication Date: 2025-07-04GUANGDONG POWER GRID CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510400726.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing data query tools based on AI language model are difficult to accurately understand and process professional vocabulary in data queries with high professionalism, resulting in low query efficiency and difficult to guarantee the accuracy and compliance of query results.

Method used

The hierarchical processing method is used to process the data query requests through the AI language model of the basic language layer, the general knowledge layer and the professional knowledge layer in turn, generate SQL query statements, and audit and verify the preliminary query results to ensure the accuracy and compliance of the results.

Benefits of technology

It improves the accuracy and efficiency of professional data query, meets the needs of complex and changeable data query scenarios, ensures the integrity, compliance and reliability of query results, and is suitable for application scenarios such as bidding, which require high data quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256448A_ABST
    Figure CN120256448A_ABST
Patent Text Reader

Abstract

The invention discloses a data query tool method and system based on an AI language model. The method comprises the steps that a data query request input by a target user is obtained; inputting the data query request into a basic language layer of the language analysis model to obtain basic semantic request data; inputting the basic semantic data into a general knowledge layer of a language analysis model to obtain fused semantic request data; inputting the fused semantic request data into a professional knowledge layer of a language analysis model to obtain professional semantic request data; mapping the professional semantic request data into a database to obtain an SQL (Structured Query Language) query statement; based on the SQL query statement, accessing a database to obtain a preliminary query result; and auditing and verifying the preliminary query result to obtain a target query result. According to the method, the accuracy of the query result is improved, and comprehensive consideration and design are carried out in the aspect of specialty, so that the complex and changeable data query scene requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a data query tool method and system based on an AI language model. Background Art

[0002] With the advent of the big data era, the amount of data faced by enterprises and individuals is increasing day by day. How to efficiently extract valuable information from the vast amount of data has become an important issue. Traditional data query methods usually rely on structured query languages (such as SQL) or specific search tools, which require users to have certain technical knowledge and query syntax. In recent years, the development of artificial intelligence technology, especially natural language processing (NLP), has provided new solutions for data query. A data query tool based on an AI language model can understand the natural language input of users and convert it into effective query statements, thus realizing more convenient and intuitive data retrieval.

[0003] Currently, data query tools based on AI language models mainly adopt the following methods, including natural language understanding, query generation, result presentation, and interactive feedback, etc. These methods enable non-technical personnel to easily perform complex data query operations. Although existing data query tools based on AI language models perform well in many application scenarios. However, in data queries involving higher professionalism, such as bidding queries, existing general-purpose data query tools are difficult to accurately understand and process professional terms in bidding, resulting in the need for multiple queries to obtain satisfactory data sometimes. And during the bidding process, the accuracy and compliance of data are crucial. However, the data determined by the current query methods still need to be judged by users themselves based on professional knowledge before they can be used, leading to low query efficiency. Summary of the Invention

[0004] The object of the present invention is to solve the problems in the prior art, and a data query tool method and system based on an AI language model are proposed.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] A data query tool method system based on an AI language model, comprising the following steps:

[0007] Obtain a data query request input by a target user;

[0008] Input the data query request into the basic language layer of the language parsing model to obtain basic semantic request data;

[0009] Input the basic semantic data into the general knowledge layer of the language parsing model to obtain fused semantic request data;

[0010] Input the fused semantic request data into the expertise layer of the language parsing model to obtain professional semantic request data; the language parsing model is trained based on sample request data and their corresponding hierarchical semantic label results.

[0011] Map the professional semantic request data to a database to obtain an SQL query statement.

[0012] Access the database based on the SQL query statement to obtain a preliminary query result.

[0013] Audit and verify the preliminary query result to obtain a target query result.

[0014] An AI language model-based data query tool system includes:

[0015] An acquisition unit: used to acquire a data query request input by a target user.

[0016] A first processing unit: used to input the data query request into the basic language layer of the language parsing model to obtain basic semantic request data.

[0017] A second processing unit: used to input the basic semantic data into the general knowledge layer of the language parsing model to obtain fused semantic request data.

[0018] A third processing unit: used to input the fused semantic request data into the expertise layer of the language parsing model to obtain professional semantic request data; the language parsing model is trained based on sample request data and their corresponding hierarchical semantic label results.

[0019] A mapping unit: used to map the professional semantic request data to a database to obtain an SQL query statement.

[0020] A search unit: used to access the database based on the SQL query statement to obtain a preliminary query result.

[0021] A verification unit: used to audit and verify the preliminary query result to obtain a target query result.

[0022] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the above-mentioned AI language model-based data query tool method are implemented.

[0023] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the above-mentioned AI language model-based data query tool method are implemented.

[0024] The present invention has the following advantages compared with the prior art:

[0025] A data query tool method and system based on an AI language model provided by the present invention, by sequentially passing the obtained data query request through the basic language layer, general knowledge layer, and professional knowledge layer of the language model, enabling the AI language model to combine hierarchical processing with database interaction, not only improves the accuracy of the query result, but also comprehensively considers and designs in terms of professionalism to meet the requirements of complex and changeable data query scenarios, and has significant advantages in application scenarios with high data quality requirements such as bidding and professional field data retrieval; in addition, the data query request processed by the language model is first mapped in the database, and after the query is completed, it is audited and verified, further ensuring the integrity, compliance, and accuracy of the query result, improving the query efficiency and providing a reliable guarantee for the use of data. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 It is a schematic flowchart of a data query tool method based on an AI language model provided by an embodiment of the present invention;

[0028] Figure 2 It is a schematic structural diagram of a data query tool system based on an AI language model provided by an embodiment of the present invention;

[0029] Figure 3 It is a schematic structural diagram of an electronic device proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0031] The following will describe Figure 1 - Figure 3 a data query tool method and system based on an AI language model of the present invention.

[0032] Figure 1It is a schematic flow chart of a data query tool method provided by the present invention. As Figure 1 shown, the method includes:

[0033] Step 101, obtaining a data query request input by a target user.

[0034] The data query tool system obtains the data query request input by the user from the user interface, and the data query request can be any one or a combination of text, voice, and picture upload methods. When obtaining a data query request in a combined form, the following steps are specifically used to complete the splicing of the entire data query request, including:

[0035] Obtaining the text request information, voice request information, and picture request information of the target user; performing voice recognition conversion on the voice request information to obtain first conversion information; performing picture recognition conversion on the picture request information to obtain second conversion information; splicing the text request information, first conversion information, and second conversion information to obtain a data query request.

[0036] Furthermore, during the conversion process of the voice request information, the built-in integrated voice recognition technology of the data query tool system is used to perform real-time recognition conversion on the sent voice request information. For example, in one embodiment, in a mobile device application, after the user clicks the voice input button and says to search for "conference information about the power intelligent seminar published in the last month", the system converts the voice into text and then performs subsequent processing. Moreover, dialect recognition matching can be set in the voice recognition technology to adapt to users with different accents, semantics, and language habits.

[0037] In the picture upload method, the built-in optical character recognition technology in the data query tool system is used to recognize the text information in the uploaded picture and extract the text information.

[0038] Obtaining the data query request input by the target user through multiple methods can accurately capture the actual data needs of the user. Based on this, subsequent processing can be carried out targeted, and it can adapt to the usage needs and actual scenarios of different users, thereby improving the generalization ability of the system.

[0039] Step 102, inputting the data query request into the basic language layer of the language parsing model to obtain basic semantic request data.

[0040] Specifically, after the data query tool system receives the final text data query request, it first passes the original query request input by the user to the basic language layer of the language parsing model. The basic language layer will perform a preliminary language analysis on the request in natural language form, and finally sort out the natural language requests with strong randomness and diverse forms from the user into a data form with a clear structure and definite semantic components, which helps to improve the accuracy of the entire data processing process and avoid understanding deviations caused by the ambiguity and complexity of natural language.

[0041] Furthermore, within the basic language layer of the language parsing model in the data query tool system, the data query request is converted into basic semantic request data, and the specific steps are as described in 1021 - 1025.

[0042] Step 103: Input the basic semantic data into the general knowledge layer of the language parsing model to obtain the fused semantic request data.

[0043] Specifically, in the general knowledge layer of the language parsing model, the basic semantic data combines general knowledge to further improve and expand the basic semantic data, integrates relevant general knowledge into it, makes the semantics richer and more accurate, realizes making up for the possible deficiencies in understanding solely from the statement itself with the help of general knowledge, makes the semantics of the request data more in line with the conventional understanding in the actual application scenario, and enhances the rationality and accuracy of the data request in a wider context.

[0044] Furthermore, in the general knowledge layer of the large model in the data query tool system, the basic semantic data is fused into the fused semantic request data, and the specific steps are as described in 1031 - 1036.

[0045] Step 104: Input the fused semantic request data into the professional knowledge layer of the language parsing model to obtain the professional semantic request data; the language parsing model is trained based on sample request data and their corresponding hierarchical semantic label results.

[0046] Specifically, the professional knowledge layer of the large voice model focuses on the application of knowledge in a specific professional field. When receiving the fused semantic request data, according to the model parameters trained in advance through a large number of sample request data and their corresponding hierarchical semantic label results, it integrates the knowledge, rules, logic, etc. unique to the professional field, and finally outputs the professional semantic request data that meets the professional requirements to ensure the accuracy of the data request at the professional level. Thus, the data request is accurately adapted to the actual situation and requirements of the professional field, fully considering the special rules and logic within the professional field, greatly improving the accuracy and reliability of professional data queries, and meeting the needs of professional users to obtain high-quality data in specific fields.

[0047] Furthermore, in the professional knowledge layer of the language parsing model of the data query tool system, when converting the fused semantic request data into professional semantic request data, the specific steps are as described in 1041-1044.

[0048] Step 105: Map the professional semantic request data to the database to obtain an SQL query statement.

[0049] Specifically, the data query tool system converts the professional semantic request data into the corresponding SQL query statement (i.e., Structured Query Language) according to the structure of the database and the query language specification. The converted SQL query statement can facilitate the execution of query operations in the database, realizing the conversion of data requirements at the language level to actual query instructions in the database, enabling the subsequent use of the powerful storage and retrieval functions of the database to obtain the corresponding data, and ensuring the smooth progress of the entire query process in the database environment.

[0050] Step 106: Based on the SQL query statement, access the database to obtain a preliminary query result.

[0051] Specifically, use the SQL query statement generated in Step 105 to execute the statement in the corresponding database management system. The database will perform operations such as retrieving, filtering, and aggregating the stored data according to the conditions set in the statement (such as filtering conditions, joining conditions, etc.), and then return the preliminary query result that meets the conditions. By leveraging the efficient query processing mechanism of the database, it quickly and accurately extracts data that may meet the user's needs from the massive data, completing a key step from data request to actual data acquisition, and laying a foundation for presenting effective data to the user ultimately.

[0052] Step 107: Audit and verify the preliminary query result to obtain the target query result.

[0053] Specifically, conduct various inspections and verifications on the preliminary query result obtained from the database, including checking the integrity of the data (i.e., whether all the necessary data is included), accuracy (i.e., whether the data values are correct), and compliance (i.e., whether it conforms to relevant business rules, data security regulations, etc.). After review and correction, determine the final result that meets the requirements as the target query result and provide it to the user. This ensures that the data presented to the user finally is of high quality, reliable, and meets various specification requirements, avoiding the situation where the user receives incorrect or inaccurate data due to possible errors in the database itself or omissions in the previous processing links, and improving the quality and credibility of the entire data query service.

[0054] The present invention relates to the field of artificial intelligence, and provides a data query tool method and system based on an AI language model. In the present invention, the proposed data query tool method based on an AI language model can sequentially process the obtained data query request through the basic language layer, general knowledge layer, and professional knowledge layer of the speech model, enabling the AI language model to combine hierarchical processing with database interaction. This not only improves the accuracy of the query results but also comprehensively considers and designs in terms of professionalism to meet the requirements of complex and changing data query scenarios, achieving significant advantages in application scenarios with high data quality requirements such as bidding and professional field data retrieval. In addition, the data query request processed by the language model is first mapped in the database, and then audited and verified after the query is completed, further ensuring the integrity, compliance, and accuracy of the query results, improving the query efficiency while providing reliable guarantee for the use of data.

[0055] In one embodiment, the descriptions of steps 1021-1025 are as follows:

[0056] Step 1021: Filter out noise and standardize the statement of the data query request to obtain basic target request data.

[0057] It should be noted that the data query tool system of the embodiment of the present invention is embedded with a pre-trained language parsing model. Among them, the language parsing model is trained based on sample request data and its corresponding hierarchical semantic label results. The pre-trained model can be a recurrent neural network, a convolutional neural network, etc. Therefore, the data query tool system inputs the data query request into the basic language layer (i.e., the bottom layer of the model) of the language parsing model. The language parsing model performs preliminary language analysis on the request in the natural language form in the basic language layer and outputs data with a clear structure and definite semantic components (i.e., basic semantic request data).

[0058] Specifically, in the noise filtering of the data query request, elements such as special symbols, garbled characters, and redundant blank characters that may interfere with subsequent analysis are removed. Then, terms or common abbreviations in a specific field are standardized to establish a unified and accurate semantic basis to improve the accuracy of data processing, so that the identified abbreviations and terms are replaced with standard and complete expression forms. This realizes the conversion of concise but possibly ambiguous abbreviations or professional terms into easy-to-understand and clearly expressed full names, eliminating the semantic ambiguity problems caused by different users' inconsistent understandings of abbreviations, and enabling the text to be constructed based on a unified and accurate semantics in the basic language layer.

[0059] Step 1022: Split the basic target request data to obtain multiple first sub-basic data, and establish a logical association relationship between adjacent first sub-basic data; the logical association relationship includes the semantic association degree and syntactic association degree between adjacent first sub-basic data.

[0060] Specifically, after obtaining the basic target request data, split it to analyze its internal structure and semantic relationship in depth. When splitting, it can be carried out in any way such as words, sentence segmentation, or specific syntactic structures, and the word-based splitting method is preferred here. After splitting to obtain multiple first sub-basic data, when establishing the logical association relationship between the vector first sub-basic data, the determination of the semantic association degree needs to consider the closeness of the semantic connection between words, while the syntactic association degree focuses on analyzing the roles and mutual relationships of words in the sentence syntactic structure.

[0061] Furthermore, in one embodiment, the semantic association degree can be determined by calculating the word occurrence frequency and semantic vector similarity, and the syntactic association degree can be determined by using the result relationship of the syntactic analysis tree. Through the above method, the complex request data is split into finer-grained sub-data, which is convenient for detailed analysis of its internal semantics and syntactic structure, can deeply explore the internal structure of the request data, and provides a detailed information basis for more accurate semantic role annotation and semantic relationship mapping in the future.

[0062] Step 1023: Perform semantic role annotation on multiple first sub-basic data to obtain multiple second sub-basic data respectively.

[0063] Specifically, semantic role annotation is to determine the semantic role played by each first sub-basic data in the entire query request statement. For example, the split "query" is labeled as an action role, indicating the operation type of the entire request; "in the past year" is labeled as a time role, indicating the time range limit; "sales amount" is labeled as an object role, which is the main data object targeted by the query. Through such semantic role annotation, each first sub-basic data is given a clear semantic function identifier, thus forming multiple second sub-basic data. In the process of realizing this, it can be achieved through a trained semantic role annotation model, and the semantic role annotation model is trained based on a labeled corpus and can accurately judge the semantic role according to the context environment and language knowledge of words.

[0064] By clarifying the semantic function of each sub-data in the entire query request, the semantic expression of the data becomes clearer and more structured, providing more accurate semantic information for subsequent semantic relationship mapping and data integration.

[0065] Step 1024: Map multiple second sub-basic data in a pre-constructed semantic relationship model to obtain multiple third sub-basic data respectively.

[0066] Specifically, the pre-constructed semantic relationship model here is a model containing rich semantic relationship knowledge, which is constructed based on knowledge graph and semantic network technologies. For example, in a semantic relationship model in the power field, sales amount is the second sub-basic data, which is mapped to the financial indicator category in the semantic relationship model, and is associated with other financial indicators such as profit and cost. Finally, a third sub-basic data containing richer semantic information and domain knowledge is obtained. Thus, it realizes integrating the sub-data after semantic role annotation into a wider semantic knowledge system. By mapping with the pre-constructed semantic relationship model, the sub-data can obtain more domain knowledge and semantic association information, enriching its semantic connotation to better meet the processing requirements of data query requests.

[0067] Step 1025: According to the logical association relationship between adjacent first sub-basic data, splice multiple third sub-basic data to obtain basic semantic request data.

[0068] Specifically, after obtaining multiple third sub-basic data, it is necessary to splice them according to the logical association relationship established between adjacent first sub-basic data before. For example, according to the previously determined syntactic association degree and semantic association degree, place the third sub-basic data representing an action (such as the result after semantic relationship mapping of "query") at the front, and then arrange the third sub-basic data related to time, object, condition, subject, result, etc. in sequence according to the logical relationship. The spliced basic semantic request data may be a structured expression form, which clearly reflects the organic combination of information such as query action, time range, object condition, subject object, and expected result, and each part contains rich semantic knowledge and association information obtained from the semantic relationship model. And splicing is not a simple string connection, but a reasonable organization and integration according to the logical relationship, which may involve some operations such as format conversion and semantic integration marking to ensure that the spliced basic semantic request data can accurately and completely express the semantic intention of the original query request.

[0069] By splicing according to the logical association relationship, the accuracy, integrity, and logic of the basic semantic request data are ensured, making it contain both the core information of the original request and rich semantic knowledge and association information, providing a high-quality semantic basis for subsequent data processing.

[0070] In summary, in steps 1021 - 1025, data quality and processing efficiency are improved through noise filtering and statement standardization; the internal structure of the data is deeply explored by splitting and establishing logical association relationships; semantic role annotation clarifies the semantic functions of the data; semantic relationship model mapping enriches the semantic connotations of the data; and splicing based on logical association relationships ensures the integrity and logic of the data. The entire process improves the accuracy, reliability, and scalability of data parsing, and can better meet the processing requirements of data query requests in different fields and with different levels of complexity.

[0071] In one embodiment, steps 1031 - 1036 are described as follows:

[0072] In step 1031, based on the basic semantic request data, a knowledge graph is constructed, and unique identifiers for the basic semantic request data and the nodes in the knowledge graph are generated through semantic fingerprints.

[0073] Specifically, first, for the basic semantic request data, a related knowledge graph is constructed. During the construction of the knowledge graph, relevant entities, relationships, and attributes are extracted from various data sources. These entities, relationships, and attributes form the nodes and edges of the knowledge graph. Then, semantic fingerprint technology is used to generate unique identifiers. A semantic fingerprint is a method of converting text information into a fixed - length digital or string identifier. For the basic semantic request data, through a specific hash function or feature extraction algorithm, it is converted into a unique semantic fingerprint. For example, information such as keywords and semantic structures in the request data can be encoded to generate a string like "f89a3bc1" as its unique identifier. For each node in the knowledge graph, a corresponding semantic fingerprint is also generated according to the characteristics of the entity or relationship it represents. The purpose is to be able to quickly and accurately identify and compare different elements during subsequent matching processes.

[0074] Furthermore, through the knowledge graph, rich semantic association information can be provided, enabling the basic semantic request data to be connected to a wider knowledge system and enhancing the depth of understanding of the request data. Semantic fingerprint technology can quickly locate and compare elements, reducing the computational amount and time complexity during the matching process and improving the efficiency of the entire data processing flow.

[0075] In step 1032, based on the unique identifiers, the basic semantic request data is matched with the knowledge graph to obtain the first graph region.

[0076] Specifically, using the unique identifier generated previously, perform a quick match in the knowledge graph. For example, compare the semantic fingerprint "f89a3bc1" of the basic semantic request data with the semantic fingerprints of each node in the knowledge graph. When it is found that the semantic fingerprints of some nodes have a high similarity or relevance to the semantic fingerprint of the request data, the area formed by these nodes and their related edges is determined as the first graph area. During the matching process, use a similarity calculation algorithm, such as the cosine similarity formula: where and represent the feature vectors of the request data and the node respectively, and the matching degree is judged by calculating the cosine similarity between them.

[0077] Through the matching based on the unique identifier, it is possible to quickly locate the knowledge area that may be related to the request data, avoid a comprehensive and time-consuming search of the entire knowledge graph, improve the speed and efficiency of data processing, and reduce the interference of irrelevant information.

[0078] Step 1033, perform a deep semantic match on the first graph area to obtain a second graph area; the deep semantic match includes entity names, entity attributes, relationships, and context.

[0079] Specifically, after the first graph area is determined, perform a deep semantic match. For entity name matching, not only the exact name needs to be matched, but also synonyms, abbreviations, etc. need to be considered; for entity attribute matching, check the consistency between the attribute requirements in the request data and the entity attributes in the knowledge graph; in terms of relationship matching, analyze whether the implicit relationships in the request data are consistent with the relationships in the knowledge graph; at the same time, consider the context, such as the industry background where the request data is located, the intention background of the user's question, etc. By comprehensively matching these aspects, a more accurate second graph area is further selected.

[0080] The deep semantic match can improve the matching accuracy between the knowledge graph and the request data, reduce the incorrect matches caused by semantic ambiguity or incompleteness, thereby improving the accuracy of the entire data processing, and making the finally obtained fused semantic request data better reflect the true needs of users.

[0081] Step 1034, extract the entity detailed information related to the basic semantic data from the second graph area.

[0082] Specifically, in the determined second graph region, entity detailed information related to the basic semantic request data is extracted. For example, for a company entity, detailed information such as the company's historical sales data, the sales proportion of different product lines, and the market distribution in different regions of the world is extracted; for a market entity, historical data on market size, major competitors, trends in market growth rate, etc. are extracted. This information is extracted in a structured form, such as in a table format, for subsequent processing and integration.

[0083] By extracting entity detailed information, the content of the basic semantic request data can be enriched, making it not only limited to the simple semantic expression in the original request, but incorporating more detailed knowledge from the knowledge graph, thereby improving the richness and practicality of the data.

[0084] Step 1035: Based on the entity detailed information, construct a relationship network among the entities in the second graph region.

[0085] Specifically, according to the extracted entity detailed information, construct a relationship network among the entities in the second graph region. And when constructing the relationship network, the weight and direction of the relationship are also considered. The weight of the relationship represents the tightness of different relationships, and the direction of the relationship represents the flow or dependency relationship of information. In this way, a more complete, accurate, and semantically meaningful relationship network is constructed. The relationship network can intuitively display the semantic associations and interactions among entities, helping to improve the understanding and analysis ability of the data, and providing a clear structural basis for subsequent fusion operations, making the fusion process more orderly and accurate.

[0086] Step 1036: Based on the relationship network and the basic semantic request data, perform fusion to obtain fused semantic request data.

[0087] Specifically, fuse the constructed relationship network with the basic semantic request data. For example, for the basic semantic request data "query the share and growth trend of XX Company in the power equipment market in 2024", incorporate data on the share of XX Company in the power equipment market in 2024, influencing factors related to the share (such as the situation of competitors, the overall market trend, etc.), and analysis data on the growth trend, etc. in the relationship network into the request data to form a fused semantic request data that incorporates the general knowledge of the knowledge graph. This fused semantic request data may be a structured expression, containing not only the semantic information of the original request, but also detailed knowledge and analysis results obtained from the relationship network.

[0088] The fused semantic request data has richer semantic connotations and knowledge reserves, can better handle various complex data analysis and processing tasks, improves the quality and efficiency of the entire data processing process, and makes the final data result more in line with the expectations of users and actual business needs.

[0089] In summary, in steps 1031 - 1036, constructing the knowledge graph and generating unique identifiers improve the efficiency of knowledge integration and retrieval; matching and deep semantic matching improve the accuracy of knowledge acquisition; extracting entity details and constructing relationship networks enhance the structuring and relevance of data; and the fusion operation enhances the semantic richness and practicality of the requested data. The entire process improves the quality and efficiency of data processing at the general knowledge level, enabling the data to better combine with the general knowledge system and enhancing the ability to understand and process the requested data.

[0090] In one embodiment, the description of steps 1041 - 1044 is as follows:

[0091] Step 1041, identify multiple professional fields involved in the fused semantic request data through semantic analysis to obtain multiple professional knowledge ontologies.

[0092] Specifically, for the fused semantic request data, semantic analysis must first be performed. The semantic analysis process involves lexical analysis, that is, part-of-speech tagging and semantic parsing of each word in the request data, and syntactic analysis is used to analyze the sentence structure and clarify the relationships between components. Through these analyses and matching with the professional field term library, multiple professional fields involved in the request are identified, and then based on the constructed professional knowledge system, multiple professional knowledge ontologies for the corresponding multiple professional fields are obtained. The professional ontology organizes knowledge in a structured manner, which contains elements such as classes, attributes, and instances. Classes represent concepts, attributes describe the relationships between classes, and instances are specific entities. Through semantic analysis, the professional fields can be accurately identified, avoiding blind searching in the vast knowledge system and improving the efficiency and accuracy of knowledge acquisition. The professional knowledge ontology organizes knowledge in a structured way, facilitating subsequent feature mapping and knowledge fusion operations and contributing to improving the structuring degree of the entire data processing.

[0093] Step 1042, map the features in the fused semantic request data to the features of multiple professional knowledge ontologies through an ontology mapping algorithm to obtain multiple knowledge fragments related to the fused semantic request data.

[0094] Specifically, an ontology mapping algorithm is adopted, such as a mapping algorithm based on concept similarity calculation. For the features in the fused semantic request data, search for matching or similar concept features in the professional knowledge ontology, calculate the concept similarity, adopt a method based on the vector space model, and then determine the mapping relationship by calculating the similarity between vectors. When a matching or similar concept is found, extract the knowledge fragment associated with the concept.

[0095] The ontology mapping algorithm can automatically and relatively accurately locate relevant knowledge in the professional knowledge ontology library, improving the automation and accuracy of knowledge acquisition. Compared with manual knowledge search, it greatly improves efficiency and can ensure the relevance of the acquired knowledge fragments to the request data.

[0096] Step 1043: Calculate the semantic similarity between the knowledge fragments in multiple professional knowledge ontology libraries and the fused semantic request data.

[0097] Specifically, when calculating the semantic similarity, multiple factors can be considered. First, for the similarity at the lexical level, the edit distance algorithm (such as the Levenshtein distance) can be used to measure the similarity between the words in the knowledge fragments and the words in the fused semantic request data. Second, considering the similarity at the semantic level, a semantic vector model (such as the Word2Vec or GloVe model) is used to map words into the semantic space, and the vector similarity between the knowledge fragments and the request data in the semantic space is calculated. In addition, the syntactic structure similarity can also be considered, analyzing the matching degree of the sentence structures and syntactic components of the knowledge fragments and the request data. By assigning certain weights to the lexical similarity, semantic similarity, and syntactic structure similarity respectively, and then performing a weighted sum, the comprehensive semantic similarity is obtained.

[0098] Among them, when using the semantic vector model, the formula for calculating the vector similarity between the knowledge fragments and the request data in the semantic space is:

[0099]

[0100] A = [a1, a2, …, a n

[0101] B = [b1, b2, …, b n

[0102] Among them, A represents the fused semantic request data vector, and B represents the knowledge fragment vector.

[0103] By accurately calculating the semantic similarity, the knowledge fragments can be screened and fused more scientifically and reasonably, avoiding simply treating all knowledge fragments equally, improving the accuracy and quality of the fused professional semantic request data, and making it more in line with the semantic requirements of the original request data.

[0104] Step 1044: Based on the semantic similarity, perform weighted fusion of multiple knowledge fragments and the fused semantic request data to obtain the professional semantic request data.

[0105] ​​Specifically, according to the calculated semantic similarity, the knowledge fragments are weighted and fused. Define the fused semantic request data as "Query the competitor quotes and technical advantage analysis in the bidding for the 500 kV substation construction project of a certain power engineering company". After calculating the semantic similarity between the knowledge fragments in multiple professional knowledge ontologies and the fused semantic request data, assume that a knowledge fragment about the compliance of the bidding process is obtained from the power engineering bidding regulations knowledge ontology, and its semantic similarity is 0.6. For example, this knowledge fragment contains "In power engineering bidding, the release time of the bidding announcement should be no less than the legal working days, the deadline for submitting tender documents should be clear and compliant, and the bid opening process requires the participation of a notary institution or relevant supervision department, etc.".

[0106] A knowledge fragment about the cost estimation and quotation strategy of the substation construction project is obtained from the power engineering cost knowledge ontology, and the semantic similarity is 0.7. It may include "The main cost components of the 500 kV substation construction project include land acquisition, equipment procurement, construction and installation works, etc. The quotation strategy can be determined according to the cost-plus method combined with the market competition situation, and the equipment procurement cost accounts for about 40% - 50%, etc.".

[0107] A knowledge fragment about the technical specifications and advantage comparison of the 500 kV substation is obtained from the power engineering technical standards knowledge ontology, and the semantic similarity is 0.8. For example, "The main technical standards of the 500 kV substation include that the electrical main wiring design should meet reliability, flexibility and economy, the transformer should have characteristics such as high efficiency, energy saving and low noise, and it has obvious advantages in terms of transmission capacity and coverage compared with substations of other voltage levels".

[0108] Perform weighted fusion. Taking text fusion as an example, the key sentences in the knowledge fragments are incorporated into the fused semantic request data according to their semantic similarity weights. For the knowledge fragment about the compliance of the bidding process, after incorporating it with a weight of 0.6, we get "Query the competitor quotes and technical advantage analysis in the bidding for the 500 kV substation construction project of a certain power engineering company". For the knowledge fragment about the cost estimation and quotation strategy of the substation construction project, after incorporating it with a weight of 0.7, it becomes "Query the competitor quotes and technical advantage analysis in the bidding for the 500 kV substation construction project of a certain power engineering company". For the knowledge fragment about the technical specifications and advantage comparison of the 500 kV substation, after incorporating it with a weight of 0.8, we get "Query the competitor quotes and technical advantage analysis in the bidding for the 500 kV substation construction project of a certain power engineering company", and finally form the complete professional semantic request data.

[0109] In summary, in steps 1041-1044, semantic analysis to identify the professional field improves the pertinence of knowledge acquisition; the ontology mapping algorithm improves the accuracy and automation of knowledge extraction; semantic similarity calculation provides a scientific basis for knowledge fusion; weighted fusion enhances the quality and practicality of professional semantic request data. The entire process improves the accuracy, efficiency, and reliability of data processing at the professional knowledge level, enabling data to better combine with the professional knowledge system and enhancing the ability to handle complex professional data requests.

[0110] Figure 2 is a schematic structural diagram of a data query tool system based on an AI language model provided by the present invention, as Figure 2 shown, the device includes: an acquisition unit 10: used to acquire a data query request input by a target user; a first processing unit 20: used to input the data query request into the basic language layer of the language parsing model to obtain basic semantic request data; a second processing unit 30: used to input the basic semantic data into the general knowledge layer of the language parsing model to obtain fused semantic request data; a third processing unit 40: used to input the fused semantic request data into the professional knowledge layer of the language parsing model to obtain professional semantic request data; the language parsing model is trained based on sample request data and its corresponding hierarchical semantic label results; a mapping unit 50: used to map the professional semantic request data to a database to obtain an SQL query statement; a search unit 60: used to access the database based on the SQL query statement to obtain a preliminary query result; a verification unit 70: used to perform audit verification on the preliminary query result to obtain a target query result.

[0111] Figure 3 is a schematic structural diagram of an electronic device provided by the present invention, as Figure 3As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340. Among them, the processor 33, the communications interface 320, and the memory 330 complete communication with each other through the communication bus 340. The processor 33 may call the logical instructions in the memory 330 to execute a data query tool method and its system based on an AI language model. The method includes: obtaining a data query request input by a target user; inputting the data query request into the basic language layer of a language parsing model to obtain basic semantic request data; inputting the basic semantic data into the general knowledge layer of the language parsing model to obtain fused semantic request data; inputting the fused semantic request data into the professional knowledge layer of the language parsing model to obtain professional semantic request data; the language parsing model is trained based on sample request data and its corresponding hierarchical semantic label results; mapping the professional semantic request data to a database to obtain an SQL query statement; accessing the database based on the SQL query statement to obtain a preliminary query result; auditing and verifying the preliminary query result to obtain a target query result.

[0112] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, and other various media that can store program codes.

[0113] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute a data query tool method and its system based on an AI language model provided by each of the above methods. The method includes: obtaining a data query request input by a target user; inputting the data query request into the basic language layer of a language parsing model to obtain basic semantic request data; inputting the basic semantic data into the general knowledge layer of the language parsing model to obtain fused semantic request data; inputting the fused semantic request data into the professional knowledge layer of the language parsing model to obtain professional semantic request data; the language parsing model is trained based on sample request data and its corresponding hierarchical semantic label results; mapping the professional semantic request data into a database to obtain an SQL query statement; accessing the database based on the SQL query statement to obtain a preliminary query result; and performing audit verification on the preliminary query result to obtain a target query result.

[0114] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute a data query tool method and its system based on an AI language model provided by each of the above. The method includes: obtaining a data query request input by a target user; inputting the data query request into the basic language layer of a language parsing model to obtain basic semantic request data; inputting the basic semantic data into the general knowledge layer of the language parsing model to obtain fused semantic request data;

[0115] inputting the fused semantic request data into the professional knowledge layer of the language parsing model to obtain professional semantic request data; the language parsing model is trained based on sample request data and its corresponding hierarchical semantic label results; mapping the professional semantic request data into a database to obtain an SQL query statement; accessing the database based on the SQL query statement to obtain a preliminary query result; and performing audit verification on the preliminary query result to obtain a target query result.

[0116] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0118] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A data query tool method based on an AI language model, characterized in that, It includes the following steps: Obtain a data query request input by a target user; Input the data query request into the basic language layer of a language parsing model to obtain basic semantic request data; Input the basic semantic data into the general knowledge layer of the language parsing model to obtain fused semantic request data; Input the fused semantic request data into the professional knowledge layer of the language parsing model to obtain professional semantic request data; the language parsing model is trained based on sample request data and their corresponding hierarchical semantic label results; Map the professional semantic request data to a database to obtain an SQL query statement; Access the database based on the SQL query statement to obtain a preliminary query result; Audit and verify the preliminary query result to obtain a target query result.

2. The method of the data query tool based on the AI language model according to claim 1, wherein The obtaining of the data query request input by the target user includes: Obtain the text request information, voice request information, and picture request information of the target user; Perform speech recognition conversion on the voice request information to obtain first conversion information; Perform picture recognition conversion on the picture emotion information to obtain second conversion information; Concatenate the text request information, first conversion information, and second conversion information to obtain a data query request.

3. The method of the data query tool based on the AI language model according to claim 1, characterized in that, The inputting of the data query request into the basic language layer of the language parsing model to obtain basic semantic request data includes: Perform noise filtering and statement standardization on the data query request to obtain basic target request data; Split the basic target request data to obtain multiple first sub-basic data, and establish a logical association relationship between adjacent first sub-basic data; the logical association relationship includes the semantic association degree and syntactic association degree between adjacent first sub-basic data; Perform semantic role annotation on the multiple first sub-basic data to obtain multiple second sub-basic data respectively; Map the multiple second sub-basic data in a pre-constructed semantic relationship model to obtain multiple third sub-basic data respectively; Based on the logical association relationship between adjacent first sub-basic data, concatenate the multiple third sub-basic data to obtain basic semantic request data.

4. The method of the data query tool based on the AI language model according to claim 1, wherein, The inputting of the basic semantic request data into the general knowledge layer of the language parsing model to obtain fused semantic request data includes: Based on the basic semantic request data, construct a knowledge graph, and generate unique identifiers for the basic semantic request data and the nodes in the knowledge graph through semantic fingerprints; Based on the unique identifier, match the basic semantic request data with the knowledge graph to obtain a first graph region; Perform deep semantic matching on the first graph region to obtain a second graph region; the deep semantic matching includes entity names, entity attributes, relationships, and context; Extract entity detailed information related to the basic semantic data from the second graph region; Based on the entity detailed information, construct a relationship network between the entities in the second graph region; Based on the fusion of the relationship network and the basic semantic request data, obtain fused semantic request data.

5. The method of the data query tool based on the AI language model according to claim 1, wherein, The inputting of the fused semantic request data into the professional knowledge layer of the language parsing model to obtain professional semantic request data includes: Identify multiple professional fields involved in the fused semantic request data through semantic analysis to obtain multiple professional knowledge ontologies; Map the features in the fused semantic request data to the features of the multiple professional knowledge ontologies through an ontology mapping algorithm to obtain multiple knowledge fragments related to the fused semantic request data; Calculate the semantic similarity between the knowledge fragments in the multiple professional knowledge ontologies and the fused semantic request data; Based on the semantic similarity, perform weighted fusion of the multiple knowledge fragments and the fused semantic request data to obtain professional semantic request data.

6. The method of the data query tool based on the AI language model according to claim 5, wherein, The calculation formula for the semantic similarity is: A = [a1, a2, …, a n ​ B = [b1, b2, …, b n ​ where A represents the fused semantic request data vector and B represents the knowledge fragment vector.

7. The method of the data query tool based on the AI language model according to claim 1, wherein Audit and verify the preliminary query result, including: checking whether the preliminary query result contains all the required data, verifying the data accuracy of the query result, and reviewing whether the preliminary query result complies with regulations.

8. A data query tool system based on an AI language model, characterized in that, Including: an acquisition unit: used to acquire the data query request input by the target user; A first processing unit: used to input the data query request into the basic language layer of the language parsing model to obtain basic semantic request data; A second processing unit: used to input the basic semantic data into the general knowledge layer of the language parsing model to obtain fused semantic request data; A third processing unit: used to input the fused semantic request data into the professional knowledge layer of the language parsing model to obtain professional semantic request data; the language parsing model is trained based on sample request data and their corresponding hierarchical semantic label results; A mapping unit: used to map the professional semantic request data to the database to obtain an SQL query statement; A search unit: used to access the database based on the SQL query statement to obtain a preliminary query result; A verification unit: used to audit and verify the preliminary query result to obtain a target query result.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for a data query tool based on an AI language model according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for a data query tool based on an AI language model according to any one of claims 1 to 7.

Citation Information

Cited By

  • Data compliance intelligent auditing method and device based on AI, equipment and medium

    CN121413754A

  • Information retrieval method and device, electronic equipment, storage medium and product

    CN121434454A