Data whole network search method and device based on artificial intelligence, equipment and medium
By adopting artificial intelligence-based data search methods in the entire network search of government data, the problem of government data being unable to be scaled, processed and standardized is solved, efficient and accurate data query and management is achieved, and the needs of intelligent processing and intelligent applications are adapted.
Patent Information
- Application Number
- CN202510109823.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-16
AI Technical Summary
Government data cannot be scaled, processed and standardized in search applications across the entire network, resulting in the speed of achieving results far from keeping up with the needs of intelligent processing and intelligent applications of government data.
The data search method based on artificial intelligence is adopted, including obtaining metadata from the data resource pool for preprocessing and standardization, using the knowledge graph to correlate the metadata with business entities and relationships, using the natural language processing model to parse the natural query language input by the user into structured query statements, and using the query language and inference model of the graph database to search in the metadata knowledge graph.
By clearing noise and redundant information in the data, correcting the wrong data format, clarifying the logical relationship between metadata, improving the integrity and consistency of data, greatly facilitating user operations, improving query efficiency and accuracy, and adapting to changing business needs and technological developments.
Smart Images

Figure CN120011610A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data processing technology, and in particular to an artificial intelligence-based data full-network search method, device, equipment and medium. Background Art
[0002] As digital transformation is further promoted in government departments, Yunshang Guizhou and various government departments have accumulated massive amounts of data covering different fields and business systems. These data come from a wide range of sources and in various formats. Metadata contains key information such as source, format, and update time, making it difficult to manage. For example, different departments may use different data formats to record dates, and the frequency of data updates is also uneven, which requires effective metadata collection and integration technology to ensure orderly management of data.
[0003] The inability to scale, streamline, and standardize is the fundamental pain point in the current application of government data search across the entire network. That is, the speed of achieving results is far behind the needs of intelligent processing and smart application of government data.
[0004] Therefore, there is an urgent need to propose an artificial intelligence-based data full-network search method to solve the technical problems that government data cannot be scaled, streamlined and standardized in full-network search applications. Summary of the invention
[0005] In order to overcome the problems existing in the related technologies, the present invention provides a method, device, equipment and medium for searching the entire network of data based on artificial intelligence to solve the technical problems in the related technologies that government data cannot be scaled, streamlined and standardized in the application of searching the entire network.
[0006] One or more embodiments of this specification provide a method for searching the entire network of data based on artificial intelligence, including the following steps:
[0007] Obtaining metadata from a data resource pool, and preprocessing and standardizing the metadata;
[0008] Using the knowledge graph to associate the metadata with business entities and relationships, and construct a metadata knowledge graph;
[0009] A natural language processing model is used to parse the natural query language input by the user into structured query statements;
[0010] The query language and reasoning model of the graph database are used to search in the metadata knowledge graph according to the parsed query statement, and the search results are returned.
[0011] Preferably, the method further comprises the following steps:
[0012] Visually present the search results according to the data type and characteristics of the search results;
[0013] Generate an intelligent report based on natural language according to the user's search requirements and the search results.
[0014] Preferably, the method further includes constructing a knowledge graph, which specifically includes the following steps:
[0015] Using a deep learning model to perform entity recognition and relationship extraction on the text content in the data resource pool to construct a knowledge graph;
[0016] When the data in the data resource pool is updated, the knowledge graph is automatically updated using an incremental update algorithm.
[0017] Preferably, the method of using a natural language processing model to parse the natural query language input by the user into a structured query statement specifically includes the following steps:
[0018] A natural language processing model is used to pre-process the natural query language input by the user, wherein the pre-processing includes word segmentation, part-of-speech tagging, and stop word removal, and the natural query language input by the user includes voice and / or text;
[0019] A sequence-to-sequence model based on deep learning is used to convert the natural query language input by users into structured query statements.
[0020] Preferably, the method of preprocessing the natural query language input by the user using a natural language processing model further includes the following steps:
[0021] Collect dialect speech and text data, build speech recognition models, and perform annotation and training;
[0022] Migrate the natural language processing model trained on Mandarin to dialects.
[0023] Preferably, the query language and reasoning model of the graph database is used to search in the metadata knowledge graph according to the parsed query statement and return the search results, and further includes the following steps:
[0024] According to the user's search history and preference information, collaborative filtering and / or content-based recommendation algorithms are used to sort and recommend the search results.
[0025] One or more embodiments of this specification provide a data network-wide search device based on artificial intelligence, including a metadata acquisition module, a metadata knowledge graph construction module, a parsing module, and a search module;
[0026] The metadata acquisition module is used to acquire metadata from the data resource pool and perform preprocessing and standardization on the metadata;
[0027] The metadata knowledge graph construction module is used to associate the metadata with business entities and relationships using the knowledge graph to construct a metadata knowledge graph;
[0028] The parsing module is used to parse the natural query language input by the user into a structured query statement using a natural language processing model;
[0029] The search module is used to use the query language and reasoning model of the graph database to search in the metadata knowledge graph according to the parsed query statement and return the search results.
[0030] Preferably, it also includes a display module, which is used to visually present the search results according to the data type and characteristics of the search results;
[0031] Generate an intelligent report based on natural language according to the user's search requirements and the search results.
[0032] One or more embodiments of the present specification provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned artificial intelligence-based data full-network search method when executing the computer program.
[0033] One or more embodiments of the present specification provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned artificial intelligence-based data full-network search method are implemented.
[0034] The present disclosure provides an artificial intelligence-based data full-network search method, device, equipment and medium, which have the advantages of obtaining metadata from a data resource pool, preprocessing and standardizing the metadata, effectively removing noise and redundant information in the data, and correcting erroneous data formats; using a knowledge graph to associate the metadata with business entities and relationships, constructing a metadata knowledge graph, clarifying the logical relationship between the metadata, making the organization and management of data more orderly, and improving the integrity and consistency of the data; using a natural language processing model to parse the natural query language input by the user into a structured query statement, which greatly facilitates user operation. The user does not need to master complex query syntax, but only needs to express the needs in daily language, and the system can accurately understand and convert them into executable query instructions; using the query language and reasoning model of the graph database to search in the metadata knowledge graph according to the parsed query statement, and returning the search results, it can quickly locate the relevant nodes and relationships, and the reasoning model can also mine implicit information, quickly and accurately return the results, improve the query efficiency and accuracy, and adopt a modular design. Each part is relatively independent and cooperates with each other, so that the system has good flexibility and scalability and can adapt to changing business needs and technological development. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0036] Figure 1 A flowchart of a method for searching the entire network of data based on artificial intelligence provided for one or more embodiments of this specification;
[0037] Figure 2 A system flow chart for one or more embodiments of this specification;
[0038] Figure 3 A schematic diagram of the structure of an artificial intelligence-based data network-wide search device provided in one or more embodiments of this specification;
[0039] Figure 4 A schematic diagram of the structure of a computer device provided for one or more embodiments of this specification. DETAILED DESCRIPTION
[0040] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0041] The present invention is described in detail below in conjunction with specific implementation methods and the accompanying drawings.
[0042] Method Embodiment
[0043] According to an embodiment of the present invention, a method for searching the entire network of data based on artificial intelligence is provided. Figure 1 As shown, it is a schematic diagram of the process of the data full network search method based on artificial intelligence provided in this embodiment. According to the data full network search method based on artificial intelligence in this embodiment of the present invention, the following steps are included:
[0044] S110. Obtain metadata from the data resource pool, pre-process and standardize the metadata. Specifically, automatically extract metadata from the three major systems of Yunshang Guizhou and the data areas of various government departments, including the source, format, update time, data owner and other information of the data. Integrate these metadata into a unified metadata warehouse and use distributed databases such as Apache Cassandra to store them to ensure high availability and scalability. Use data cleaning algorithms to remove noise and redundant information in metadata. For example, in the case of inconsistent data formats, they are standardized into a unified format through regular expression matching and conversion rules. At the same time, establish metadata standard specifications to ensure the semantic and structural consistency of metadata from different sources.
[0045] S120, using the knowledge graph to associate the metadata with business entities and relationships, and construct a metadata knowledge graph. For example, the data of government departments are associated with corresponding business processes, policies and regulations, etc., to enrich the connotation of metadata and provide richer background information for subsequent search and analysis.
[0046] S130: Use a natural language processing model to parse the natural query language input by the user into a structured query statement.
[0047] S140, using the query language and reasoning model of the graph database to search in the metadata knowledge graph according to the parsed query statement, and return the search results. For example, when a user queries "which companies are affected by a certain policy", the reasoning function of the knowledge graph can quickly find the corporate entities related to the policy and the relationship between them, providing the user with an accurate answer.
[0048] Specifically, based on distributed computing frameworks such as Apache Hadoop or Apache Spark, a distributed index is constructed for the data resource pool and the data areas of various government departments. Technologies such as inverted index are used to map the keywords of the data with the storage location of the data, thereby improving the speed and accuracy of the search.
[0049] At the same time, when new data enters the system, the index can be quickly updated and included in the search scope. The incremental indexing technology is used to only index the newly added data instead of rebuilding the entire index, which improves the real-time performance and efficiency of the system.
[0050] The method provided in this embodiment can effectively remove noise and redundant information in the data and correct erroneous data formats by obtaining metadata from a data resource pool and preprocessing and standardizing the metadata. The metadata is associated with business entities and relationships by using a knowledge graph to construct a metadata knowledge graph, clarifying the logical relationship between metadata, making data organization and management more orderly, and improving data integrity and consistency. The natural language processing model is used to parse the natural query language input by the user into a structured query statement, which greatly facilitates user operation. The user does not need to master complex query syntax, but only needs to express the requirements in daily language, and the system can accurately understand and convert them into executable query instructions. The query language and reasoning model of the graph database are used to search in the metadata knowledge graph according to the parsed query statement, and the search results are returned, which can quickly locate relevant nodes and relationships. The reasoning model can also mine implicit information, quickly and accurately return results, and improve query efficiency and accuracy. The modular design is adopted, and each part is relatively independent and cooperative, so that the system has good flexibility and scalability and can adapt to changing business needs and technological development.
[0051] In one embodiment, the following steps are also included:
[0052] S150. According to the data type and characteristics of the search results, develop rich data visualization components, including bar charts, line charts, pie charts, maps, word clouds, etc., to visualize the search results. For time series data, use line charts to show its changing trends; for geographic distribution data, use maps for visualization.
[0053] Generate intelligent reports based on natural language according to the user's search needs and search results. The report content includes data overview, analysis of key indicators, trend forecasts, etc. For example, for the search "query economic development data of a certain region in the past five years", the generated report can include analysis of the region's GDP, industrial structure, growth rate and other data in the past five years, and predict future economic development trends.
[0054] Special difference analysis algorithms are developed for difference analysis based on three types of comparison objects: time, division, and entity. For example, for difference analysis of time series data, year-on-year and month-on-month methods are used to calculate the difference values, and the difference points are highlighted through visualization; for difference analysis of division data, cluster analysis, principal component analysis and other methods are used to find the difference characteristics between different divisions.
[0055] The method provided in this embodiment quickly conveys information through visualization based on data characteristics, generates intelligent reports in natural language, lowers the threshold for understanding, clearly presents the overall picture and analysis of data to users, assists in efficient and accurate decision-making, meets the habits of different users, provides multiple data acquisition forms, and improves the user experience.
[0056] In one embodiment, it also includes constructing a knowledge graph, which specifically includes the following steps:
[0057] A deep learning model, such as BERT (Bidirectional Encoder Representations from Transformers), is used to perform entity recognition and relationship extraction on the text content in the data resource pool. For example, entities such as institution names, personnel names, and events are identified in government documents, and their hierarchical relationships, participation relationships, etc. are extracted. The extracted entities and relationships are stored in the graph database Neo4j to build a knowledge graph.
[0058] A real-time monitoring mechanism is established. When the data in the data resource pool is updated, the update process of the knowledge graph is automatically triggered. The incremental update algorithm is used to automatically update the knowledge graph, and only the changed parts are updated to improve the update efficiency. At the same time, the quality of the knowledge graph is regularly evaluated to detect and repair errors and missing information in the graph.
[0059] The method provided in this embodiment uses deep learning to mine text, extract entities and relationships, construct a knowledge graph, and present the intrinsic connections of data. It can efficiently utilize computing resources, maintain the timeliness and accuracy of knowledge, and continue to provide reliable support for various applications.
[0060] In one embodiment, a natural language processing model is used to parse the natural query language input by the user into a structured query statement, which specifically includes the following steps:
[0061] A natural language processing model is used to preprocess the natural query language input by the user, and the preprocessing includes word segmentation, part-of-speech tagging, and stop word removal. The natural query language input by the user includes voice and / or text, wherein the natural language processing model is a function implemented using an open source natural language processing toolkit such as NLTK (Natural Language Toolkit) or Stanford NLP. An end-to-end voice wake-up model based on deep learning is used, such as a combination model of convolutional neural networks (CNN) and long short-term memory networks (LSTM). Through a large amount of voice data training, the model can accurately identify specific wake-up words, such as "Guizhou Search", etc., to achieve a low-power voice wake-up function.
[0062] A sequence-to-sequence model based on deep learning is used to convert the natural query language input by the user into a structured query statement. For example, a natural language query such as "query the fiscal revenue of Guiyang City last year and compare it with the year before last" is parsed into a structured query containing information such as time (last year, the year before last), region (Guiyang City), indicator (fiscal revenue) and operation (query, comparison).
[0063] The following steps are also included:
[0064] For dialects such as Guizhou dialect and Sichuan dialect, we collect dialect speech and text data, build speech recognition models, and perform annotation and training. The speech recognition model uses a combination of acoustic models and language models. The acoustic model is used to extract the acoustic features of speech, and the language model is used to model the semantics of speech. Through training with a large amount of dialect speech data, the accuracy of dialect speech recognition is improved. For example, in Guizhou dialect speech recognition, the special vocabulary and pronunciation characteristics in Guizhou dialect can be accurately identified.
[0065] Using the transfer learning method, the natural language processing model trained on Mandarin can be migrated to dialects to improve the recognition and processing capabilities of dialects.
[0066] The method provided in this embodiment purifies input information and improves data quality by performing word segmentation, part-of-speech tagging and stop word removal on natural query language in speech and text forms, and accurately converts natural query language into structured query statements with the help of a deep learning sequence-to-sequence model, which facilitates system understanding and execution and realizes efficient human-computer interaction.
[0067] In one embodiment, searching in the metadata knowledge graph according to the parsed query statement using the query language and reasoning model of the graph database and returning the search results also includes the following steps:
[0068] According to the user's search history and preference information, collaborative filtering and / or content-based recommendation algorithms are used to sort and recommend the search results. For example, when the user often searches for education-related content, when the search results are returned, education-related data is ranked first and related education data resources are recommended.
[0069] At the same time, deep learning-based speech synthesis technology, such as Tacotron and WaveNet models, is used to convert search results into natural and fluent voice broadcasts. It supports a variety of timbres and speech speeds to meet the needs of different users. For example, users can choose male or female voices to broadcast search results, and can adjust the speech speed.
[0070] like Figure 2 The flowchart of the system provided in this embodiment is shown in FIG. 1 , which is further described by a specific implementation case below:
[0071] First, metadata including population information, enterprise information, financial data, and land resource data were extracted from the business systems of various departments of the municipal government through metadata collection tools, involving more than 20 data sources. These metadata were integrated into the metadata warehouse, and after cleaning and standardization, a unified metadata specification was formed. For example, in the population information data, the age, gender, and other information recorded by different departments were unified in format and deduplicated.
[0072] Then, using knowledge graph technology, we linked population information with enterprise information, land resource information, etc., and built a government knowledge graph. For example, we linked the registered address of a company with the corresponding land resource information, and linked the corporate legal person with population information, forming a comprehensive government data association network.
[0073] Citizens can use voice or text to input natural language queries, such as "Query the tax payment of industrial enterprises in this city last year and compare it with the year before." The natural language processing engine parses the query into a structured query statement, and the intelligent search engine searches the government data resource pool based on the parsed query and returns relevant results. For example, the search results include the total tax payment of industrial enterprises in this city last year, the number of tax-paying enterprises, the list of major tax-paying enterprises, and other information, and compares and analyzes the data with the previous year.
[0074] The data interpretation and intelligent reporting module automatically generates an intelligent report based on the search results. The report shows the comparison of the total tax paid by industrial enterprises last year and the year before in the form of a bar chart, and analyzes the reasons for the change in tax amount, such as industrial policy adjustments, business conditions, etc. In addition, the report also predicts the tax trend of industrial enterprises in the city in the future and puts forward relevant policy recommendations.
[0075] Citizens can use voice or text to input natural language queries, such as "Query the tax payment of industrial enterprises in this city last year and compare it with the year before." The natural language processing engine parses the query into a structured query statement, and the intelligent search engine searches the government data resource pool based on the parsed query and returns relevant results. For example, the search results include the total tax payment of industrial enterprises in this city last year, the number of tax-paying enterprises, the list of major tax-paying enterprises, and other information, and compares and analyzes the data with the previous year.
[0076] The data interpretation and intelligent reporting module automatically generates an intelligent report based on the search results. The report shows the comparison of the total tax paid by industrial enterprises last year and the year before in the form of a bar chart, and analyzes the reasons for the change in tax amount, such as industrial policy adjustments, business conditions, etc. In addition, the report also predicts the tax trend of industrial enterprises in the city in the future and puts forward relevant policy recommendations.
[0077] Device Embodiment
[0078] According to an embodiment of the present invention, a data network-wide search device based on artificial intelligence is provided. Figure 3 As shown, it is a structural schematic diagram of the artificial intelligence-based data full-network search device provided in this embodiment. According to the artificial intelligence-based data full-network search device of the embodiment of the present invention, it includes a metadata acquisition module 31, a metadata knowledge graph construction module 32, a parsing module 33 and a search module 34.
[0079] The metadata acquisition module 31 is used to acquire metadata from the data resource pool and perform preprocessing and standardization on the metadata.
[0080] The metadata knowledge graph construction module 32 is used to use the knowledge graph to associate the metadata with business entities and relationships to construct a metadata knowledge graph.
[0081] The parsing module 33 is used to parse the natural query language input by the user into a structured query statement using a natural language processing model.
[0082] The search module 34 is used to use the query language and reasoning model of the graph database to search in the metadata knowledge graph according to the parsed query statement and return the search results.
[0083] In the device provided in this embodiment, the metadata acquisition module 31 obtains metadata from the data resource pool, pre-processes and standardizes the metadata, and can effectively remove noise and redundant information in the data and correct erroneous data formats; the metadata knowledge graph construction module 32 uses the knowledge graph to associate the metadata with business entities and relationships, constructs the metadata knowledge graph, clarifies the logical relationship between the metadata, makes the organization and management of the data more orderly, and improves the integrity and consistency of the data; the parsing module 33 uses the natural language processing model to parse the natural query language input by the user into a structured query statement, which greatly facilitates the user operation. The user does not need to master complex query syntax, but only needs to express the needs in daily language, and the system can accurately understand and convert them into executable query instructions; the search module 34 uses the query language and reasoning model of the graph database to search in the metadata knowledge graph according to the parsed query statement, and returns the search results, which can quickly locate the relevant nodes and relationships. The reasoning model can also mine implicit information, quickly and accurately return the results, and improve the query efficiency and accuracy. The modular design is adopted, and each part is relatively independent and cooperates with each other, so that the system has good flexibility and scalability, and can adapt to the ever-changing business needs and technological development.
[0084] In one embodiment, a display module 35 is also included, which is used to visualize the search results according to the data type and characteristics of the search results, and generate intelligent reports based on natural language according to the user's search needs and search results.
[0085] The device provided in this embodiment quickly conveys information through visualization based on data characteristics, generates intelligent reports in natural language, lowers the threshold for understanding, clearly presents the overall picture and analysis of data to users, assists in efficient and accurate decision-making, meets the habits of different users, provides multiple data acquisition forms, and improves the user experience.
[0086] The embodiment of the present invention is an apparatus embodiment corresponding to the above-mentioned method embodiment. The specific operations of the processing steps of each module can be understood by referring to the description of the method embodiment, which will not be repeated here.
[0087] like Figure 4 As shown, the present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for searching the entire network for data in the above-mentioned embodiment is implemented; or when the computer program is executed by a processor, the method for searching the entire network for data based on artificial intelligence in the above-mentioned embodiment is implemented.
[0088] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0089] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and the contents not described in detail in the specification of the present invention belong to the common knowledge of those skilled in the art.
Claims
1. A data network-wide search method based on artificial intelligence, characterized in that: The following steps are involved: Obtaining metadata from a data resource pool, and preprocessing and standardizing the metadata; Using the knowledge graph to associate the metadata with business entities and relationships, and construct a metadata knowledge graph; A natural language processing model is used to parse the natural query language input by the user into structured query statements; The query language and reasoning model of the graph database are used to search in the metadata knowledge graph according to the parsed query statement, and the search results are returned.
2. The method for searching the entire network of data based on artificial intelligence according to claim 1, characterized in that: The following steps are also included: Visually present the search results according to the data type and characteristics of the search results; Generate an intelligent report based on natural language according to the user's search requirements and the search results.
3. The method for searching the entire network of data based on artificial intelligence according to claim 1, characterized in that: It also includes building a knowledge graph, which includes the following steps: Using a deep learning model to perform entity recognition and relationship extraction on the text content in the data resource pool to construct a knowledge graph; When the data in the data resource pool is updated, the knowledge graph is automatically updated using an incremental update algorithm.
4. The method for searching the entire network of data based on artificial intelligence according to claim 1, characterized in that: The method of using a natural language processing model to parse the natural query language input by the user into a structured query statement specifically includes the following steps: A natural language processing model is used to pre-process the natural query language input by the user, wherein the pre-processing includes word segmentation, part-of-speech tagging, and stop word removal, and the natural query language input by the user includes voice and / or text; A sequence-to-sequence model based on deep learning is used to convert the natural query language input by users into structured query statements.
5. The method for searching the entire network of data based on artificial intelligence according to claim 4, characterized in that: The method of preprocessing the natural query language input by the user using a natural language processing model also includes the following steps: Collect dialect speech and text data, build speech recognition models, and perform annotation and training; Migrate the natural language processing model trained on Mandarin to dialects.
6. The method for searching the entire network of data based on artificial intelligence according to claim 1, characterized in that: The query language and reasoning model of the graph database are used to search in the metadata knowledge graph according to the parsed query statement, and the search results are returned, and the following steps are also included: According to the user's search history and preference information, collaborative filtering and / or content-based recommendation algorithms are used to sort and recommend the search results.
7. A data network-wide search device based on artificial intelligence, characterized in that: It includes metadata acquisition module, metadata knowledge graph construction module, parsing module and search module; The metadata acquisition module is used to acquire metadata from the data resource pool and perform preprocessing and standardization on the metadata; The metadata knowledge graph construction module is used to associate the metadata with business entities and relationships using the knowledge graph to construct a metadata knowledge graph; The parsing module is used to parse the natural query language input by the user into a structured query statement using a natural language processing model; The search module is used to use the query language and reasoning model of the graph database to search in the metadata knowledge graph according to the parsed query statement and return the search results.
8. The data network-wide search device based on artificial intelligence as claimed in claim 7, characterized in that: It also includes a display module for visually presenting the search results according to the data type and characteristics of the search results; Generate an intelligent report based on natural language according to the user's search requirements and the search results.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the artificial intelligence-based data full-network search method as described in any one of claims 1 to 6.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the artificial intelligence-based data full-network search method as described in any one of claims 1 to 6 are implemented.