Structured Data Processing Method, Apparatus and Server Based on Large Model

By introducing large models into data processing, automatically analyzing natural language queries and generating structured query languages, the problems of low efficiency and high cost of structured data processing in the existing technology are solved, and efficient and accurate data retrieval and processing are achieved.

CN119807270BActive Publication Date: 2025-06-17INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510286906.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-17
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The prior art is inefficient and costly when processing structured data, especially due to changes in user requirements that require frequent development and debugging of application interfaces.

Method used

The structured data processing method based on the big model is adopted, and the user's data query request is received, and the natural language description is analyzed using the big model, and the semantic representation data is automatically obtained and a structured query language is generated, and then searched in the pre-built database to obtain the query results.

Benefits of technology

This method quickly converts user query intentions into database query statements through the combination of large models and databases, significantly improving the efficiency of processing structured query languages, reducing investment costs, and improving data accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807270B_ABST
    Figure CN119807270B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, and server for processing structured data based on a large model, relating to the technical field of data processing. Extract a structured data set from the collected metadata to be processed, and import the structured data set into a database. Analyze the data query text in the form of natural language description input by the user through a large model to automatically obtain semantic representation data and generate a structured query language. Use the structured query language to retrieve in the pre-constructed database to obtain a query result. Through the combination of the large model and the database, utilize the ability of the database to process structured data to improve the accuracy of data, convert the user's query intention into a database query statement, and quickly retrieve the query result from the database, omitting the research and development work of the front-end application program, improving the efficiency of processing the structured query language, and reducing the input cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular to a method, apparatus, and server for processing structured data based on a large model. Background Art

[0002] With the rapid development of information technology, big data has been deeply integrated into all walks of life and has become an indispensable important resource in all walks of life. Data has various forms and complex structures, covering various types such as structured, semi-structured, and unstructured. How to extract the structured data required by users from the vast and diverse data has become an important challenge currently faced.

[0003] Currently, related technologies develop application interfaces through R & D personnel to meet the structured data required by users. However, user needs are constantly changing. To cope with the changing needs, R & D personnel need to develop application interfaces and debug them irregularly, resulting in low efficiency in processing structured data and high input costs. Summary of the Invention

[0004] This application provides a method, apparatus, and server for processing structured data based on a large model, so as to at least solve the problems of low efficiency in processing structured data and high input costs in related technologies.

[0005] This application provides a method for processing structured data based on a large model, including:

[0006] Receiving a data query request sent by a user terminal; the data query request includes a data query text and a data result display method;

[0007] Obtaining text data to be parsed according to the data query text;

[0008] Inputting the text data request to be parsed into a large model, so that the large model performs the following steps:

[0009] Judging whether there is text data to be parsed in the cache;

[0010] If it is determined that there is no text data to be parsed in the cache, obtaining semantic representation data according to the text data to be parsed;

[0011] Identifying key data required to generate a structured query language according to the semantic representation data;

[0012] Structurally converting the key data according to a preset template to obtain structured key data;

[0013] Generating a structured query language corresponding to the semantic representation data according to the structured key data;

[0014] Retrieve in a pre-built database using Structured Query Language to obtain the query result corresponding to the data query request;

[0015] Send the query result to the client according to the data result display method so that the client can display the query result.

[0016] This application also provides a structured data processing device based on a large model, including:

[0017] A receiving module for receiving a data query request sent by the client; the data query request includes a data query text and a data result display method;

[0018] An obtaining module for obtaining text data to be parsed according to the data query text;

[0019] An input module for inputting the text data request to be parsed into the large model, enabling the large model to perform the following steps:

[0020] Among them, the input module includes:

[0021] A judgment unit for judging whether there is text data to be parsed in the cache;

[0022] A first obtaining unit for, if it is determined that there is no text data to be parsed in the cache, obtaining semantic representation data according to the text data to be parsed;

[0023] An identification unit for identifying the key data required to generate Structured Query Language according to the semantic representation data;

[0024] A second obtaining unit for structurally converting the key data according to a preset template to obtain structured key data;

[0025] A generating unit for generating Structured Query Language corresponding to the semantic representation data according to the structured key data;

[0026] A retrieval module for retrieving in a pre-built database using Structured Query Language to obtain the query result corresponding to the data query request;

[0027] A sending module for sending the query result to the client according to the data result display method so that the client can display the query result.

[0028] This application also provides a server, including: a memory for storing a computer program; a processor for implementing the steps of any of the above-mentioned structured data processing methods based on a large model when executing the computer program.

[0029] The present application also provides a computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps of any of the above-mentioned large model-based structured data processing methods.

[0030] The present application also provides a computer program product including a computer program, where the computer program, when executed by a processor, implements the steps of any of the above-mentioned large model-based structured data processing methods.

[0031] Through the present application, a structured data set is extracted from the metadata to be processed, and the structured data set is imported into a database. By utilizing the database's ability to process structured data, the accuracy of the data is improved. The large model analyzes the data query text in the form of natural language description input by the user, automatically obtains semantic representation data, and generates a structured query language. Using the structured query language, a search is performed in the pre-constructed database to obtain a query result. By combining the large model with the database, the user's query intention is converted into a database query statement, and the query result is quickly retrieved from the database, omitting the research and development work of the front-end application program, improving the efficiency of processing the structured query language, and reducing the input cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 It is a schematic diagram of the scenario of the large model-based structured data processing method provided by the embodiment of the present application;

[0034] Figure 2 It is a schematic flow diagram of the large model-based structured data processing method provided by the embodiment of the present application;

[0035] Figure 3 It is a schematic structural diagram of the large model-based structured data processing device provided by the embodiment of the present application;

[0036] Figure 4 It is a schematic structural diagram of the server provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0038] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0039] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0040] First, the nouns involved in the present application are explained:

[0041] Structured Query Language: A standard programming language for managing and operating relational databases, which can access and process databases, including data insertion, query, update and deletion.

[0042] In order to solve the problem that the requirements for structured data needed by users in the prior art are constantly changing, which requires R & D personnel to develop application interfaces irregularly and conduct debugging, resulting in low efficiency in processing structured data and high input costs. The embodiments of the present application propose the following technical concept: The inventor thought of the method of combining a database and a large model. Extract a structured data set from the metadata to be processed and import the structured data set into the database. Analyze the data query text in the form of natural language description input by the user through the large model, automatically obtain semantic representation data, and generate a structured query language. Use the structured query language to retrieve in the pre-constructed database to obtain the query result. By combining the large model with the database, the user's query intention is converted into a database query statement, and the query result is quickly retrieved from the database, omitting the research and development work of the front-end application program, improving the efficiency of processing the structured query language, and reducing the input cost.

[0043] To enable those skilled in the art of the present technology to better understand the solution of this application, the following further elaborates on this application in combination with the accompanying drawings and specific implementation manners.

[0044] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the structured data processing method based on the large model depends, the specific application environment architecture or specific hardware architecture is described herein.

[0045] Refer to Figure 1 , Figure 1 which is a schematic diagram of the scenario of the structured data processing method based on the large model provided by the embodiment of this application. The server provided by this embodiment includes: a receiving device 101, a processor 102, and a display device 103.

[0046] It can be understood that the structure schematically shown in the embodiment of this application does not constitute a specific limitation on the structured data processing method based on the large model. In some other feasible implementation manners of this application, the above architecture may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements, which can be specifically determined according to the actual application scenario and are not limited herein. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0047] In the specific implementation process, the receiving device 101 can be an input / output interface or a communication interface, and can receive the data query request sent by the user end.

[0048] The processor 102 can perform a series of processes on the data query text in the data query request to obtain the query result.

[0049] The display device 103 can be used to display the query result.

[0050] It should be understood that the above processor can be implemented by the processor reading instructions in the memory and executing the instructions, or can be implemented by a chip circuit.

[0051] In addition, the network architecture and business scenario described in the embodiment of this application are for more clearly explaining the technical solution of the embodiment of this application, and do not constitute a limitation on the technical solution provided by the embodiment of this application. Those skilled in the art know that with the evolution of the network architecture and the emergence of new business scenarios, the technical solution provided by the embodiment of this application is equally applicable to similar technical problems.

[0052] Figure 2 which is a schematic flowchart of the structured data processing method based on the large model provided by the embodiment of this application, as Figure 2As shown, the embodiment of the present application provides a structured data processing method based on a large model, and the method is described in detail as follows:

[0053] S201: receiving a data query request sent by a user terminal; wherein the data query request includes a data query text and a data result display method.

[0054] S202: Acquire text data to be parsed according to the data query text.

[0055] Specifically, noise data in the data query text is deleted to obtain initial text data to be parsed; the initial text data to be parsed is divided into multiple words; stop words in the multiple words are removed to obtain the text data to be parsed.

[0056] In this embodiment, once the user enters and submits a data query request in the front-end interface of the user end, the data query text will be captured immediately. The data query text is denoised, segmented, and stop words are removed. The data query text may contain various irrelevant characters and noise. The noise data in the data query text is deleted to obtain the initial text data to be parsed; through the word segmentation technology, the initial text data to be parsed is divided into meaningful vocabulary units to obtain multiple words. In the word segmentation results, there are some words that have no substantial contribution to subsequent analysis, that is, stop words, such as "的", "了" and "在". Although these words appear frequently in the data query text, they do not carry information themselves.

[0057] Optionally, word segmentation techniques include but are not limited to dictionary-based word segmentation, statistics-based word segmentation, and deep learning word segmentation methods.

[0058] S203: Input the text data request to be parsed into the big model, so that the big model executes the following steps.

[0059] Optionally, the big model supports the fusion analysis of structured and unstructured data, provides a unified data processing framework, and enhances the comprehensiveness of the analysis.

[0060] Specifically, step S203 includes S2031 to S2035:

[0061] S2031: Determine whether there is text data to be parsed in the cache.

[0062] Specifically, a hash calculation is performed on the text data to be parsed to obtain a hash identifier of the text data to be parsed; it is determined whether the hash identifier exists in the cache; if the hash identifier does not exist in the cache, it is determined that the text data to be parsed does not exist in the cache.

[0063] In this embodiment, if the hash identifier exists in the cache, it is determined that there is text data to be parsed in the cache; and the structured query language corresponding to the hash identifier is obtained from the cache.

[0064] In this embodiment, key-value pairs are stored in the cache, where the key is the text data to be parsed and the value is the structured query language corresponding to the text data to be parsed. If there is a hash identifier corresponding to the text data to be parsed, it indicates that there is the text data to be parsed and the corresponding structured query language in the cache, and then the corresponding structured query language can be directly obtained without regenerating the structured query language.

[0065] S2032: If it is determined that the text data to be parsed does not exist in the cache, semantic representation data is obtained according to the text data to be parsed.

[0066] Specifically, the formula for obtaining semantic representation data according to the text data to be parsed is:

[0067]

[0068] In the formula, represents the semantic representation data, represents the text data to be parsed, represents the parameters of the large model; represents the probability of obtaining semantic representation data when the text data to be parsed and the parameters of the large model are input; represents selecting the semantic representation data with the highest probability.

[0069] In this embodiment, the large model will generate multiple semantic representation data based on the text data to be parsed and the parameters of the large model, and find the semantic representation data with the highest probability from all the semantic representation data as the final semantic representation data.

[0070] S2033: Identify the key data required to generate the structured query language according to the semantic representation data.

[0071] Among them, the key data includes keywords, operation types, and query conditions.

[0072] Specifically, step S2033 includes Sa~Se:

[0073] Sa: Segment the semantic representation data to obtain multiple semantic representation words.

[0074] Sb: Perform part-of-speech tagging on the multiple semantic representation words to obtain multiple part-of-speech-tagged semantic representation words.

[0075] In this embodiment, the part of speech of each semantic representation word is tagged to determine categories such as nouns, verbs, adjectives, and adverbs. Through part-of-speech tagging, the role of the semantic representation word in the sentence can be initially judged. For example, a noun may be a table name, a field name, or an entity in a condition, and a verb may represent an operation type.

[0076] Sc: Identify keywords from multiple labeled semantic representation words according to the pre-built database keywords.

[0077] In this embodiment, the pre-built database keywords include all table names and field names. Compare the labeled semantic representation words with the pre-built database keywords. If there is a match, identify it as a table name or a field name.

[0078] Sd: Identify the operation type from multiple labeled semantic representation words according to the pre-built operation type list.

[0079] In this embodiment, the pre-built operation type list includes queries, insertions, updates, deletions, etc., which are used to determine the operation type that the user wants to perform. When the user input contains "query", the operation type can be identified as a query.

[0080] Se: Identify the query conditions from multiple labeled semantic representation words according to the pre-built query condition thesaurus.

[0081] In this embodiment, the pre-built query condition thesaurus includes words such as "greater than", "less than", "equal to", "contains", "maximum", and "minimum" that represent conditional relationships or limited ranges. Match the labeled semantic representation words with the pre-built query condition thesaurus to identify the query conditions for constructing structured query language.

[0082] S2034: Structurally transform the key data according to a preset template to obtain structured key data.

[0083] Optionally, the preset template can be JSON, Schema, Pydantic models, etc.

[0084] S2035: Generate the structured query language corresponding to the semantic representation data according to the structured key data.

[0085] Specifically, the formula for generating the structured query language corresponding to the semantic representation data according to the structured key data is:

[0086]

[0087] In the formula, represents the structured query language, represents the semantic representation data, represents the structured key data, represents the parameters of the function for generating the structured query language, represents the probability of generating the structured query language when inputting the semantic representation data, the structured key data, and the parameters of the function for generating the structured query language; Represents the Structured Query Language with the highest selection probability.

[0088] In this embodiment, the formulaic steps for generating the Structured Query Language are as follows:

[0089] SQL = "Operation + " + "Columns + "FROM" + "Table + "WHERE" + "Conditions"

[0090] Wherein, Operation represents the operation type, Columns represents the field name to be queried, Table represents the table name to be queried, and Conditions represents the query condition.

[0091] Exemplarily, the user input is: What is the best speccpu score released by Manufacturer A? The identified key data includes speccpu, score, Manufacturer A, and best. Among them, the table name is speccpu, the field name is score, the operation type is query, and the query conditions are Manufacturer A and best. The generated Structured Query Language is SELECT max(score) FROM speccpu WHERE manufacturer = "A".

[0092] In this embodiment, based on the input semantic representation data, structured key data, and the parameters of the function for generating the Structured Query Language, multiple Structured Query Languages are generated, and the Structured Query Language with the highest probability is found from all the Structured Query Languages as the final Structured Query Language.

[0093] In this embodiment, after generating the Structured Query Language, hash calculation is performed on the text data to be parsed to obtain a hash identifier. According to the hash identifier, the text data to be parsed and the Structured Query Language are saved in the cache in the form of key-value pairs.

[0094] S204: Use the Structured Query Language to retrieve in the pre-constructed database to obtain the query result corresponding to the data query request.

[0095] Specifically, the formula for using the Structured Query Language to retrieve in the pre-constructed database to obtain the query result corresponding to the data query request is:

[0096]

[0097] In the formula, represents the query result, represents the pre-constructed database, represents the Structured Query Language, represents the database execution function.

[0098] In this embodiment, the database execution function sends the Structured Query Language to a pre-built database for execution to obtain the query result.

[0099] Optionally, an efficient protocol or interface can be designed to achieve a seamless connection between the large model and the database.

[0100] Specifically, the steps for the pre-built database include Sf~Sh:

[0101] Sf: Collect the metadata to be processed.

[0102] In this embodiment, the metadata to be processed is widely collected. These data have a wide range of sources, covering multiple channels, including but not limited to various websites, etc. Taking the statistics of server performance data as an example, the data may come from the official websites of server manufacturers, etc. The collection process needs to adapt to the formats and interfaces of different data sources. When scraping metadata from web pages, it is necessary to parse HTML or XML structures, etc.

[0103] Sg: Extract the structured data set from the metadata to be processed.

[0104] In this embodiment, the data formats of the metadata to be processed collected from different data sources may vary greatly. For the convenience of subsequent unified processing, the data needs to be converted into structured data. For example, the time format may be expressed differently in different data sources. Some are in the form of YYYY-MM-DD HH:MM:SS, and some are in the form of timestamps. All time data needs to be unified into one format, and at the same time, other data types such as numbers and strings are also structured.

[0105] In this embodiment, before extracting the structured data set, the metadata to be processed needs to be data-cleaned and de-duplicated. The metadata to be processed often contains noise data, such as irrelevant special characters and garbled codes, which will interfere with subsequent analysis and need to be identified and removed; duplicate data may appear during the collection process, and these duplicate data will occupy storage space and affect the accuracy and efficiency of data analysis. It is necessary to find duplicate records and retain only unique copies.

[0106] In this embodiment, to ensure the timeliness and freshness of the data, the metadata to be processed is refreshed daily. Ensure that the latest data acquisition situation can be reflected, so that subsequent steps can understand and process data based on accurate metadata.

[0107] Sh: Create an initial database and import the structured data set into the initial database to obtain the pre-built database.

[0108] In this embodiment, a database type is selected, a table structure is designed, and an index is created to create an initial database. Among them, the initial database also provides data retrieval functions such as querying, classification, statistics, and sorting, making data access fast and convenient. Among them, according to the characteristics of the structured data set and the query pattern, an optimized index is dynamically generated. The database also provides query performance monitoring and tuning functions.

[0109] Optionally, the database type includes but is not limited to MySQL, PostgreSQL, Oracle, etc.

[0110] In this embodiment, the structured data set is stored in the initial database in the form of fields according to the designed table structure, realizing the smooth connection and processing of the data stream.

[0111] In this embodiment, encryption and desensitization are required during data transmission and processing to ensure data security.

[0112] S205: Send the query result to the user side according to the data result display method so that the user side can display the query result.

[0113] Optionally, the data result display methods include tables, charts, natural language answers, etc.

[0114] In summary, a structured data set is extracted from the metadata to be processed and imported into the database. According to the data query text in the form of natural language description input by the user, the text data to be parsed is obtained, and the text data to be parsed is analyzed by a large model to automatically obtain semantic representation data, and then a structured query language is generated. Using the structured query language, retrieval is performed in the pre-constructed database to obtain the query result. Through the combination of the large model and the database, the ability of the database to process structured data is utilized to improve the accuracy of the data; the distributed computing ability of the large model is utilized to perform parallel processing on large-scale structured data, significantly improving the processing speed. The user's query intention is converted into a database query statement, and the query result is quickly retrieved from the database, omitting the research and development work of the front-end application program, improving the efficiency of processing the structured query language, and reducing the input cost. In addition, the large model automatically completes data processing and generates a structured query language without human intervention, further reducing the human input cost.

[0115] Optionally, before retrieving in the pre-constructed database using the structured query language in step S204 to obtain the query result corresponding to the data query request, an access control mechanism is added, which is described in detail as follows:

[0116] S301: Obtain the user identifier carried in the data query request.

[0117] S302: Obtain the user role according to the user identifier.

[0118] In this embodiment, users corresponding to different user identifiers have different roles, such as administrator, ordinary user, auditor, etc. Each role has different responsibilities and operation scopes. For example, the administrator role may have all operation permissions for all tables; an ordinary user may only be able to query data from some tables; an auditor may only be able to view the audit log table.

[0119] S303: Determine whether there is permission to execute Structured Query Language according to the user role.

[0120] In this embodiment, for different roles in the pre-built database, define their access permissions to various resources in the database, including operations such as query, insert, update, and delete.

[0121] Specifically, check the Structured Query Language according to the user role. The check content includes whether there is the execution permission of the operation type, and whether there is permission to access the tables and fields involved in the Structured Query Language. If the user role has the execution permission of this operation type, further determine whether there is permission to access the tables and fields involved in the Structured Query Language. If there is access permission, then there is permission to execute the Structured Query Language.

[0122] If there is no execution permission for the operation type and no permission to access the tables and fields involved in the Structured Query Language, reject the execution of the Structured Query Language and return a prompt message of insufficient permissions to the user side.

[0123] In summary, by adding an access control mechanism to determine whether there is permission to execute Structured Query Language according to the user role, unauthorized access is prevented, and the security of data is increased.

[0124] Based on the above embodiment, in this embodiment, the optimization of the generated Structured Query Language is introduced in detail as follows:

[0125] S401: Calculate the query complexity of the Structured Query Language.

[0126] Specifically, the formula for calculating the query complexity of the Structured Query Language is:

[0127]

[0128] In the formula, represents the query complexity of the Structured Query Language, represents the number of tables involved in the Structured Query Language query, represents the weight of the table; represents the number of fields involved in the Structured Query Language query, Represents the weight of a field; Represents the number of conditions in a Structured Query Language (SQL) query, Represents the weight of a condition.

[0129] In this embodiment, the query complexity is a quantified value used to measure the difficulty and resource consumption of executing SQL. The higher the complexity, the more computing resources and time are required to complete the query.

[0130] In this embodiment, the more tables involved in an SQL query, the more complex the table join operations the database needs to perform when executing the query, which may lead to a decline in query performance. For example, if an SQL needs to retrieve data from 2 tables, then ; the more fields involved in an SQL query, the greater the amount of data the database needs to read and process, which also increases the query complexity. For example, if the query needs to retrieve 3 fields, then ; the more conditions involved in an SQL query, the more judgment and comparison operations the database needs to perform when filtering data, and the query complexity will also increase accordingly. For example, the query condition is "age greater than 25 and department is department B", and there are two conditions here, .

[0131] In this embodiment, , and : are the weights of the table, field, and condition respectively. These weights are adjusted according to the specific business scenario, reflecting the relative importance of the number of tables, fields, and conditions on the query complexity in that scenario. For example, in a business with high requirements for table association, the value of may be set relatively large, and in some scenarios mainly focusing on condition filtering, the weight of may be higher.

[0132] S402: Optimize the Structured Query Language according to the query complexity to obtain an optimized Structured Query Language.

[0133] In this embodiment, according to the query complexity, the Structured Query Language can be optimized to select a more efficient execution method, thereby improving data processing efficiency and reducing query execution time and resource consumption. For example, if it is found that the complexity of a certain Structured Query Language is relatively high and mainly due to the excessive number of tables involved, the tables can be considered for splitting or merging to reduce the query complexity.

[0134] Optionally, the database design can also be optimized according to the query complexity, such as adjusting the table structure, creating appropriate indexes, etc.

[0135] In summary, the complexity of different Structured Query Languages can be evaluated through query complexity, thereby optimizing the Structured Query Language and improving data processing efficiency.

[0136] Optionally, after sending the query result to the client, the query result feedback information of the user can be collected, which is described in detail as follows:

[0137] S501: Receive the query result feedback information sent by the client.

[0138] In this embodiment, the query result feedback information is cleaned to remove duplicate, invalid or incomplete query result feedback information. For example, for some content filled in or filled in randomly by users, as well as the same feedback information submitted by the same user multiple times, screening and processing are performed.

[0139] In this embodiment, the query result feedback information is classified according to different themes or types, such as performance issues and data accuracy, etc. At the same time, each feedback is marked for subsequent analysis and processing. For example, the inaccurate query result feedback by the user is marked as the data accuracy category. The slow query speed feedback by the user is marked as the performance issue category.

[0140] S502: Optimize the large model and / or the pre-built database according to the result feedback information.

[0141] In this embodiment, for the optimization of the large model, according to the results marked in step S501, the data related to the large model is sorted out. For example, if the user feedbacks that the answers of the large model in certain specific fields are inaccurate, the correct data and examples in the relevant fields can be collected for expanding and optimizing the training data set of the large model. The sorted feedback data is used to fine-tune the large model. The parameters of the large model are adjusted so that it can better handle the questions and requirements raised by users. For example, for the problems in natural language understanding feedback by users, the semantic understanding part of the large model is fine-tuned specifically. After training the large model with the feedback data, the effect is evaluated. By comparing the performance of the large model on relevant tasks before and after training, it is judged whether the training is effective. For example, compare whether the accuracy rate and recall rate of the large model in handling common problems in user feedback have improved.

[0142] In this embodiment, for the optimization of the database, according to the results marked in step S501, the data related to the database is sorted out. For example, if the user feedbacks that the query speed is slow, it is necessary to analyze the index structure, query statement, etc. of the database to find out the reasons for the slow query. Optimize the structure of the database, such as adjusting the table structure, adding or deleting indexes, etc. For example, if it is found that the data redundancy of a certain table is relatively high, the table can be split or merged to improve the efficiency of data storage and query.

[0143] In summary, by collecting the query result feedback information sent by the user terminal and optimizing the large model and / or the pre-built database according to the query result feedback information, the performance of the large model and the performance of the pre-built database can be improved, thereby enhancing the user experience.

[0144] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0145] Figure 3 This is a schematic structural diagram of a structured data processing device based on a large model provided by an embodiment of the present application. As Figure 3 shown, the embodiment of the present application also provides a structured data processing device based on a large model, including: a receiving module 301, an obtaining module 302, an input module 303, a retrieval module 304, and a sending module 305. The input module 303 includes a judgment unit 3031, a first obtaining unit 3032, an identification unit 3033, a second obtaining unit 3034, and a generating unit 3035.

[0146] The receiving module 301 is configured to receive a data query request sent by the user terminal; the data query request includes a data query text and a data result display method;

[0147] The obtaining module 302 is configured to obtain text data to be parsed according to the data query text;

[0148] The input module 303 is configured to input the text data request to be parsed into the large model, so that the large model performs the following steps:

[0149] Among them, the input module 303 includes:

[0150] The judgment unit 3031 is configured to judge whether there is text data to be parsed in the cache;

[0151] The first obtaining unit 3032 is configured to, if it is determined that there is no text data to be parsed in the cache, obtain semantic representation data according to the text data to be parsed;

[0152] The identification unit 3033 is configured to identify the key data required to generate a structured query language according to the semantic representation data;

[0153] The second obtaining unit 3034 is configured to perform structured conversion on the key data according to a preset template to obtain structured key data;

[0154] A generating unit 3035, configured to generate a structured query language corresponding to semantic representation data according to structured key data;

[0155] A retrieval module 304, configured to retrieve in a pre-constructed database by using the structured query language to obtain a query result corresponding to a data query request;

[0156] A sending module 305, configured to send the query result to a user terminal according to a data result display mode so that the user terminal can display the query result.

[0157] In a possible implementation manner, the structured data processing device based on a large model further includes: a construction module. The construction module includes:

[0158] An acquisition unit, configured to acquire metadata to be processed;

[0159] An extraction unit, configured to extract a structured data set from the metadata to be processed;

[0160] A creation unit, configured to create an initial database and import the structured data set into the initial database to obtain a pre-constructed database.

[0161] In a possible implementation manner, the acquisition module 302 includes:

[0162] A deletion unit, configured to delete noise data in the data query text to obtain initial text data to be parsed.

[0163] A splitting unit, configured to split the initial text data to be parsed into multiple words.

[0164] A rejection unit, configured to reject stop words in the multiple words to obtain text data to be parsed.

[0165] In a possible implementation manner, the formula for the first acquisition unit 3032 to acquire semantic representation data according to the text data to be parsed is:

[0166]

[0167] In the formula, represents semantic representation data, represents the text data to be parsed, represents the parameters of the large model; represents the probability of obtaining semantic representation data when the text data to be parsed and the parameters of the large model are input; represents selecting the semantic representation data with the maximum probability.

[0168] In a possible implementation manner, the key data includes keywords, operation types, and query conditions;

[0169] Correspondingly, the recognition unit 3033 includes:

[0170] The word segmentation sub-unit is used to segment the semantic representation data to obtain multiple semantic representation words;

[0171] The annotation sub-unit is used to perform part-of-speech annotation on multiple semantic representation words to obtain multiple annotated semantic representation words;

[0172] The first recognition sub-unit is used to identify keywords from multiple annotated semantic representation words according to the pre-constructed database keywords.

[0173] The second recognition sub-unit is used to identify the operation type from multiple annotated semantic representation words according to the pre-constructed operation type list;

[0174] The third recognition sub-unit is used to identify query conditions from multiple annotated semantic representation words according to the pre-constructed query condition thesaurus.

[0175] In a possible implementation manner, the generation unit 3035 generates the structured query language corresponding to the semantic representation data according to the structured key data, and the formula is:

[0176]

[0177] In the formula, represents the structured query language, represents the semantic representation data, represents the structured key data, represents the parameter of the function for generating the structured query language, represents the probability of generating the structured query language when inputting the semantic representation data, the structured key data, and the parameter of the function for generating the structured query language; represents selecting the structured query language with the highest probability.

[0178] In a possible implementation manner, the judgment unit 3031 includes:

[0179] The hash calculation sub-unit is used to perform hash calculation on the text data to be parsed to obtain the hash identifier of the text data to be parsed;

[0180] The judgment sub-unit is used to judge whether the hash identifier exists in the cache;

[0181] The first determination sub-unit is used to determine that the text data to be parsed does not exist in the cache if the hash identifier does not exist in the cache.

[0182] In a possible implementation manner, the structured data processing device based on the large model further includes: a judgment module, and the judgment module includes:

[0183] A second determination subunit, configured to determine that the text data to be parsed exists in the cache if a hash identifier exists in the cache;

[0184] An acquisition subunit, configured to acquire a Structured Query Language corresponding to the hash identifier from the cache.

[0185] In a possible implementation manner, the structured data processing apparatus based on a large model further includes: a first optimization module, and the first optimization module includes:

[0186] A complexity calculation subunit, configured to calculate the query complexity of the Structured Query Language;

[0187] A first optimization subunit, configured to optimize the Structured Query Language according to the query complexity to obtain an optimized Structured Query Language.

[0188] In a possible implementation manner, the formula for the complexity calculation subunit to calculate the query complexity of the Structured Query Language is:

[0189]

[0190] In the formula, represents the query complexity of the Structured Query Language, represents the number of tables involved in the Structured Query Language query, represents the weight of the table; represents the number of fields involved in the Structured Query Language query, represents the weight of the field; represents the number of conditions in the Structured Query Language query, represents the weight of the condition.

[0191] In a possible implementation manner, the formula for the retrieval module 304 to perform retrieval in the pre-constructed database by using the Structured Query Language to obtain a query result corresponding to a data query request is:

[0192]

[0193] In the formula, represents the query result, represents the pre-constructed database, represents the Structured Query Language, represents the database execution function.

[0194] In a possible implementation manner, the structured data processing apparatus based on a large model further includes: a second optimization module, and the second optimization module includes:

[0195] A receiving subunit, configured to receive query result feedback information sent by a user terminal;

[0196] A second optimization subunit, configured to optimize the large model and / or a pre-constructed database according to the result feedback information.

[0197] For the descriptions of the features in the corresponding embodiments of the structured data processing device based on the large model, reference may be made to the relevant descriptions in the corresponding embodiments of the structured data processing method based on the large model, which will not be elaborated here one by one.

[0198] Figure 4 This is a schematic structural diagram of the server provided in the embodiment of the present application. As Figure 4 shown, the server provided in this embodiment includes at least one processor 401 and a memory 402. Optionally, the server further includes a communication component 403. Among them, the processor 401, the memory 402, and the communication component 403 are connected through a bus.

[0199] In a specific implementation process, at least one processor 401 executes the computer execution instructions stored in the memory 402, so that at least one processor 401 executes the above-mentioned embodiment of the structured data processing method based on the large model.

[0200] For the specific implementation process of the processor 401, reference may be made to the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0201] In the above embodiment, it should be understood that the processor may be a central processing unit (Central Processing Unit, abbreviated as CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0202] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.

[0203] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0204] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any of the above-described embodiments of the structured data processing method based on a large model when running.

[0205] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc.

[0206] An embodiment of the present application also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the structured data processing method based on a large model.

[0207] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the structured data processing method based on a large model.

[0208] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. For the sake of clearly illustrating the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0209] The above has introduced in detail a method for processing structured data based on a large model provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A structured data processing method based on a large model, characterized in that: include: Receive a data query request sent by a user; The data query request includes a data query text and a data result display method; According to the data query text, obtain the text data to be parsed; The text data request to be parsed is input into the big model, so that the big model performs the following steps: Determine whether the text data to be parsed exists in the cache; If it is determined that the text data to be parsed does not exist in the cache, acquiring semantic representation data according to the text data to be parsed; According to the semantic representation data, identifying key data required for generating a structured query language, the key data including keywords, operation types and query conditions; The key data is structured according to a preset template to obtain structured key data; Generating a structured query language corresponding to the semantic representation data according to the structured key data; Using the structured query language, searching in a pre-built database to obtain a query result corresponding to the data query request; Sending the query result to the user terminal according to the data result display method, so that the user terminal displays the query result; The step of obtaining the text data to be parsed according to the data query text includes: Deleting noise data in the data query text to obtain initial text data to be parsed; Segmenting the initial text data to be parsed into multiple words; Stop words are eliminated from the multiple words to obtain text data to be parsed.

2. The method according to claim 1, characterized in that: Before receiving the data query request sent by the user terminal, the method further includes: Collect metadata to be processed; extracting a structured data set from the metadata to be processed; An initial database is created, and the structured data set is imported into the initial database to obtain a pre-constructed database.

3. The method according to claim 1, characterized in that The formula for obtaining semantic representation data according to the text data to be parsed is: In the formula, Representation semantics represents data, represents the text data to be parsed, Represents the parameters of the large model; Indicates the probability of obtaining the semantic representation data when the text data to be parsed and the parameters of the large model are input; Indicates the semantic representation data with the highest probability of selection.

4. The method according to claim 1, characterized in that: The step of identifying key data required for generating a structured query language according to the semantic representation data includes: Segmenting the semantic representation data to obtain a plurality of semantic representation words; Performing part-of-speech tagging on the multiple semantic representation words to obtain multiple tagged semantic representation words; According to the pre-constructed database keywords, identifying the keywords from the multiple annotated semantic representation words; According to a pre-built operation type list, identifying the operation type from the multiple annotated semantic representation words; The query condition is identified from the plurality of annotated semantic representation words according to a pre-constructed query condition word library.

5. The method according to claim 1, characterized in that The formula for generating the structured query language corresponding to the semantic representation data according to the structured key data is: In the formula, stands for Structured Query Language, represents the semantic representation data, represents the structured key data, Indicates the parameters of the generated structured query language function. Indicates the probability of generating the structured query language when the semantic representation data, the structured key data and the parameters of the function for generating the structured query language are input; It represents the structured query language with the highest selection probability.

6. The method according to claim 1, characterized in that The determining whether the to-be-parsed text data exists in the cache comprises: Performing hash calculation on the text data to be parsed to obtain a hash identifier of the text data to be parsed; Determine whether the hash identifier exists in the cache; If the hash identifier does not exist in the cache, it is determined that the text data to be parsed does not exist in the cache.

7. The method according to claim 6, characterized in that After determining whether the hash identifier exists in the cache, the method further includes: If the hash identifier exists in the cache, determining that the text data to be parsed exists in the cache; The structured query language corresponding to the hash identifier is obtained from the cache.

8. The method according to claim 1, characterized in that After obtaining the query result corresponding to the data query request, the method further includes: Calculating the query complexity of the structured query language; The structured query language is optimized according to the query complexity to obtain an optimized structured query language.

9. The method according to claim 8, characterized in that The formula for calculating the query complexity of the structured query language is: In the formula, represents the query complexity of the structured query language, represents the number of tables involved in the structured query language query, Indicates the weight of the table; represents the number of fields involved in the structured query language query, Indicates the weight of the field; represents the number of conditions in the structured query language query, Indicates the weight of the condition.

10. The method according to any one of claims 1 to 9, characterized in that: The formula for using the structured query language to search in a pre-built database to obtain the query result corresponding to the data query request is: In the formula, Indicates the query result. Represents a pre-built database, represents the structured query language, Represents a database execution function.

11. The method according to any one of claims 1 to 9, characterized in that: After sending the query result to the user terminal according to the data result display method so that the user terminal displays the query result, the method further includes: Receiving query result feedback information sent by the user terminal; The large model and / or the pre-built database are optimized according to the result feedback information.

12. A server, characterized in that: include: Memory for storing computer programs; A processor, used to implement the steps of the large model-based structured data processing method as described in any one of claims 1 to 11 when executing the computer program.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the large model-based structured data processing method as described in any one of claims 1 to 11.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the structured data processing method based on a large model as described in any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Database query method and device based on natural language and electronic equipment

    CN118535679A