A Natural Language Enhanced Data Analysis Method and System for Big Data

By classifying and matching analysis of user history search information and natural language text after speech conversion, the preferred natural language text is selected and their selection method is optimized, the problem of mismatch between natural language text and user needs in the prior art is solved, and the efficiency and accuracy of information retrieval is improved.

CN119396891BActive Publication Date: 2025-05-27JIANGSU ZHENYUN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411482722.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-05-27
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

In the prior art, when users enter information through voice input, they cannot effectively judge the matching between natural language text and user needs in advance, resulting in improper use of the text to be converted, and the retrieved information does not meet user needs, which reduces the efficiency of information retrieval.

Method used

By collecting user's historical search information and speech-transformed natural language text, performing classification preprocessing and matching analysis, selecting preferred natural language texts and generating preferred databases, optimizing the selection method of text to be converted, and using the T5 model to be converted into SQL language for information retrieval.

Benefits of technology

It effectively reduces the probability that the retrieved information does not meet user needs due to improper use of the text to be converted, and improves the efficiency and accuracy of information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119396891B_ABST
    Figure CN119396891B_ABST
Patent Text Reader

Abstract

The present invention discloses a natural language enhanced data analysis method and system for big data, which relates to the technical field of natural language conversion management, including a historical query data collection module, a language text screening module, a query information collection module, an information retrieval planning module, and a retrieval planning optimization module; the historical query data collection module collects the user's historical retrieval information and historical natural language texts, and performs classification preprocessing on the historical natural language texts; the language text screening module screens and selects preferred natural language texts and generates a preferred database; the query information collection module recognizes the current user's speech and converts it into a natural language text; the information retrieval planning module plans the selection method of the text to be converted, converts the text to be converted into SQL and performs information retrieval; the retrieval planning optimization module optimizes the selection method of the text to be converted, reducing the probability that the information does not meet the user's needs due to improper use of the text to be converted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language conversion management, and specifically to a natural language enhanced data analysis method and system for big data. Background Art

[0002] Data augmentation in natural language processing aims to better meet specific requirements by performing various transformations on text data. The purpose of converting natural language into structured query language SQL is to improve the efficiency and convenience of database queries, and further help users query the required information using natural language;

[0003] However, when a user queries the required information by voice input, first, the voice input by the user needs to be converted into natural language text form through voice recognition, and then SQL conversion is performed. Since the natural language text obtained by voice recognition conversion has diversity, there may be multiple natural language texts that can express the same meaning as the voice input by the user. However, for a specific SQL language conversion model, there may be a problem that the information retrieved after converting some language texts in the text expressing the same meaning into SQL does not meet the user's needs. At this time, the user needs to input the voice for retrieval again, which reduces the information retrieval efficiency. The prior art fails to pre-judge the matching situation between the natural language text and the user's needs to perform replacement processing on the natural language to be converted in advance, that is, perform natural language augmentation processing, but directly converts the recognized natural language text into SQL and performs information retrieval, and cannot effectively reduce the probability that the retrieved information does not meet the user's needs due to improper use of the text to be converted. Summary of the Invention

[0004] The purpose of the present invention is to provide a natural language enhanced data analysis method and system for big data to solve the problems raised in the prior art.

[0005] To achieve the above object, the present invention provides the following technical solution: A natural language enhanced data analysis system for big data, comprising a historical query data collection module, a language text screening module, a query information collection module, an information retrieval planning module, and a retrieval planning optimization module; the historical query data collection module collects the historical retrieval information of the user and the historical natural language text generated after converting the voice input by the user who needs to perform information retrieval, and performs classification preprocessing on the historical natural language text; the language text screening module analyzes the matching situation between the preprocessed historical natural language text and the user's needs, screens and selects the preferred natural language text and generates a preferred database; the query information collection module collects the voice information of the user who currently needs to perform information retrieval, and converts the voice information into the current natural language text; the information retrieval planning module plans the selection method of the text to be converted, converts the selected text to be converted into SQL language, and uses the SQL language to perform information retrieval; the retrieval planning optimization module establishes an optimization model for the selection of the text to be converted to optimize the selection method of the text to be converted.

[0006] Further, the historical query data collection module includes a natural language text collection unit, a historical text classification processing unit, and a retrieval information collection unit; the natural language text collection unit collects the historical natural language text generated after recognizing the voice input by the user through voice recognition technology after each previous voice input by the user who needs to retrieve information; the historical text classification processing unit classifies and processes the historical natural language text according to semantic similarity; the retrieval information collection unit collects the number of times the user re-enters the voice and the interval time information of the re-entered voice after transmitting the information retrieved by using the SQL language to the user terminal after converting the historical natural language text into the SQL language.

[0007] Further, the language text screening module includes a demand matching degree analysis unit, a preferred text screening unit, and a preferred database generation unit; the demand matching degree analysis unit analyzes the matching coefficient between the historical natural language text classified into the same category and the user's needs; the preferred text screening unit compares the matching coefficients between the historical natural language text classified into the same category and the user's needs, and screens out the preferred natural language text from the historical natural language text; the preferred database generation unit screens out the preferred natural language text in each category and generates a preferred database, and the preferred database includes the preferred natural language text in each category screened out; the query information collection module includes a voice receiving and recognizing unit and a natural language text generating unit, the voice receiving and recognizing unit receives the voice and performs voice recognition on the voice input by the user after the user currently inputs the voice; the natural language text generating unit generates the current natural language text after completing the voice recognition.

[0008] Further, the information retrieval planning module includes a library text matching unit, a text to be transformed selection unit, an SQL transformation unit, and an information retrieval unit; the current natural language text is matched with the preferred natural language text in the preferred database through the library text matching unit: the semantic similarity between the current natural language text and the preferred natural language text is analyzed, and the semantic similarity and the semantic similarity threshold are compared; through the text to be transformed selection unit, if there is no preferred natural language text in the preferred database whose semantic similarity with the current natural language text is higher than the threshold, the current natural language text is selected as the text to be transformed; if there is a preferred natural language text in the preferred database whose semantic similarity with the current natural language text is higher than the threshold, the preferred natural language text with the highest semantic similarity with the current natural language text is selected as the text to be transformed; through the SQL transformation unit, the text to be transformed is transformed into SQL language by using the T5 model; through the information retrieval unit, information retrieval is performed using the SQL language and the retrieved information is sent to the user.

[0009] Further, the retrieval planning optimization module includes a selection model establishment unit and a selection method optimization unit; after performing a number of text to be transformed selection operations through the selection model establishment unit, the preferred natural language texts selected as the text to be transformed are screened out, the semantic similarity between the screened-out texts and the natural language texts generated after recognizing the user input voice corresponding to each time is collected, the deviation value between the semantic similarity and the semantic similarity threshold is calculated, the matching coefficient between the screened-out texts and the user requirements of the input voice corresponding to each time is analyzed, and the deviation value and the matching coefficient are subjected to fitting training to establish an optimized model for selecting the text to be transformed; through the selection method optimization unit, the selection method of the text to be transformed is optimized: the voice input by the user is recognized to generate a natural language text, if there is no preferred natural language text in the preferred database whose semantic similarity with the corresponding generated natural language text is higher than the threshold, the corresponding generated natural language text is selected as the text to be transformed; if there is a preferred natural language text in the preferred database whose semantic similarity with the corresponding generated natural language text is higher than the threshold, the deviation value between the semantic similarity and the threshold is calculated, the deviation value is input into the optimized model for selecting the text to be transformed, the matching coefficient between the preferred natural language text whose semantic similarity with the corresponding generated natural language text is higher than the threshold and the corresponding user is output, and the preferred natural language text with the highest matching coefficient is selected as the text to be transformed.

[0010] A natural language enhanced data analysis method for big data, characterized by comprising the following steps:

[0011] Z100: Collect the historical retrieval information of users and the historical natural language text generated after converting the voice input by users who need to perform information retrieval, and perform classification preprocessing on the historical natural language text;

[0012] Z101: Analyze the matching situation between the preprocessed historical natural language text and the user's needs, screen out the preferred natural language text and generate a preferred database;

[0013] Z102: Collect the voice information of users who currently need to perform information retrieval, and convert the voice information into the current natural language text;

[0014] Z103: Plan the selection method of the text to be converted, convert the selected text to be converted into SQL language, and use the SQL language to perform information retrieval;

[0015] Z104: Establish an optimization model for the selection of the text to be converted, and optimize the selection method of the text to be converted.

[0016] Further, step Z100 includes: collecting the historical natural language text generated after recognizing the voice input by users through voice recognition technology each time in the past when information needs to be retrieved, analyzing the semantic similarity between historical natural language texts, setting the semantic similarity threshold as S, randomly extracting a text from the historical natural language texts, analyzing the semantic similarity between the remaining texts and the extracted text one by one, classifying the texts with a semantic similarity higher than S with the extracted text into the same category, and after the first classification is completed, performing classification preprocessing on the remaining unclassified texts in the same way until there are no unclassified texts, converting the historical natural language text into SQL language, using the SQL language to perform information retrieval and transmitting the retrieved information to users, and collecting the number of times users re-enter voice and the interval time between re-entering voice in the past.

[0017] Further, step Z101 includes: counting that there are n historical natural language texts classified into a random category, retrieving the set of the number of times users re-enter voice L = {L1, L2,..., Ln} after receiving the information retrieved by converting the corresponding n historical natural language texts into SQL language, retrieving the number of times Li that users re-enter voice after receiving the information retrieved by converting the i-th historical natural language text into SQL language, obtaining the set of the interval time between re-entering voice T = {T 1 , T 2 ,..., T Li}, and calculating the matching coefficient Ri between the i-th historical natural language text and the user's needs according to the following formula:

[0018] Ri = (1 / Li) + [(∑Li j=1 (T j )) / (Li);

[0019] j represents the j-th time the user re-enters the voice after receiving the information retrieved after converting the i-th historical natural language text into SQL language. The matching coefficients between n historical natural language texts and the user's requirements are calculated in the same way. The n matching coefficients are compared, and the n historical natural language texts are arranged in descending order of the matching coefficients. The first m historical natural language texts after the arrangement are selected as the preferred natural language texts, where 1 < m < n. The preferred natural language texts in each category are selected to generate a preferred database.

[0020] Using big data technology to collect historical information query data, analyze the number of times and time interval information of the user re-entering the voice for retrieval after receiving the information retrieved after converting the input voice into a natural language text and then converting the natural language text into SQL language. The more times of re-retrieval, the less the retrieved information meets the user's requirements. The purpose of analyzing the matching coefficient in combination with the time interval information is to consider that the user may have another retrieval requirement when entering the voice after a long interval, and the previous retrieval requirement may have been met. In this case, if the matching coefficient is analyzed only based on the number of times, the obtained matching coefficient has an error. Combining the number of times and the interval time to analyze the matching coefficient can effectively avoid this error. The preferred natural language texts are selected according to the matching coefficient and a preferred database is generated. Considering the problem of insufficient data in the previous preferred database, there is more than one preferred text in each category selected here.

[0021] Further, step Z103 includes: after obtaining the current natural language text in step Z102, analyzing the semantic similarity between the current natural language text and the preferred natural language texts, and comparing the semantic similarity with S: if there is no preferred natural language text in the preferred database with a semantic similarity higher than S to the current natural language text, select the current natural language text as the text to be converted; if there is a preferred natural language text in the preferred database with a semantic similarity higher than S to the current natural language text, select the preferred natural language text with the highest semantic similarity to the current natural language text as the text to be converted, use the T5 model to convert the text to be converted into SQL language, perform information retrieval using the SQL language and send the retrieved information to the user.

[0022] After generating the preferred database, perform a preliminary selection plan for the text to be transformed: select the text to be transformed by confirming whether there is a preferred natural language text in the preferred database that matches the current text. If there is a matching preferred text in the preferred database, select to replace the current text with the most matching preferred text as the text to be transformed. After replacement, perform SQL transformation and retrieve information. Timely replacement when necessary can effectively reduce the probability that the retrieved information does not meet the user's needs due to improper use of the text to be transformed.

[0023] Further, step Z104 includes: after performing k selection operations on the text to be transformed using the selection method planned in step Z103, a total of r preferred natural language texts are selected as the text to be transformed. The set of semantic similarity between the r texts and the natural language texts generated after recognizing the user's input voice for the corresponding times is {F1, F2,..., Fr}. The set of deviation values between the calculated semantic similarity and S is {F1 - S, F2 - S,..., Fr - S}. After counting that after converting the e-th selected text into SQL language for information retrieval and transmitting the retrieved information to the user, the number of times the user re-enters the voice is g, and the set of interval times for the g re-entries of the voice is {t 1 , t 2 ,..., t g}}, the matching coefficient We between the e-th selected text and the user's needs is obtained. We = (1 / g) + [(∑ g u=1 (t u )) / g], where u represents the u-th re-entry of the voice. The set of matching coefficients between the r selected texts and the user's needs is {W1, W2,...We,...Wr}, and the training data {(F1 - S, W1), (F2 - S, W2),...(Fr - S, Wr)} is generated. After linearly fitting the training data, an optimization model for text selection to be transformed is established:

[0024] Y = δ1*X + δ2;

[0025] Among them, δ1 and δ2 represent fitting coefficients, X represents the variable referring to the deviation value in the model, Y represents the variable referring to the matching coefficient in the model, and the optimization selection method is as follows: After performing speech recognition on the user input voice to generate natural language text, if there is no preferred natural language text in the preferred database whose semantic similarity to the corresponding generated natural language text is higher than S, select the corresponding generated natural language text as the text to be transformed; if there are H preferred natural language texts in the preferred database whose semantic similarity to the corresponding generated natural language text is higher than S, calculate the deviation value A between the semantic similarity between the a-th text among the H texts and the corresponding generated natural language text and S, let X = A, and output the matching coefficient between the a-th text and the user's requirement as δ1*A + δ2. Input the H deviation values into the model one by one, output H matching coefficients, and select the preferred natural language text with the highest matching coefficient among the H texts as the text to be transformed;

[0026] After performing multiple selections of the text to be transformed according to the preliminary selection planning method, considering the variability of users who need to retrieve information, the information retrieved by selecting the text to be transformed according to the preliminary planning method may also not match the changing user requirements. Further analyze the situation after selecting the text to be transformed according to the preliminary planning method to establish an optimization model for selecting the text to be transformed, and use the model as reference data for optimizing the selection of the text to be transformed, further optimizing the selection method of the text to be transformed, and further improving the matching degree between the retrieved information and the user.

[0027] Compared with the prior art, the beneficial effects of the present invention are:

[0028] The present invention collects historical information query data through big data technology, analyzes the number of times and time intervals of the user to re-enter the voice for retrieval after receiving the information retrieved after converting the input voice into natural language text and then converting the natural language text into SQL language, combines the number of times and the interval time to analyze the matching coefficient between the historical natural language text and the user's requirements, screens out the preferred natural language text and generates a preferred database. After generating the preferred database, conduct a preliminary selection planning for the text to be transformed, and timely perform replacement processing on the text to be transformed, effectively reducing the probability that the retrieved information does not meet the user's requirements due to improper use of the text to be transformed, and using the optimization model for selecting the text to be transformed to further optimize the selection method of the text to be transformed, and further improving the matching degree between the retrieved information and the user. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a schematic flowchart of a natural language enhanced data analysis method for big data according to the present invention;

[0030] Figure 2Schematic diagram of the structure of a natural language enhanced data analysis system for big data according to the present invention. Detailed implementation manners

[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0032] Embodiment 1: As Figure 1 shown, this embodiment provides a natural language enhanced data analysis method for big data, including the following steps:

[0033] Z100: Collect the historical retrieval information of the user and the historical natural language text generated after converting the voice input by the user who needs to perform information retrieval. Perform classification preprocessing on the historical natural language text: Collect the historical natural language text generated after identifying the voice input by the user through speech recognition technology every time the user needs to retrieve information in the past. Analyze the semantic similarity between the historical natural language texts. Set the semantic similarity threshold as S. Randomly extract a text from the historical natural language texts, and analyze the semantic similarity between the remaining texts and the extracted text one by one. Classify the texts with a semantic similarity higher than S to the same category as the extracted text. After the first classification is completed, classify and preprocess the remaining unclassified texts in the same way until there are no unclassified texts. Convert the historical natural language text into sql language, use the sql language to perform information retrieval, and transmit the retrieved information to the user. Collect the number of times the user re-enters the voice and the interval time between re-entering the voice after receiving the information in the past. The historical retrieval information refers to the number of times and interval time information;

[0034] Z101: Analyze the matching situation between the pre - processed historical natural language text and the user's needs, screen the preferred natural language text and generate a preferred database: It is statistically found that there are n historical natural language texts classified as random. After retrieving the information that the user retrieved in the past after converting the corresponding n historical natural language texts into SQL language, the set of the number of times of re - inputting voice is L = {L1, L2,..., Ln}. When the retrieved information is not the information the user needs, the user will re - input voice to retrieve information. When the re - retrieved information is still not the information the user needs, the user will re - input voice to retrieve information until the required information is obtained. Therefore, the number of times of re - inputting voice may be more than once. The number of times of re - inputting voice after retrieving the information that the user retrieved after converting the i - th historical natural language text into SQL language is Li, and the set of the time intervals between re - inputting voice is T = {T 1 ,T 2 ,...,T Li}, T 1 represents the time interval between the time when the user re - inputs voice for the first time and the time when the information is received, and T 2 represents the time interval between the time when the user re - inputs voice for the second time and the time when the user re - inputs voice for the first time. Calculate the matching coefficient Ri of the i - th historical natural language text and the user's needs according to the following formula:

[0035] Ri=(1 / Li)+[(∑ Li j=1 (T j )) / (Li)];

[0036] j represents the j - th time the user re - inputs voice after retrieving the information that the user retrieved after converting the i - th historical natural language text into SQL language. Calculate the matching coefficients of the n historical natural language texts and the user's needs in the same way, compare the n matching coefficients, arrange the n historical natural language texts in descending order of the matching coefficients, screen out the first m historical natural language texts after arrangement as the preferred natural language text, 1 < m < n, screen out the preferred natural language texts in each category, and generate a preferred database;

[0037] Z102: Collect the voice information of the user who needs to perform information retrieval currently, and convert the voice information into the current natural language text;

[0038] Z103: Plan the selection method of the text to be transformed, transform the selected text to be transformed into SQL language, and use the SQL language for information retrieval: After obtaining the current natural language text in step Z102, analyze the semantic similarity between the current natural language text and the preferred natural language text, and compare the semantic similarity with S. If there is no preferred natural language text in the preferred database whose semantic similarity with the current natural language text is higher than S, select the current natural language text as the text to be transformed. If there is a preferred natural language text in the preferred database whose semantic similarity with the current natural language text is higher than S, select the preferred natural language text with the highest semantic similarity with the current natural language text as the text to be transformed, transform the text to be transformed into SQL language using the T5 model, use the SQL language for information retrieval, and send the retrieved information to the user. The T5 model is a pre-trained language model based on the Transformer architecture, aiming to explore transfer learning of natural language processing tasks through a unified text-to-text conversion framework. The T5 model can achieve the transformation of natural language to structured query language SQL;

[0039] Z104: Establish an optimization model for the selection of the text to be transformed to optimize the selection method of the text to be transformed: After performing k times of text selection operations for the text to be transformed using the selection method planned in step Z103, a total of r preferred natural language texts selected as the text to be transformed are screened out. The set of semantic similarities between the r texts and the natural language texts generated after recognizing the user's input voice for the corresponding times is {F1, F2,..., Fr}, and the set of deviation values between the calculated semantic similarity and S is {F1 - S, F2 - S,..., Fr - S}. After counting that after transforming the e-th text screened out into SQL language for information retrieval and transmitting the retrieved information to the user, the number of times the user re-enters the voice is g, and the set of interval times for the g times of re-entering the voice is {t 1 , t 2 ,... t g}, t 1 represents the interval time between the time when the user first re-enters the voice after receiving the information retrieved using the e-th text after transformation and the time when the information is received, and the matching coefficient We between the e-th text screened out and the user's needs is obtained. We = (1 / g) + [(∑ g u=1 (t u )) / g], u represents the u-th time of re-entering the voice, and the set of matching coefficients between the r texts screened out and the user's needs is {W1, W2,... We,... Wr}, and the training data {(F1 - S, W1), (F2 - S, W2),...(Fr - S, Wr)} is generated. After linearly fitting the training data, an optimization model for the selection of the text to be transformed is established:

[0040] Y = δ1 * X + δ2;

[0041] Wherein, * represents the multiplication sign, δ1 and δ2 represent fitting coefficients, X represents the variable referring to the deviation value in the model, Y represents the variable referring to the matching coefficient in the model. Optimization selection method: After performing speech recognition on the speech input by the user to generate natural language text, if there is no preferred natural language text in the preferred database whose semantic similarity to the corresponding generated natural language text is higher than S, select the corresponding generated natural language text as the text to be transformed; if there are H preferred natural language texts in the preferred database whose semantic similarity to the corresponding generated natural language text is higher than S, calculate that the deviation value A between the semantic similarity between the a-th text among the H texts and the corresponding generated natural language text and S, let X = A, and output the matching coefficient between the a-th text and the user's requirement as δ1 * A + δ2. Input the H deviation values into the model one by one, output H matching coefficients, and select the preferred natural language text with the highest matching coefficient among the H texts as the text to be transformed;

[0042] For example: Establish an optimization model for selecting the text to be transformed: Y = 66.7X - 5. After performing speech recognition on the speech input by the user to generate natural language text, it is statistically found that there are 3 preferred natural language texts in the preferred database whose semantic similarity to the corresponding generated natural language text is higher than the threshold. Calculate that the deviation values between the semantic similarity between the 3 texts and the corresponding generated natural language text and the semantic similarity threshold are 0.11, 0.26, and 0.22 respectively. Let X be equal to 0.11, 0.26, and 0.22 respectively, and output the matching coefficients between the 3 texts and the user's requirement as 2.34, 12.34, and 9.67 respectively. The highest matching coefficient is 12.34. Select the 2nd text as the text to be transformed, use the T5 model to transform the text to be transformed into SQL language, perform information retrieval using the SQL language, and send the retrieved information to the user.

[0043] Example 2: As Figure 2As shown in the figure, this embodiment provides a natural language enhanced data analysis system for big data, which is implemented based on the data analysis method in the embodiment, and specifically includes: a historical query data collection module, a language text screening module, a query information collection module, an information retrieval planning module, and a retrieval planning optimization module; the historical query data collection module collects the historical retrieval information of the user and the historical natural language text generated after converting the voice input by the user who needs to perform information retrieval, and performs classification preprocessing on the historical natural language text; the language text screening module analyzes the matching situation between the preprocessed historical natural language text and the user's needs, screens out the preferred natural language text and generates a preferred database; the query information collection module collects the voice information of the user who needs to perform information retrieval currently, and converts the voice information into the current natural language text; the information retrieval planning module plans the selection method of the text to be converted, converts the selected text to be converted into sql language, and performs information retrieval using the sql language; the retrieval planning optimization module establishes an optimization model for the selection of the text to be converted, and optimizes the selection method of the text to be converted.

[0044] The historical query data collection module includes a natural language text collection unit, a historical text classification processing unit, and a retrieval information collection unit; the natural language text collection unit collects the historical natural language text generated after identifying the voice input by the user through voice recognition technology after each time the user needs to retrieve information in the past; the historical text classification processing unit classifies the historical natural language text according to semantic similarity; the retrieval information collection unit collects the number of times the user re-enters the voice and the interval time information of the re-entered voice after transmitting the information retrieved by using the sql language to the user terminal after converting the historical natural language text into the sql language.

[0045] The language text screening module includes a demand matching degree analysis unit, a preferred text screening unit, and a preferred database generation unit; the demand matching degree analysis unit analyzes the matching coefficient between the historical natural language text classified into the same category and the user's needs; the preferred text screening unit compares the matching coefficient between the historical natural language text classified into the same category and the user's needs, and screens out the preferred natural language text from the historical natural language text; the preferred database generation unit screens out the preferred natural language text in each category and generates a preferred database, and the preferred database includes the preferred natural language text in each category screened out; the query information collection module includes a voice receiving and recognizing unit and a natural language text generating unit, the voice receiving and recognizing unit receives the voice and performs voice recognition on the voice input by the user after the user inputs the voice currently; the natural language text generating unit generates the current natural language text after completing the voice recognition.

[0046] The information retrieval planning module includes a library text matching unit, a text to be transformed selection unit, an SQL transformation unit, and an information retrieval unit; the library text matching unit matches the current natural language text with the preferred natural language text in the preferred database: analyzes the semantic similarity between the current natural language text and the preferred natural language text, and compares the semantic similarity with the semantic similarity threshold; through the text to be transformed selection unit, if there is no preferred natural language text in the preferred database with a semantic similarity higher than the threshold to the current natural language text, the current natural language text is selected as the text to be transformed; if there is a preferred natural language text in the preferred database with a semantic similarity higher than the threshold to the current natural language text, the preferred natural language text with the highest semantic similarity to the current natural language text is selected as the text to be transformed; through the SQL transformation unit, the text to be transformed is transformed into SQL language using the T5 model; through the information retrieval unit, information retrieval is performed using the SQL language and the retrieved information is sent to the user.

[0047] The retrieval planning optimization module includes a selection model establishment unit and a selection method optimization unit; through the selection model establishment unit, after performing a number of text to be transformed selection operations, the preferred natural language texts selected as the text to be transformed are screened out, the semantic similarity between the screened-out texts and the natural language texts generated after recognizing the user input voice corresponding to each time is collected, the deviation value between the semantic similarity and the semantic similarity threshold is calculated, the matching coefficient between the screened-out texts and the user requirements of the input voice corresponding to each time is analyzed, and the deviation value and the matching coefficient are subjected to fitting training to establish an optimization model for text to be transformed selection; through the selection method optimization unit, the selection method of the text to be transformed is optimized: the voice input by the user is recognized to generate a natural language text, if there is no preferred natural language text in the preferred database with a semantic similarity higher than the threshold to the corresponding generated natural language text, the corresponding generated natural language text is selected as the text to be transformed; if there is a preferred natural language text in the preferred database with a semantic similarity higher than the threshold to the corresponding generated natural language text, the deviation value between the semantic similarity and the threshold is calculated, the deviation value is input into the optimization model for text to be transformed selection, the matching coefficient between the preferred natural language text with a semantic similarity higher than the threshold to the corresponding generated natural language text and the corresponding user is output, and the preferred natural language text with the highest matching coefficient is selected as the text to be transformed.

[0048] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. A natural language enhanced data analysis system for big data, characterized by: It includes historical query data collection module, language text screening module, query information collection module, information retrieval planning module and retrieval planning optimization module; The historical query data collection module collects the historical search information of the user and the historical natural language text generated by converting the voice input by the user who needs to perform information retrieval, and classifies and pre-processes the historical natural language text; The language text screening module analyzes the matching between the pre-processed historical natural language text and the user's needs. The matching is obtained by analyzing the number of times the user has re-entered the voice and the interval time between the re-entered voices after receiving the information retrieved by converting the historical natural language text into the SQL language, and screening the preferred natural language text and generating the preferred database; The query information collection module collects voice information of the user who currently needs to perform information retrieval, and converts the voice information into current natural language text; The information retrieval planning module is used to plan the selection method of the text to be converted, the selected text to be converted is converted into SQL language, and the information retrieval is performed using SQL language; The information retrieval planning module includes a text matching unit in the library, a text selection unit to be converted, a SQL conversion unit and an information retrieval unit; Matching the current natural language text with the preferred natural language text in the preferred database by the text matching unit in the library: analyzing the semantic similarity between the current natural language text and the preferred natural language text, and comparing the semantic similarity with the semantic similarity threshold; If there is no preferred natural language text in the preferred database whose semantic similarity with the current natural language text is higher than the threshold value, the current natural language text is selected as the text to be converted by the text selection unit to be converted; if there is a preferred natural language text in the preferred database whose semantic similarity with the current natural language text is higher than the threshold value, the preferred natural language text with the highest semantic similarity with the current natural language text is selected as the text to be converted; The sql conversion unit converts the text to be converted into sql language using the T5 model; The information retrieval unit uses the SQL language to retrieve information and sends the retrieved information to the user; Establishing a text selection optimization model to be converted through the search planning optimization module to optimize the selection method of the text to be converted; The retrieval planning optimization module includes a selection model building unit and a selection mode optimization unit; After performing several operations of selecting texts to be converted, the selection model building unit screens out preferred natural language texts selected as texts to be converted, collects semantic similarity between the screened texts and natural language texts generated after the corresponding recognition of user input speech, calculates the deviation value between the semantic similarity and the semantic similarity threshold, analyzes the matching coefficient between the screened texts and the user requirements of the corresponding input speech, performs fitting training on the deviation value and the matching coefficient, and establishes a text selection optimization model to be converted; The selection method of the text to be converted is optimized by the selection method optimization unit: natural language text is generated after voice recognition is performed on the voice input by the user; if there is no preferred natural language text in the preferred database whose semantic similarity with the corresponding generated natural language text is higher than a threshold, the corresponding generated natural language text is selected as the text to be converted; if there is a preferred natural language text in the preferred database whose semantic similarity with the corresponding generated natural language text is higher than the threshold, the deviation value between the semantic similarity and the threshold is calculated, the deviation value is input into the text to be converted selection optimization model, the matching coefficient between the preferred natural language text whose semantic similarity with the corresponding generated natural language text is higher than the threshold and the corresponding user is output, and the preferred natural language text with the highest matching coefficient is selected as the text to be converted.

2. The natural language enhanced data analysis system for big data according to claim 1, characterized in that: The historical query data collection module includes a natural language text collection unit, a historical text classification processing unit and a search information collection unit; The natural language text collection unit collects the historical natural language text generated by recognizing the user input voice by the voice recognition technology after the user input voice is recognized each time the user needs to retrieve information in the past; The historical text classification processing unit classifies the historical natural language texts according to semantic similarity; The retrieval information collection unit collects the number of times the user re-enters the voice and the interval time information of the re-entering voice after the historical natural language text is converted into the SQL language and the information retrieved by the SQL language is transmitted to the user terminal.

3. The natural language enhanced data analysis system for big data according to claim 2, characterized in that: The language text screening module includes a demand matching degree analysis unit, a preferred text screening unit and a preferred database generation unit; The matching coefficients of historical natural language texts classified into the same category and user requirements are analyzed by the demand matching degree analysis unit; the matching coefficients of historical natural language texts classified into the same category and user requirements are compared by the preferred text screening unit, and preferred natural language texts are screened out from the historical natural language texts; the preferred natural language texts in each category are screened out by the preferred database generation unit, and a preferred database is generated, wherein the preferred database includes the screened preferred natural language texts in each category; The query information collection module includes a speech receiving and recognition unit and a natural language text generation unit. After the user currently inputs speech, the speech receiving and recognition unit receives the speech and performs speech recognition on the user's input speech; and the natural language text generation unit generates the current natural language text after completing the speech recognition.

4. A natural language enhanced data analysis method for big data, characterized by: The following steps are involved: Z100: collects historical search information of users and converts the voice input of users who need to search for information into historical natural language text, and performs classification preprocessing on the historical natural language text; Z101: Analyze the matching between the pre-processed historical natural language text and the user's needs. The matching is obtained based on the number of times the user has re-entered the voice after receiving the retrieved information after the historical natural language text is converted into SQL language, and the interval time between the re-entry of the voice, and select the preferred natural language text and generate the preferred database; Z102: Collect the voice information of the user who currently needs to retrieve information, and convert the voice information into the current natural language text; Z103: Plan the method of selecting the text to be converted, convert the selected text to be converted into SQL language, and use SQL language to perform information retrieval; Z104: Establish an optimization model for selecting texts to be converted and optimize the selection method of texts to be converted; Step Z103 includes: after executing step Z102 to obtain the current natural language text, analyzing the semantic similarity between the current natural language text and the preferred natural language text, and comparing the semantic similarity with S: if there is no preferred natural language text in the preferred database whose semantic similarity with the current natural language text is higher than S, selecting the current natural language text as the text to be converted; if there is a preferred natural language text in the preferred database whose semantic similarity with the current natural language text is higher than S, selecting the preferred natural language text with the highest semantic similarity with the current natural language text as the text to be converted, converting the text to be converted into SQL language using the T5 model, performing information retrieval using the SQL language and sending the retrieved information to the user, S is the set semantic similarity threshold; Step Z104 includes: after performing k times of text selection operations to be converted using the selection method planned in step Z103, a total of r preferred natural language texts selected as texts to be converted are screened out, the semantic similarity set between the r texts and the natural language texts generated after the corresponding recognition of the user input voice is collected as {F1, F2, ..., Fr}, the deviation value set between the semantic similarity and S is calculated as {F1-S, F2-S, ..., Fr-S}, and it is counted that after the e-th text selected is converted into SQL language for information retrieval and the retrieved information is transmitted to the user, the number of times the user re-enters the voice is g, and the interval time set of the g times of re-entering the voice is {t1, t2, ..., t g }, and obtain the matching coefficient We between the filtered e-th text and the user's needs, We=(1 / g)+[(∑ g u=1 (t u )) / g], u represents the u-th re-input of speech, and the set of matching coefficients between the selected r texts and user needs is {W1, W2, ...We, ...Wr}, generating training data {(F1-S, W1), (F2-S, W2), ...(Fr-S, Wr)}, and establishing the optimization model for selecting texts to be converted after linear fitting of the training data: Y = δ1*X+δ2; Among them, δ1 and δ2 represent fitting coefficients, X represents the variable referring to the deviation value in the model, and Y represents the variable referring to the matching coefficient in the model. The optimization selection method is: after performing speech recognition on the user's input voice, a natural language text is generated. If there is no preferred natural language text with a semantic similarity higher than S with the corresponding generated natural language text in the preferred database, the corresponding generated natural language text is selected as the text to be converted; if there are H preferred natural language texts with a semantic similarity higher than S with the corresponding generated natural language text in the preferred database, the deviation value between the semantic similarity between the a-th text in the H texts and the corresponding generated natural language text and S is calculated to be A, let X=A, and the matching coefficient between the a-th text and the user's demand is output as δ1*A+δ2, the H deviation values ​​are input into the model one by one, and H matching coefficients are output. The preferred natural language text with the highest matching coefficient is selected from the H texts as the text to be converted.

5. The natural language enhanced data analysis method for big data according to claim 4, characterized in that: Step Z100 includes: collecting historical natural language texts generated after recognizing the user's input speech through speech recognition technology every time information needs to be retrieved in the past, analyzing the semantic similarity between the historical natural language texts, setting the semantic similarity threshold as S, randomly extracting a text from the historical natural language texts, analyzing the semantic similarity between the remaining texts and the extracted text one by one, classifying the texts with a semantic similarity higher than S with the extracted text into the same category, after the first classification is completed, preprocessing the remaining unclassified texts in the same way until there are no unclassified texts, converting the historical natural language texts into SQL language, using the SQL language to retrieve information and transmitting the retrieved information to the user, collecting the number of times the user re-entered speech and the interval time between re-entering speech after receiving the information in the past.

6. The natural language enhanced data analysis method for big data according to claim 5, characterized in that: Step Z101 includes: counting that there are n historical natural language texts classified as random, retrieving the number of times the user re-entered the voice after receiving the information retrieved by converting the corresponding n historical natural language texts into SQL language as L = {L1, L2, ..., Ln}, retrieving the number of times the user re-entered the voice after receiving the information retrieved by converting the i-th historical natural language text into SQL language as Li, and obtaining the interval time set of re-entering the voice as T = {T1, T2, ..., T Li }, T1 represents the interval between the time when the user re-enters the voice for the first time and the time when the information is received. The matching coefficient Ri between the i-th historical natural language text and the user's needs is calculated according to the following formula: Ri=(1 / Li)+[(∑ Li j=1 (T j )) / (Li)]; j represents the jth time the user re-enters speech after receiving the information retrieved by converting the ith historical natural language text into SQL language. Calculate the matching coefficients between n historical natural language texts and the user's needs in the same way, compare the n matching coefficients, arrange the n historical natural language texts in descending order of the matching coefficients, screen out the first m historical natural language texts after the arrangement as the preferred natural language texts, 1 < m < n, screen out the preferred natural language texts in each category, and generate a preferred database.

Citation Information

Patent Citations

  • Voice interaction method, vehicle, server, system and storage medium

    CN115457961A

  • Retrieval method and device based on user requirements and electronic equipment

    CN116244344A

  • Log query statement generation method and device, equipment and storage medium

    CN118152341A