Data interaction display method and system based on voice interaction and medium

Through the data interaction display method based on voice interaction, the time-consuming and labor-intensive and unintuitive problems of traditional data retrieval and analysis methods are solved, efficient and accurate data retrieval and intuitive display are achieved, and employees' work efficiency and decision-making ability are improved.

CN120164456APending Publication Date: 2025-06-17CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510304546.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Traditional data retrieval and analysis methods rely on manual input, require high SQL database operation skills, are prone to errors, and are time-consuming and labor-intensive in the face of massive data, and data display is not intuitive, which affects decision-making efficiency.

Method used

The data interaction display method based on voice interaction is adopted to realize intuitive display and interaction of data by obtaining user voice data, preprocessing, feature extraction, voice recognition and intelligent data assistant retrieval.

Benefits of technology

It improves the efficiency and accuracy of data retrieval, reduces manual processing time, realizes intuitive display and interaction of data, and improves employees' work efficiency and decision-making capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164456A_ABST
    Figure CN120164456A_ABST
Patent Text Reader

Abstract

The invention discloses a data interaction display method and system based on voice interaction and a medium, and the method comprises the steps: obtaining user voice data, carrying out the preprocessing of the user voice data, and obtaining a first voice signal; performing feature extraction on the first voice signal to obtain Mel filter bank features, and inputting the Mel filter bank features into a pre-trained voice recognition model to obtain a voice recognition result; performing retrieval and case search on the voice recognition result through a preset intelligent data assistant to obtain an access result required by the user; displaying the access result required by the user according to the numerical value type corresponding to the access result required by the user; wherein the intelligent data assistant comprises a retrieval enhancement generation framework and a large language model. The method improves the efficiency and accuracy of data retrieval, achieves the visual display and interaction of data, can greatly reduce the manual processing time, effectively improves the working efficiency of employees, and can be widely applied to the technical field of data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a data interaction display method, system and medium based on voice interaction. Background Art

[0002] In a rapidly changing business environment, enterprise employees (such as marketing service managers) need to grasp business data in real time in order to make rapid and accurate decisions and achieve better benefits for the enterprise. However, traditional data retrieval and analysis methods usually rely on manual input, require users to have relatively high SQL database operation skills, and are very error-prone. Even for experienced users, facing a large amount of business data, manual input and screening are still a time-consuming and laborious task. Traditional data display methods, such as spreadsheets or static reports, often have problems such as untimely data updates and unintuitive information display, which limit the decision-making efficiency of employees. Summary of the Invention

[0003] To solve the above technical problems, the purpose of the present invention is to provide a data interaction display method, system and medium based on voice interaction, which can improve the efficiency and accuracy of data retrieval and achieve intuitive display and interaction of data.

[0004] To achieve the above purpose, one aspect of the embodiments of the present application proposes a data interaction display method based on voice interaction, including the following steps:

[0005] Obtain user voice data, preprocess the user voice data to obtain a first voice signal;

[0006] Extract features from the first voice signal to obtain Mel filter bank features, and input the Mel filter bank features into a pre-trained speech recognition model to obtain a speech recognition result;

[0007] Retrieve and search for cases on the speech recognition result through a preset intelligent data assistant to obtain the data extraction result required by the user;

[0008] Display the data extraction result required by the user according to the numerical type corresponding to the data extraction result required by the user;

[0009] Wherein, the intelligent data assistant includes a retrieval enhanced generation framework and a large language model.

[0010] In some embodiments, the preprocessing of the user voice data to obtain a first voice signal specifically includes:

[0011] Normalize the user voice data to obtain a second voice signal;

[0012] Perform pre-emphasis processing on the second voice signal to obtain the first voice signal.

[0013] In some embodiments, the feature extraction of the first voice signal to obtain Mel filter bank features specifically includes:

[0014] Perform frame addition and windowing on the first voice signal in the time domain to obtain a time-domain signal;

[0015] Perform short-time Fourier transform on the time-domain signal to obtain a short-time Fourier transform magnitude spectrum;

[0016] Set up a Mel filter bank, input the short-time Fourier transform magnitude spectrum into the Mel filter bank to obtain a Mel short-time Fourier transform magnitude spectrum;

[0017] Take the logarithm modulus of the Mel short-time Fourier transform magnitude spectrum to obtain a logarithmic Mel time-frequency spectrogram;

[0018] Perform feature extraction on the logarithmic Mel time-frequency spectrogram to obtain the Mel filter bank features.

[0019] In some embodiments, the data interaction display method further includes the step of pre-training the speech recognition model. The pre-training of the speech recognition model specifically includes:

[0020] Obtain a public speech dataset and a local speech dataset;

[0021] Perform feature extraction on the public speech dataset to obtain a first Mel filter bank feature set;

[0022] Perform feature extraction on the local speech dataset to obtain a second Mel filter bank feature set;

[0023] Perform parameter migration on the pre-trained Paraformer model, and input the first Mel filter bank feature set and the second Mel filter bank feature set into the Paraformer model after parameter migration;

[0024] Perform transfer learning on the Paraformer model after parameter migration training through cross-entropy loss, mean absolute error loss, and minimum word error rate loss to obtain the trained speech recognition model;

[0025] Set a confidence threshold. When the confidence of the speech recognition result is higher than the confidence threshold, input the corresponding user speech data into the local speech dataset.

[0026] In some embodiments, the retrieval and case search of the speech recognition result through a preset intelligent data assistant to obtain the data extraction result required by the user specifically includes:

[0027] Construct the retrieval-augmented generation framework, and retrieve the speech recognition result through the retrieval-augmented generation framework to obtain multiple similar cases;

[0028] Construct a question text, and input the question text and each of the similar cases into the large language model for case search to obtain the most similar case;

[0029] Construct a user data extraction case library, and extract data from the user data extraction case library according to the SQL code of the most similar case to obtain the data extraction result required by the user.

[0030] In some embodiments, the retrieval-augmented generation framework includes a bidirectional encoding representation model and a text retrieval algorithm. The retrieving the speech recognition result through the preset retrieval-augmented generation framework to obtain multiple similar cases specifically includes:

[0031] Construct a data query case library;

[0032] Input the speech recognition result into the bidirectional encoding representation model for vector encoding to obtain a speech problem vector and a vector encoding corresponding to the speech problem vector;

[0033] Through the text retrieval algorithm, perform retrieval according to the data query case library, the speech problem vector, and the vector encoding to obtain multiple similar cases.

[0034] In some embodiments, the presenting the data extraction result required by the user according to the numerical type corresponding to the data extraction result required by the user specifically includes:

[0035] Construct a data presentation template, and determine the numerical type corresponding to the data extraction result required by the user;

[0036] Present the data extraction result required by the user according to the data presentation template and the numerical type.

[0037] To achieve the above object, another aspect of the embodiments of the present application proposes a data interaction display system based on voice interaction, including:

[0038] A data preprocessing module, configured to obtain user voice data and preprocess the user voice data to obtain a first voice signal;

[0039] A voice enhancement and recognition module, configured to extract features from the first voice signal to obtain Mel filter bank features, and input the Mel filter bank features into a pre-trained speech recognition model to obtain a speech recognition result;

[0040] The data extraction result determination module is used to retrieve and perform case search on the speech recognition result through a preset intelligent data assistant to obtain the data extraction result required by the user;

[0041] The data interaction display module is used to display the data extraction result required by the user according to the numerical type corresponding to the data extraction result required by the user;

[0042] Wherein, the intelligent data assistant includes a retrieval enhanced generation framework and a large language model.

[0043] To achieve the above object, on the other hand, an embodiment of the present application proposes an electronic device, which includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing the connection and communication between the processor and the memory. When the program is executed by the processor, it realizes the data interaction display method based on voice interaction as described above.

[0044] To achieve the above object, on the other hand, an embodiment of the present application proposes a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to realize the data interaction display method based on voice interaction as described above.

[0045] The beneficial effects of the present invention are as follows: For the data interaction display method, system and medium based on voice interaction of the present invention, first, the user voice data is preprocessed to obtain a first voice signal, then the first voice signal is subjected to feature extraction to obtain Mel filter bank features, and the Mel filter bank features are input into a pre-trained speech recognition model to obtain a speech recognition result. Furthermore, an intelligent data assistant including a retrieval enhanced generation framework and a large language model is used to retrieve and perform case search on the speech recognition result to obtain the data extraction result required by the user. Finally, the data extraction result required by the user is displayed according to the numerical type corresponding to the data extraction result required by the user. The present invention constructs an intelligent data assistant based on a retrieval enhanced generation framework and a large language model, and through the voice interaction function and a pre-trained language recognition model, effectively identifies user voice commands to improve data access efficiency, and finally realizes the intuitive understanding and rapid analysis of data, improves the efficiency and accuracy of data retrieval, displays data according to the numerical type corresponding to the data extraction result required by the user, realizes the intuitive display and interaction of data, can greatly reduce the manual processing time, and effectively improve the work efficiency of employees. Description of the Drawings

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduces the accompanying drawings required to be used in the embodiments of the present invention. It should be understood that the accompanying drawings in the following introduction only facilitate the clear expression of some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0047] Figure 1 It is a flowchart of steps of a method for data interaction display based on voice interaction provided by an embodiment of the present invention;

[0048] Figure 2 It is a flowchart of steps of data preprocessing provided by an embodiment of the present invention;

[0049] Figure 3 It is a flowchart of steps of voice feature extraction provided by an embodiment of the present invention;

[0050] Figure 4 It is a schematic flowchart of the training of a voice recognition model provided by an embodiment of the present invention;

[0051] Figure 5 It is a schematic flowchart of the processing of an intelligent data assistant provided by an embodiment of the present invention;

[0052] Figure 6 It is a schematic flowchart of the processing of a retrieval enhanced generation framework provided by an embodiment of the present invention;

[0053] Figure 7 It is a schematic flowchart of the process of determining the most similar case based on a large language model provided by an embodiment of the present invention;

[0054] Figure 8 It is a schematic diagram of a user data extraction case library provided by an embodiment of the present invention;

[0055] Figure 9 It is an example diagram of the display of the data extraction results required by the user provided by an embodiment of the present invention;

[0056] Figure 10 It is a schematic structural diagram of a data interaction display system based on voice interaction provided by an embodiment of the present invention;

[0057] Figure 11 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0058] To make the objectives, technical solutions, and advantages of this application more clearly understood, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this application. They are merely examples of devices and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0059] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".

[0060] The terms "at least one", "a plurality of", "each", "any one", etc. used in this application, "at least one" includes one, two, or more than two, "a plurality of" includes two or more than two, "each" refers to each of the corresponding plurality, and "any one" refers to any one of the plurality.

[0061] Before elaborating in detail on the embodiments of this application, some nouns and terms involved in the embodiments of this application are first explained. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.

[0062] Fbank (Mel Filter Bank, Mel filter bank coefficients): Features extracted from a speech signal by applying a set of triangular filters over different frequency bandwidths. These filters are typically distributed according to the Mel scale, which is a frequency scale based on human auditory perception and simulates the sensitivity of the human ear to different frequencies.

[0063] STFT (Short-Time Fourier Transform): A mathematical tool for analyzing non-stationary signals (i.e., signals whose characteristics change over time). It is an extension of the Fourier transform and is particularly suitable for processing audio signals or any signal that changes over time. STFT analyzes the frequency content of a signal within a local time range by dividing the signal into shorter segments (usually called frames) and applying the Fourier transform to each frame.

[0064] RAG (Retrieval-Augmented Generation): A natural language processing (NLP) method that combines information retrieval and text generation techniques. It aims to enhance the capabilities of a generation model by introducing an external knowledge base, enabling the model to utilize a wider range of information when generating text.

[0065] CE (Cross-Entropy Loss): The cross-entropy loss function is a metric used to measure the difference between predicted results and true results. Based on the concepts of information theory, it measures the performance of a model by measuring the difference between probability distributions. Specifically, the CE loss function calculates the cross-entropy value between the probability distribution of the predicted results and the probability distribution of the true results. When the two probability distributions are closer, the cross-entropy value is smaller, indicating that the model's prediction results are more accurate.

[0066] MAE (Mean Absolute Error): A loss function for regression models and also a metric for evaluating the quality of a model's prediction results. It represents the average of the absolute values of the differences between the predicted values and the true values.

[0067] WEER (Minimum Word Error Rate): A key performance metric in automatic speech recognition and other text generation tasks. It represents the minimum word-level error rate between the text recognized by the system and the reference text (i.e., the true or expected text). This error rate is typically obtained by calculating the total number of words with insertion errors, deletion errors, and substitution errors, and then dividing it by the total number of words in the reference text.

[0068] In a rapidly changing business environment, enterprise employees (such as marketing service managers) need to have real-time access to business data in order to make quick and accurate decisions and achieve better benefits for the enterprise. However, traditional data retrieval and analysis methods usually rely on manual input, require users to have a high level of SQL database operation skills, and are very error-prone. Even for experienced users, manually inputting and filtering through massive amounts of business data is still a time-consuming and laborious task. Traditional data display methods, such as spreadsheets or static reports, often suffer from problems such as untimely data updates and unintuitive information display, which limit the decision-making efficiency of employees. In addition, traditional data query interfaces are often complex in design and have numerous functions, making it difficult for non-professionals to get started and prone to data misinterpretation or omission due to improper operation, thus affecting subsequent data analysis and decision-making.

[0069] Meanwhile, traditional data display and analysis tools also have obvious deficiencies in terms of interactivity. These tools often simply present data, lacking interaction and feedback mechanisms with users, which makes it impossible for users to flexibly adjust the data display method and analysis dimension according to their own needs. This one-way data transmission method not only limits users' in-depth understanding and exploration of data, but also reduces users' usage experience and satisfaction.

[0070] The data query methods in the prior art mainly adopt the method of general SQL statements and the matching query method based on AI. The main differences between the two are the computing power requirements and query time. The former has lower computing power requirements but longer query time. When facing a large amount of business data, manual input and screening are extremely time-consuming and laborious tasks; the latter has higher computing power requirements, but can achieve fast and effective data search and return the data required by users for display.

[0071] Therefore, the embodiments of the present invention propose a data interaction display method based on voice interaction. First, the user voice data is preprocessed to obtain the first voice signal. Then, feature extraction is performed on the first voice signal to obtain the Mel filter bank features. The Mel filter bank features are input into a pre-trained speech recognition model to obtain the speech recognition result. Furthermore, the speech recognition result is retrieved and case searched through an intelligent data assistant including a retrieval-augmented generation framework and a large language model to obtain the data extraction result required by the user. Finally, the data extraction result required by the user is displayed according to the numerical type corresponding to the data extraction result required by the user. The present invention constructs an intelligent data assistant based on a retrieval-augmented generation framework and a large language model, and through the voice interaction function and a pre-trained language recognition model, effectively identifies user voice commands to improve data access efficiency, ultimately realizes the intuitive understanding and rapid analysis of data, improves the efficiency and accuracy of data retrieval, realizes the intuitive display and interaction of data based on the data extraction result required by the user, can greatly reduce the manual processing time, and effectively improve the work efficiency of employees.

[0072] Refer to Figure 1 , Figure 1 FIG. is a step flow chart of a data interaction display method based on voice interaction provided by an embodiment of the present invention. The embodiments of the present invention propose a data interaction display method based on voice interaction, and this method includes steps S101 to S104:

[0073] S101. Obtain user voice data, and preprocess the user voice data to obtain the first voice signal;

[0074] Specifically, the user voice data is processed through data normalization and data pre-emphasis to obtain the preprocessed voice data, that is, the noisy first voice signal, to prepare for the next step of voice enhancement and recognition.

[0075] Reference Figure 2 , Figure 2 FIG. Figure 2 is a flowchart of steps for data preprocessing provided by an embodiment of the present invention. Further, as an optional implementation, the step of preprocessing user voice data to obtain a first voice signal can be specifically divided into the following steps S1011 and S1012:

[0076] S1011. Normalize the user voice data to obtain a second voice signal;

[0077] In some optional embodiments, in order to eliminate the signal amplitude differences caused by recording with different devices of different users and improve the accuracy of the final model prediction and recognition results, it is necessary to perform data standardization on the sound signals collected by each independent device. Two optional data standardization methods are provided in the embodiments of the present invention:

[0078] 1) Amplitude normalization method:

[0079]

[0080] where x norm is the signal after amplitude normalization, and x input , x max are the collected sound signal and its maximum amplitude respectively.

[0081] 2) Data standardization method for compressing the amplitude into the interval [a, b]:

[0082]

[0083] where x norm is the signal after amplitude compression, and x input , x min , x max are the collected signal, the minimum value and the maximum value of the signal respectively.

[0084] S1012. Perform pre-emphasis processing on the second voice signal to obtain a first voice signal.

[0085] In some optional embodiments, the purpose of pre-emphasis is to enhance the high-frequency components in the voice signal. Since the high-frequency part of the recorded voice signal decays too fast, resulting in large signal energy in the low-frequency band and small energy in the high-frequency band of the voice. To compensate for the loss of the high-frequency part, a first-order high-pass digital filter is used to perform pre-emphasis on the second voice signal. Its calculation formula is as follows:

[0086] y(n) = x(n) - ax(n - 1);

[0087] H(z) = 1 - az -1 ;

[0088] Among them, x(n) is the speech signal after amplitude normalization, and a is the coefficient of pre-emphasis, generally taken between 0.9 and 1. In the embodiment of the present invention, a = 0.95 is taken.

[0089] S102. Extract features from the first speech signal to obtain Mel filter bank features, and input the Mel filter bank features into a pre-trained speech recognition model to obtain a speech recognition result;

[0090] Specifically, extract speech features from the preprocessed first speech signal to obtain Mel filter bank features. Then, input the Mel filter bank features into the speech recognition model trained by transfer learning for recognition. Finally, the model returns the speech recognition result and sends it to the next intelligent data assistant.

[0091] Refer to Figure 3 , Figure 3 , which is a step flowchart of speech feature extraction provided by the embodiment of the present invention. Further as an optional implementation manner, the step of extracting features from the first speech signal to obtain Mel filter bank features can be specifically divided into the following steps S1021 to S1025:

[0092] S1021. Frame and window the first speech signal in the time domain to obtain a time-domain signal;

[0093] It should be noted that when a clean speech signal is input into the speech recognition module, feature extraction needs to be performed first. In the embodiment of the present invention, the Fbank feature extraction method (Filter Bank, Fbank) is used. Since the human ear's response to the sound spectrum is non-linear, Fbank is a front-end processing algorithm that processes audio in a manner similar to the human ear and can improve the performance of speech recognition.

[0094] Specifically, any speech signal is non-stationary as a whole, but can be regarded as a stationary signal locally. In subsequent audio processing, a stable signal needs to be input. The principle is that within a range of 10 - 30 milliseconds, the characteristic parameters of the audio signal change little, and the change can be ignored, and it is treated as a steady state. Therefore, it is necessary to frame and window the speech signal x(n), and its formula is as follows:

[0095] N = F s ×T;

[0096] M = F s ×T s ;

[0097] Among them, F s is the sampling frequency, T is the frame length time, N is the number of sampling points, T sis the frame shift time, M is the number of frame shifts. In the embodiments of the present invention, F is taken as s 16 kHz, T is 25 ms, and T s is 10 ms.

[0098] Common window functions include rectangular window, Hanning window, and Hamming window. The Hamming window is an improvement of the Hanning window. From the perspective of the frequency response expression, it is only a change in coefficients. Compared with the Hanning window, its side lobes are smaller, so that the energy is concentrated in the main lobe, and it can better overcome the problem of spectral leakage. Therefore, in the embodiments of the present invention, the Hamming window is used to window the first speech signal.

[0099] The time-domain definition and frequency response of the Hamming window are as follows:

[0100]

[0101] In the time domain, the formula for windowing the first speech signal to obtain the i-th frame is:

[0102] x i (n) = x(n + l)w Ham (n);

[0103] where l is the total frame shift of the i-th frame relative to the starting point, and x(n + l) is the signal of the i-th frame.

[0104] S1022. Perform short-time Fourier transform on the time-domain signal to obtain the short-time Fourier transform magnitude spectrum;

[0105] Specifically, perform short-time Fourier transform on the time-domain signal after frame division and windowing, and obtain the short-time Fourier transform magnitude spectrum. The formula is as follows:

[0106]

[0107] S(n, f) = |X(n, f)| 2 ;

[0108] where x(m) is the time-domain signal after frame division and windowing, w(n - m) is the window function with an offset of n sample points, f is the frequency, and S(n, f) is the short-time Fourier transform magnitude spectrum.

[0109] S1023. Set up a Mel filter bank, input the short-time Fourier transform magnitude spectrum into the Mel filter bank, and obtain the Mel short-time Fourier transform magnitude spectrum;

[0110] Specifically, design a Mel filter bank, pass the short-time Fourier transform magnitude spectrum through the Mel filter bank to obtain the Mel short-time Fourier transform magnitude spectrum. The formula is as follows:

[0111] MelS(n, f) = mel(S(n, f));

[0112] Among them, mel(·) is the Mel filter bank function, and MelS(n,f) is the Mel short-time Fourier transform magnitude spectrum.

[0113] S1024. Take the logarithm modulus of the Mel short-time Fourier transform magnitude spectrum to obtain the logarithmic Mel spectrogram;

[0114] Specifically, take the logarithm to obtain the logarithmic Mel short-time Fourier transform magnitude spectrum, that is, the logarithmic Mel spectrogram. The formula is as follows:

[0115] LogMelS(n,f) = 20log 10 MelS(n,f);

[0116] Among them, LogMelS(n,f) is the logarithmic Mel spectrogram.

[0117] S1025. Extract features from the logarithmic Mel spectrogram to obtain Mel filter bank features.

[0118] Specifically, cut the speech into speech chunks, take 600ms for each speech chunk, and then use streaming batch to extract Mel filter bank (Fbank) features and send them into the speech recognition model. The model outputs the speech recognition result to achieve streaming real-time speech recognition.

[0119] Furthermore, as an optional implementation manner, the data interaction display method further includes the step of pre-training a speech recognition model. This step of pre-training a speech recognition model can be further divided into the following steps A1021 to A1026:

[0120] A1021. Obtain a public speech dataset and a local speech dataset;

[0121] A1022. Extract features from the public speech dataset to obtain the first Mel filter bank feature set;

[0122] A1023. Extract features from the local speech dataset to obtain the second Mel filter bank feature set;

[0123] A1024. Perform parameter migration on the pre-trained Paraformer model, and input the first Mel filter bank feature set and the second Mel filter bank feature set into the Paraformer model after parameter migration;

[0124] A1025. Perform transfer learning on the Paraformer model after parameter migration training through cross-entropy loss, mean absolute error loss, and minimum word error rate loss to obtain a trained speech recognition model;

[0125] A1026. Set a confidence threshold. When the confidence of the speech recognition result is higher than the confidence threshold, input the corresponding user speech data into the local speech data set.

[0126] Optionally, as Figure 4 shown in a schematic diagram of a process for training a speech recognition model. In the embodiments of the present invention, a speech recognition model is obtained through transfer learning based on the Paraformer model. Other models can also be used for transfer training to obtain a speech recognition model according to data processing requirements, which is not limited herein.

[0127] Specifically, use a publicly available speech data set and a local speech data set collected locally for transfer learning. First, perform parameter transfer on the pre-trained Paraformer model and extract features from the two speech data sets. The feature extraction process also adopts the Fbank feature extraction method, as shown in the above steps S1021 to S1025. Then, input the first Mel filter bank feature set and the second Mel filter bank feature set obtained by feature extraction into training, and perform transfer learning on the Paraformer model after parameter transfer. The formula of the loss function used in the transfer learning process is as follows:

[0128]

[0129] It should be noted that the embodiments of the present invention use cross-entropy loss mean absolute error loss and minimum word error rate loss for joint loss training, where γ is a hyperparameter. The mean absolute error loss can effectively improve the ability of the model to extract speech features, while the cross-entropy loss and the minimum word error rate loss can effectively improve the speech recognition performance of the model. By combining the training of the three loss functions, the performance of the speech recognition model is effectively improved.

[0130] The cross-entropy loss has the following formula:

[0131]

[0132] where y i is the actual label of the i-th sample, is the predicted probability of the i-th sample, and N is the total number of samples.

[0133] The mean absolute error loss has the following formula:

[0134]

[0135] where yi is the actual label of the i-th sample, is the predicted probability of the i-th sample, and N is the total number of samples. is the sum of the absolute values of the differences between the predicted and actual values of all samples.

[0136] Minimum Word Error Rate Loss The formula is as follows:

[0137]

[0138] where y i is the number of word errors for the i-th sample, y * is the actual sentence label, and x is the input feature vector. is the average number of word errors in the sample, subtracting helps reduce the variance of gradient estimation and is important for stable training.

[0139] Furthermore, after the model transfer learning training is completed and it is put into use in the speech recognition process, if the confidence of the user's speech recognition result is higher than the preset confidence threshold, the user's speech is labeled and uploaded to update the local speech dataset, further enriching the local speech dataset for the next model transfer learning update.

[0140] It should be noted that in the embodiments of the present invention, a speech recognition model is obtained through a transfer learning method, reducing the dependence on a large amount of speech datasets required for model training, further reducing the model training time through model parameter transfer, effectively improving the model training effect, and making the communication between humans and machines more convenient by introducing a speech recognition interaction function, which can greatly reduce the manual processing time and effectively improve the work efficiency of employees.

[0141] S103. Retrieve and search for cases based on the speech recognition result through a preset intelligent data assistant to obtain the data extraction result required by the user;

[0142] Among them, the intelligent data assistant includes a retrieval-augmented generation framework and a large language model.

[0143] In some alternative embodiments, the speech recognition result is sent into a retrieval-augmented generation (RAG) framework based on a pre-trained model to return k similar cases to the large language model. The large language model finds the most similar case according to semantic analysis and retrieves the user's data extraction case library according to the data extraction SQL code corresponding to the most similar case to obtain the data extraction result required by the user.

[0144] Refer to Figure 5 , Figure 5This is a schematic diagram of a processing flow of the intelligent data assistant provided by the embodiments of the present invention. Further, as an optional implementation, the step of retrieving and searching for cases based on the speech recognition result through a preset intelligent data assistant to obtain the data extraction result required by the user can be further divided into the following steps S1031 to S1033:

[0145] S1031. Construct a retrieval augmented generation framework, and retrieve the speech recognition result through the retrieval augmented generation framework to obtain multiple similar cases;

[0146] Further, as an optional implementation, the retrieval augmented generation framework includes a bidirectional encoding representation model and a text retrieval algorithm. The step of retrieving the speech recognition result through the preset retrieval augmented generation framework to obtain multiple similar cases can be further divided into the following steps S10311 to S10313:

[0147] S10311. Construct a data query case library;

[0148] Specifically, as Figure 6 shown is a schematic diagram of a processing flow of the retrieval augmented generation framework. The retrieval augmented generation (RAG) framework includes a bidirectional encoding representation model (i.e., the BERT pre-trained model, Bidirectional Encoder Representations from Transformers - the bidirectional encoder based on Transformers) and the DPR text retrieval algorithm. Input approximately 3 million data query cases into the bidirectional encoding representation model (i.e., the BERT pre-trained model), and the model outputs data case query vectors and generates a data query case library for DPR text query. The formula is as follows:

[0149] Case_vector{p1,p2,...,p N}=BERT(Case(c1,c2,...,c N ));

[0150] Among them, c is the data query case, Case{·} is the data query case library, v is the data query case vector, Case_vector{·} is the data query case vector library, and BERT(·) is the BERT pre-trained model.

[0151] S10312. Input the speech recognition result into the bidirectional encoding representation model for vector encoding to obtain the speech problem vector and the corresponding vector encoding of the speech problem vector;

[0152] S10313. Through the text retrieval algorithm, retrieve according to the data query case library, the speech problem vector, and the vector encoding to obtain multiple similar cases.

[0153] Specifically, as Figure 6 shown, the speech recognition result is input into a bidirectional encoding representation model (i.e., the BERT pre-trained model) for vector encoding to obtain the user's speech question vector q and vector encoding V q = E Q (q), and then the DPR text retrieval algorithm is used to retrieve the top k similar cases closest to V q . Its algorithm formula is as follows:

[0154]

[0155] sim(q,p) = E Q (q) T E P (p);

[0156] where sim(q,p) is to calculate the dot product of the speech question vector q and the case vector p, that is, the similarity. And the algorithm L(·) is to calculate and obtain the top k similar cases according to the similarity, is the first similar case, is the second similar case, is the kth similar case.

[0157] Furthermore, during the retrieval, FAISS (Facebook AI Similarity Search) is used to index them. FAISS is a very efficient open-source library for similarity search and clustering, which can be easily applied to billions of vectors.

[0158] S1032. Construct the question text, input the question text and each similar case into the large language model for case search to obtain the most similar case;

[0159] Specifically, as Figure 7 shown is a schematic flow diagram for determining the most similar case based on the large language model. First, construct the question text, for example, "You are a data analyst, please analyze and return the most similar case according to the following k cases", and send the k similar cases and the question text obtained in the previous step into the large language model and return the most similar case. Among them, the large language model can select models with the ability of general agent question answering such as Telechat and Qwen according to requirements such as deployment cost, number of parameters, and model answering ability.

[0160] S1033. Construct the user data extraction case library, extract data from the user data extraction case library according to the SQL code of the most similar case to obtain the data extraction result required by the user.

[0161] Specifically, as Figure 8The following is a schematic diagram of a user data extraction case library. A user data extraction case library is established based on approximately 3 million existing cases. This user data extraction case library is a case library based on Parquet (a columnar storage format for Apache Hadoop), which helps to speed up the search speed of the embodiments of the present invention. The data framework established is as follows:

[0162] The i-th case = {the description of the i-th case, the PGSQL statement of the i-th case};

[0163] According to the PGSQL statement corresponding to the most similar case obtained in the previous step, it is sent to the user data extraction case library (Postgresql, PG) database to execute data query and return the data extraction result required by the user.

[0164] It should be noted that the embodiments of the present invention effectively analyze the user's voice instructions through a retrieval-enhanced generation framework and a large language model based on a pre-trained model, quickly search and obtain the most similar cases from millions of case libraries, and extract data according to the PGSQL statements behind the most similar cases, returning the data required for different user instructions, further improving the data access efficiency, and being able to help users with the intuitive understanding and quick analysis of data.

[0165] S104. Display the data extraction result required by the user according to the numerical type corresponding to the data extraction result required by the user.

[0166] Further as an optional implementation manner, the step of displaying the data extraction result required by the user according to the numerical type corresponding to the data extraction result required by the user can be further divided into the following steps S1041 and S1042:

[0167] S1041. Construct a data display template and determine the numerical type corresponding to the data extraction result required by the user;

[0168] S1042. Display the data extraction result required by the user according to the data display template and the numerical type.

[0169] Specifically, since the data extraction result required by the user is a numerical table, and the numerical tables extracted by different users are quite different, certain judgments are needed to effectively display the business data. Therefore, judgments and corresponding displays are made according to the numerical type corresponding to the data extraction result required by the user.

[0170] Exemplarily, as Figure 9 shown in the example diagram of the display of the data extraction result required by the user, when the numerical type corresponding to the data extraction result required by the user is a single value, the single value is directly displayed; when the numerical type corresponding to the data extraction result required by the user is multiple values, the numerical values of the time type are displayed in a line chart, the numerical values of the region type are displayed in a bar chart, and the numerical values of the person type are displayed in a table.

[0171] The above description of the method for data interaction display based on voice interaction in the embodiments of the present invention has been given. It can be recognized that compared with the existing data query technologies, the embodiments of the present invention have the following advantages:

[0172] First, the Paraformer speech recognition model based on transfer learning is adopted. In existing data query technologies, manual input queries or manual screening are used. In the face of massive business data queries, only manual input queries often consume time and effort, reducing the interaction and experience of users. In the embodiments of the present invention, by introducing the voice recognition interaction function, the communication between humans and machines becomes more convenient, which can greatly reduce the manual processing time and effectively improve the work efficiency of employees.

[0173] Second, a joint loss function used for speech recognition training for transfer learning is provided. Through the cross-entropy loss, mean absolute error loss, and the speech recognition model trained by loss function transfer learning, the accuracy of speech recognition is effectively improved, enhancing the user experience effect.

[0174] Third, a retrieval-enhanced generation framework based on a pre-trained model and an intelligent data assistant of a large language model are provided. Through the retrieval-enhanced generation framework based on the pre-trained model, k similar cases are quickly retrieved and calculated, and the large language model is used to ask questions to the intelligent agent to return the most similar case from the k similar cases to complete the query and fetch data. Finally, the data required by the user is returned, effectively realizing the need for employees to quickly view data, and based on the numerical type of the data required by the user, real-time display and analysis of the data are realized, improving the decision-making ability and work efficiency of employees.

[0175] Referring to Figure 10 , the embodiments of the present invention also provide a data interaction display system based on voice interaction, including:

[0176] A data preprocessing module, configured to obtain user voice data and preprocess the user voice data to obtain a first voice signal;

[0177] A voice enhancement and recognition module, configured to extract features from the first voice signal to obtain Mel filter bank features, and input the Mel filter bank features into a pre-trained speech recognition model to obtain a speech recognition result;

[0178] A data fetching result determination module, configured to retrieve and search for cases for the speech recognition result through a preset intelligent data assistant to obtain the data fetching result required by the user;

[0179] A data interaction display module, configured to display the data fetching result required by the user according to the numerical type corresponding to the data fetching result required by the user;

[0180] Among them, the intelligent data assistant includes a retrieval enhanced generation framework and a large language model.

[0181] The content in the above embodiments of the data interaction display method based on voice interaction is applicable to the embodiments of the data interaction display system based on voice interaction. The functions specifically implemented by the embodiments of the data interaction display system based on voice interaction are the same as those of the above embodiments of the data interaction display method based on voice interaction, and the beneficial effects achieved are also the same as those of the above embodiments of the data interaction display method based on voice interaction.

[0182] An embodiment of the present invention also provides an electronic device, which includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing the connection and communication between the processor and the memory. When the program is executed by the processor, it realizes the above-mentioned data interaction display method based on voice interaction. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0183] As Figure 11 shown is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. Referring to Figure 11 , an embodiment of the present invention provides an electronic device, including:

[0184] A processor 1001, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;

[0185] A memory 1002, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002, and the processor 1001 is called to execute the data interaction display method based on voice interaction in the embodiments of the present invention;

[0186] An input / output interface 1003, which is used to implement information input and output;

[0187] A communication interface 1004, which is used to implement communication and interaction between this device and other devices. It can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0188] A bus 1005, which transmits information between various components of the device (such as a processor 1001, a memory 1002, an input / output interface 1003, and a communication interface 1004);

[0189] Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 are communicatively connected to each other inside the device through the bus 1005.

[0190] An embodiment of the present invention also provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned data interaction display method based on voice interaction.

[0191] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0192] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.

[0193] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order noted in the operational illustrations. For example, depending upon the functionality / operation involved, two blocks shown in succession may in fact be executed substantially concurrently or the blocks may sometimes be executed in the reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example in order to provide a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.

[0194] Moreover, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above-described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Accordingly, those of ordinary skill in the art will be able to implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, the scope of which is determined by the full scope of the appended claims and their equivalents.

[0195] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0196] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0197] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the above programs can be printed, because the above programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or, if necessary, other suitable processing, and then storing them in a computer memory.

[0198] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), and the like.

[0199] In the above description of this specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0200] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0201] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without violating the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A data interactive display method based on voice interaction, characterized in that: The following steps are involved: Acquire user voice data, and preprocess the user voice data to obtain a first voice signal; Extracting features from the first speech signal to obtain Mel filter bank features, and inputting the Mel filter bank features into a pre-trained speech recognition model to obtain a speech recognition result; The voice recognition results are retrieved and case-searched through a preset intelligent data assistant to obtain the data acquisition results required by the user; Display the data acquisition result required by the user according to the numerical type corresponding to the data acquisition result required by the user; Wherein, the intelligent data assistant includes a retrieval enhancement generation framework and a large language model.

2. According to the data interactive display method based on voice interaction according to claim 1, it is characterized in that: The preprocessing of the user voice data to obtain the first voice signal specifically includes: Normalizing the user voice data to obtain a second voice signal; The second speech signal is pre-emphasized to obtain the first speech signal.

3. The data interactive display method based on voice interaction according to claim 1, characterized in that: The extracting features of the first speech signal to obtain Mel filter bank features specifically includes: Performing frame windowing on the first speech signal in the time domain to obtain a time domain signal; Performing a short-time Fourier transform on the time domain signal to obtain a short-time Fourier transform amplitude spectrum; Setting a Mel filter bank, inputting the short-time Fourier transform amplitude spectrum into the Mel filter bank, and obtaining a Mel short-time Fourier transform amplitude spectrum; Taking the logarithmic modulus of the Mel short-time Fourier transform amplitude spectrum to obtain a logarithmic Mel time-frequency spectrum diagram; Feature extraction is performed on the logarithmic Mel-spectrogram to obtain the Mel filter bank feature.

4. The data interactive display method based on voice interaction according to claim 1, characterized in that: The data interactive display method further includes the step of pre-training the speech recognition model, wherein the pre-training of the speech recognition model specifically includes: Obtain public speech datasets and local speech datasets; Performing feature extraction on the public speech data set to obtain a first Mel filter bank feature set; Performing feature extraction on the local speech data set to obtain a second Mel filter bank feature set; Performing parameter migration on a pre-trained Paraformer model, and inputting the first Mel filter group feature set and the second Mel filter group feature set into the Paraformer model after parameter migration; Performing transfer learning on the Paraformer model after parameter transfer training through cross entropy loss, mean absolute error loss, and minimum word error rate loss to obtain the trained speech recognition model; A confidence threshold is set, and when the confidence of the speech recognition result is higher than the confidence threshold, the corresponding user speech data is input into the local speech data set.

5. The data interactive display method based on voice interaction according to claim 1 is characterized in that: The preset intelligent data assistant searches the speech recognition results and case searches to obtain the data acquisition results required by the user, specifically including: Constructing the retrieval enhancement generation framework, and searching the speech recognition results through the retrieval enhancement generation framework to obtain multiple similar cases; Constructing a question text, inputting the question text and each of the similar cases into the large language model to perform case search, and obtaining the most similar case; A user data acquisition case library is constructed, and data is acquired from the user data acquisition case library according to the SQL code of the most similar case to obtain the data acquisition result required by the user.

6. A data interactive display method based on voice interaction according to claim 5, characterized in that: The retrieval enhancement generation framework includes a bidirectional encoding representation model and a text retrieval algorithm. The speech recognition result is retrieved through the preset retrieval enhancement generation framework to obtain multiple similar cases, specifically including: Build a data query case library; Inputting the speech recognition result into the bidirectional coding representation model for vector coding to obtain a speech question vector and a vector coding corresponding to the speech question vector; Through the text retrieval algorithm, a search is performed based on the data query case library, the voice question vector and the vector code to obtain a plurality of similar cases.

7. A data interactive display method based on voice interaction according to any one of claims 1 to 6, characterized in that: The displaying of the data acquisition result required by the user according to the numerical type corresponding to the data acquisition result required by the user specifically includes: Construct a data display template and determine the value type corresponding to the data acquisition result required by the user; The data acquisition result required by the user is displayed according to the data display template and the value type.

8. A data interactive display system based on voice interaction, characterized in that: include: A data preprocessing module, used to obtain user voice data, preprocess the user voice data, and obtain a first voice signal; A speech enhancement and recognition module, configured to extract features from the first speech signal to obtain Mel filter bank features, and input the Mel filter bank features into a pre-trained speech recognition model to obtain a speech recognition result; The data acquisition result determination module is used to retrieve and search the voice recognition results through a preset intelligent data assistant to obtain the data acquisition results required by the user; A data interactive display module is used to display the data acquisition result required by the user according to the numerical type corresponding to the data acquisition result required by the user; Wherein, the intelligent data assistant includes a retrieval enhancement generation framework and a large language model.

9. An electronic device, characterized in that: The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the steps of the data interaction display method based on voice interaction as described in any one of claims 1 to 7 are realized.

10. A storage medium, the storage medium being a computer-readable storage medium, used for computer-readable storage, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the data interaction display method based on voice interaction as described in any one of claims 1 to 7.