Data processing method and device, electronic equipment and storage medium

By using the target model to generate prediction statements and comparing semantic similarity with the recognition statements, the problem of insufficient real-time and lack of comprehensive context in the prior art is solved, and more efficient and accurate speech recognition is achieved.

CN120126469APending Publication Date: 2025-06-10联想诺谛(北京)智能科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510122792.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing audio data processing methods have shortcomings in real-time data processing and context comprehensive considerations, which cannot meet the real-time data processing requirements and lack comprehensive considerations for context.

Method used

By obtaining the audio data before the first moment in the current conversation, using the target model to generate prediction statements, and combining the recognition statements after the first moment to perform semantic similarity comparison, the target statements are determined.

Benefits of technology

It improves the accuracy and coherence of speech recognition, and quickly generates corresponding statements in real-time continuous dialogue scenarios, which improves the speed of data processing and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126469A_ABST
    Figure CN120126469A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring audio data before a first moment in a current dialogue; based on the audio data before the first moment, utilizing a target model to generate a prediction statement; the target model is used for predicting the following information of the audio data before the first moment; identifying the target audio data after the first moment to obtain an identification statement; and determining a target statement corresponding to the target audio data after the first moment based on the semantic similarity between the prediction statement and the recognition statement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to computer technology, and in particular to a data processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Currently, for the method of processing audio data, in the scenario of audio data recognition, the audio data is recognized by a language model. However, due to the slow sentence generation speed of the language model, it cannot meet the real-time data processing of audio data, and it can only improve the semantic coherence of single sentences in the audio data, lacking the comprehensive consideration of the context. Summary of the Invention

[0003] Embodiments of this application provide a data processing method, apparatus, electronic device, and storage medium.

[0004] According to the first aspect of this application, a data processing method is provided. The method includes: obtaining audio data before a first moment in a current conversation;

[0005] Based on the audio data before the first moment, using a target model to generate a predicted sentence; the target model is used to predict the subsequent information of the audio data before the first moment;

[0006] Identifying target audio data after the first moment to obtain an identified sentence;

[0007] Based on the semantic similarity between the predicted sentence and the identified sentence, determining a target sentence corresponding to the target audio data after the first moment.

[0008] According to an embodiment of this application, the method further includes:

[0009] Obtaining a speech recognition result corresponding to the audio data before the first moment;

[0010] Based on the speech recognition result and the set prompt word information, using the target model to generate multiple predicted sentences.

[0011] According to an embodiment of this application, the audio data before the first moment at least includes continuous conversation audio in the current conversation; the continuous conversation audio represents the conversation audio located before the first moment in the time series;

[0012] The predicted sentence at least includes multiple text representations corresponding to the subsequent information of the audio data before the first moment predicted.

[0013] According to an embodiment of this application, the identifying the target audio data after the first moment to obtain an identified sentence includes:

[0014] Performing speech recognition on the target audio data after the first moment based on a speech recognition model to obtain the recognition statement of the target audio data;

[0015] The speech recognition model is used to extract acoustic features from the target audio data and determine the corresponding text representation based on the acoustic features;

[0016] The recognition statement of the target audio data at least includes the text representation of the target audio data.

[0017] According to an embodiment of the present application, the method further includes:

[0018] Based on a semantic vectorization model, converting the prediction statement into a first text vector;

[0019] Based on the semantic vectorization model, converting the recognition statement of the target audio data into a second text vector;

[0020] Based on the first text vector and the second text vector, determining the semantic similarity between the prediction statement and the recognition statement of the target audio data.

[0021] According to an embodiment of the present application, the determining the target statement corresponding to the target audio data after the first moment based on the semantic similarity between the prediction statement and the recognition statement includes:

[0022] In response to the semantic similarity being greater than or equal to a set similarity threshold, determining the recognition statement of the target audio data as the target statement.

[0023] According to an embodiment of the present application, the determining the target statement corresponding to the target audio data after the first moment based on the semantic similarity between the prediction statement and the recognition statement includes:

[0024] In response to the semantic similarity being less than the set similarity threshold, determining the phoneme sequence of the target audio data;

[0025] Obtaining the continuous conversation text corresponding to the audio data before the first moment in the current conversation;

[0026] Based on the continuous conversation text and the phoneme sequence, using a large language model to generate the target statement corresponding to the target audio data; the large language model is used to generate the target statement corresponding to the phoneme sequence according to the context corresponding to the continuous conversation text.

[0027] According to the second aspect of the present application, there is provided a data processing device, and the data processing device includes:

[0028] An acquisition module, configured to acquire audio data before a first moment in a current conversation;

[0029] A prediction module, configured to generate a prediction statement based on the audio data before the first moment by using a target model; the target model is used to predict subsequent information of the audio data before the first moment;

[0030] An identification module, configured to identify target audio data after the first moment to obtain an identification statement;

[0031] A determination module, configured to determine a target statement corresponding to the target audio data after the first moment based on a semantic similarity between the prediction statement and the identification statement.

[0032] According to a third aspect of the present application, there is provided an electronic device, including:

[0033] At least one processor; and

[0034] A memory communicatively connected to the at least one processor; wherein,

[0035] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the present application.

[0036] According to a fourth aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the method described in the present application. Description of the Drawings

[0037] By referring to the drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present application will become easy to understand. In the drawings, several embodiments of the present application are shown in an exemplary rather than restrictive manner, where:

[0038] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.

[0039] Figure 1 Shows a schematic processing flow of the data processing method provided by an embodiment of the present application Figure 1 ;

[0040] Figure 2 Shows a schematic processing flow of the data processing method provided by an embodiment of the present application Figure 2 ;

[0041] Figure 3 Shows a schematic processing flow of the data processing method provided by an embodiment of the present application Figure 3 ;

[0042] Figure 4 shows the processing flow diagram of the data processing method provided by the embodiments of the present application Figure 4 ;

[0043] Figure 5 shows the processing flow diagram of the data processing method provided by the embodiments of the present application Figure 5 ;

[0044] Figure 6 shows an application scenario diagram of the data processing method provided by the embodiments of the present application;

[0045] Figure 7 shows an optional schematic diagram of the data processing device provided by the embodiments of the present application;

[0046] Figure 8 shows the composition structure schematic diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0047] To make the objectives, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application.

[0048] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0049] In the following description, the terms "first / second" involved are only used to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.

[0051] The processing flow in the data processing method provided by the embodiments of the present application will be described. Refer to Figure 1 , Figure 1 is the processing flow diagram of the data processing method provided by the embodiments of the present applicationFigure 1 , it will be described in conjunction with Figure 1 the steps S101 - S104 shown.

[0052] Step S101, obtain the audio data before the first moment in the current conversation.

[0053] In some embodiments, the current conversation may include: conversations in real - time speech recognition scenarios and conversations in long speeches. For example, the current conversation may include conversations in scenarios such as live TV broadcasts, video conferences, and human - machine conversations. The first moment may include: a specific time point during the current conversation, used to distinguish the front and back parts of the conversation.

[0054] Step S102, based on the audio data before the first moment, use the target model to generate a predicted statement; the target model is used to predict the subsequent information of the audio data before the first moment.

[0055] In some embodiments, the audio data before the first moment may include: audio such as conversation content, user voice commands, and speech audio in a meeting before the first moment. The audio data before the first moment may include continuous conversation audio in the current conversation; the continuous conversation audio represents the conversation audio located before the first moment in the time series. The embodiments of the present application do not limit the specific audio data. The predicted statement may include: a predicted text statement of the content of the audio data after the first moment generated by the target model according to the audio data before the first moment. The predicted statement may include multiple text representations corresponding to the subsequent information of the audio data before the first moment predicted. The embodiments of the present application do not limit the specific number of predicted statements.

[0056] Step S103, identify the target audio data after the first moment to obtain an identified statement.

[0057] In some embodiments, the target audio data after the first moment may include: audio such as conversation content, user voice commands, and speech audio in a meeting after the first moment. The embodiments of the present application do not limit the specific audio data. The identified statement may include: a text statement obtained after performing speech recognition processing on the target audio data after the first moment.

[0058] Step S104, based on the semantic similarity between the predicted statement and the identified statement, determine the target statement corresponding to the target audio data after the first moment.

[0059] In some embodiments, the semantic similarity may include: the similarity of text vectors. The embodiments of the present application do not limit the specific semantic similarity. The target statement may be a statement used to represent the content of the target audio data after the first moment.

[0060] The method of the embodiment of the present application can effectively improve the accuracy and coherence of speech recognition by using the audio data before the first moment and the target model to generate a predicted sentence, and then combining the recognized sentence after the first moment for semantic similarity comparison. In a real-time continuous dialogue scenario, the current voice segment and the following information can be combined to quickly generate the corresponding sentence, which improves the speed of data processing and enhances the user experience.

[0061] In some embodiments, the processing flow of the data processing method is shown as follows: Figure 2 ,like Figure 2 As shown, the data processing method may further include:

[0062] Step S201, obtaining a speech recognition result corresponding to the audio data before a first moment.

[0063] Step S202: Based on the speech recognition result and the set prompt word information, a plurality of prediction sentences are generated using the target model.

[0064] In this embodiment, the audio data before the first moment is recognized by an ASR (Automatic Speech Recognition) model to obtain a speech recognition result corresponding to the audio data. The speech recognition result may include a recognition sentence corresponding to the audio data before the first moment. The set prompt word information may include: a pre-set keyword for prediction. The prompt word information may be used to assist the target model in recognizing and understanding the input speech recognition result.

[0065] As an example, the ASR model is used to perform speech recognition on the audio data before the first moment, and the speech recognition result is "I bought a new computer". The set prompt word information may include: "You are an expert in audio translation. You hear [I bought a new computer]. Please predict 10 sentences that the person may say next, and each sentence is presented in a line break." Using the target model, the generated multiple prediction sentences may include:

[0066] “I like its design and performance.

[0067] The computer screen display effect is very good.

[0068] I find the computer keyboard very comfortable to type on.

[0069] The computer is also very cost-effective and worth buying.

[0070] I'm installing some prerequisite software.

[0071] But I was a little uncomfortable at the beginning and I’m still adapting.

[0072] I'm considering buying another accessory to enhance the user experience.

[0073] Are there any recommended software or setting tips?

[0074] I heard its battery life is very good and I hope to try it out soon.

[0075] Generally speaking, this purchase experience is quite satisfactory. Thank you for your advice!

[0076] The method of the embodiment of the present application generates a prediction statement through the speech recognition result and the set prompt word information, further enhancing the accuracy and coherence of the prediction statement. In a real-time continuous dialogue scenario, it can combine the current speech segment and the following information to quickly generate the corresponding statement, improving the speed of data processing and enhancing the user experience.

[0077] In some embodiments, a schematic diagram of the processing flow of the data processing method Figure 3 is shown as Figure 3 shown. Identifying the target audio data after the first moment in step S103 to obtain an identification statement may specifically include:

[0078] Step S301: Perform speech recognition on the target audio data after the first moment based on a speech recognition model to obtain an identification statement of the target audio data.

[0079] Step S302: Determine the target statement corresponding to the target audio data after the first moment based on the semantic similarity between the prediction statement and the identification statement.

[0080] As an example, in the scenario of generating subtitles from audio, the prediction statement may include: "The screen display effect of the computer is very good". Through the ASR model, perform speech recognition on the target audio data after the first moment to obtain an identification statement of the target audio data as "The screen of the computer is very large". Determine whether the semantic similarity between the identification statement and the prediction statement is greater than or equal to the set similarity threshold. In response to the semantic similarity being greater than or equal to the set similarity threshold, determine "The screen of the computer is very large" as the target statement. The target statement can be output as a subtitle.

[0081] The method of the embodiment of the present application extracts the acoustic features in the target audio data through a speech recognition model and converts them into an identification statement. According to the similarity between the identification statement and the prediction statement, determine the target statement corresponding to the target audio data. Further enhancing the accuracy and coherence of the prediction statement. In a real-time continuous dialogue scenario, it can combine the current speech segment and the following information to quickly generate the corresponding statement, improving the speed of data processing and enhancing the user experience.

[0082] In some embodiments, a schematic diagram of the processing flow of the data processing methodFigure 4 , such as Figure 4 shown, the data processing method may further include:

[0083] Step S401, based on the semantic vectorization model, convert the prediction statement into a first text vector.

[0084] Step S402, based on the semantic vectorization model, convert the recognition statement of the target audio data into a second text vector.

[0085] Step S403, based on the first text vector and the second text vector, determine the semantic similarity between the prediction statement and the recognition statement of the target audio data.

[0086] As an example, in a meeting record scenario, based on the speech audio in the meeting before the first moment in a video conference, use the target model to generate a prediction statement. Based on the semantic vectorization model, convert the prediction statement into a first text vector. Perform speech recognition on the target audio data after the first moment based on the speech recognition model to obtain the recognition statement of the target audio data. Based on the semantic vectorization model, convert the recognition statement of the target audio data into a second text vector. Calculate the cosine similarity between the first text vector and the second text vector to determine the semantic similarity between the prediction statement and the recognition statement.

[0087] The method of the embodiment of the present application determines the similarity between the recognition statement and the prediction statement through the semantic vectorization model, and then determines whether to rewrite the recognition statement according to the semantic similarity. Further enhances the accuracy and coherence of the prediction statement. In a real-time continuous conversation scenario, it can combine the current speech segment and the following information to quickly generate corresponding statements, improving the speed of data processing and enhancing the user experience.

[0088] In some embodiments, the processing flow of the data processing method is schematically shown Figure 5 , such as Figure 5 shown, based on the semantic similarity between the prediction statement and the recognition statement in step S104, determining the target statement corresponding to the target audio data after the first moment may include:

[0089] Step S501, perform recognition on the target audio data after the first moment to obtain a recognition statement.

[0090] For the specific description of step S501, it is the same as step S103 and will not be elaborated here.

[0091] Step S502a, in response to the semantic similarity being greater than or equal to the set similarity threshold, determine the recognition statement of the target audio data as the target statement.

[0092] Step S502b, in response to the semantic similarity being less than a set similarity threshold, determine the phoneme sequence of the target audio data.

[0093] Step S503b, obtain the continuous dialogue text corresponding to the audio data before the first moment in the current dialogue.

[0094] Step S504b, based on the continuous dialogue text and the phoneme sequence, use a large language model to generate the target statement corresponding to the target audio data.

[0095] In this embodiment, the similarity threshold may include: a preset threshold. The similarity threshold can be used to determine the target statement. The phoneme sequence may include: the basic phoneme units corresponding to the target audio data. The continuous dialogue text may include: the dialogue text corresponding to the audio data before the first moment, that is, the processed dialogue content in the previous dialogue. The large language model may include: the GPT (Generative Pre-trained Transformer) model and the BERT (Bidirectional Encoder Representations from Transformers) model. The large language model may also include other models, and the embodiments of the present application do not limit the specific large language model. The large language model can be used to generate the target statement corresponding to the phoneme sequence according to the context corresponding to the continuous dialogue text. The processing speed of the large language model is less than the processing speed of the semantic vectorization model.

[0096] As an example, the predicted statements may include: "I quite like its design and performance. The screen display effect of the computer is very good. I think the keyboard of the computer is very comfortable to type on. The cost performance of the computer is also very high and it is well worth buying. I am installing some necessary software. But at first, I was a little unaccustomed and still getting used to it. I am considering buying another accessory to enhance the usage experience. Are there any recommended software or setting tips? I heard that its battery life is very good and I hope to try it out as soon as possible. Generally speaking, this purchase experience is quite satisfactory. Thank you for your advice!" By using the ASR model, perform speech recognition on the target audio data after the first moment, and the recognized statement of the target audio data is "The screen of the computer is very large". Determine the semantic similarity between the recognized statement and each predicted statement. The set similarity threshold can be 0.9, and determine whether each semantic similarity is greater than or equal to 0.9. In response to the semantic similarity between the recognized statement and each predicted statement being greater than 0.9, then determine "The screen of the computer is very large" as the target statement.

[0097] As an example, through the ASR model, speech recognition is performed on the target audio data after the first moment, and the recognized sentence of the target audio data is "The screen of the pad brain is very large". For sentences with similar pronunciations but different semantics, their semantic similarity is relatively low. Determine the semantic similarity between the recognized sentence and each predicted sentence. The set similarity threshold can be 0.9, and determine whether each semantic similarity is greater than or equal to 0.9. In response to the semantic similarity between the recognized sentence and each predicted sentence being less than 0.9, determine that the phoneme sequence of the target audio data is

dian nao de ping mu hen da

[0098] Obtain the continuous dialogue text corresponding to the audio data before the first moment in the current conversation. The continuous dialogue text includes:

Summer vacation has started

This time I'm going to have a good time playing games

I bought a new computer

Summer vacation has started

This time I'm going to have a good time playing games

I bought a new computer

dian nao de ping mu hen da

[0099] The method of the embodiment of the present application determines the similarity between the recognized sentence and the predicted sentence through the semantic vectorization model, and then determines whether to rewrite the recognized sentence according to the semantic similarity. Further enhances the accuracy and coherence of the predicted sentence. In a real-time continuous dialogue scenario, use the large language model to predict the next sentence, compare the semantic similarity between the predicted sentence and the recognized sentence, and thus decide whether to rewrite. For audio recognition without context dependence, almost no recognition time is lost, ensuring the performance of data processing; for audio recognition with a small amount of context dependence, the large model can be dynamically used for context-based rewriting, so as to quickly generate the corresponding sentence while ensuring the accuracy of audio recognition, improving the speed of data processing and enhancing the user experience.

[0100] Reference Figure 6 , a scenario diagram of an application of the data processing method provided by the embodiment of the present application, which is applied to data processing in a multi-round dialogue scenario.

[0101] Obtain the audio data of the human voice sentence 1 in the current conversation. Through the ASR model, perform speech recognition on the audio data before the first moment to obtain the speech recognition result S n-1It is "I just bought a new computer". Speech recognition can include acoustic feature extraction and post-processing based on an LM (Language Model). Next sentence prediction is performed on the speech recognition result based on an LLM (Large Language Model). The prompt information set by the LLM can include: "You are a listening and translation expert. You hear [I just bought a new computer], and please predict 10 sentences that the person may say next, each presented on a new line." Using the target model, multiple predicted sentences {S n_predict} can include:

[0102] "I quite like its design and performance.

[0103] The screen display effect of the computer is very good.

[0104] I think the keyboard of the computer is very comfortable to type on.

[0105] The cost performance of the computer is also very high and it is well worth buying.

[0106] I am installing some necessary software.

[0107] However, it was a bit unfamiliar at first and I'm still getting used to it.

[0108] I'm considering buying another accessory to enhance the usage experience.

[0109] Are there any recommended software or setting tips?

[0110] I heard that its battery life is very good and I hope to try it out soon.

[0111] Generally speaking, this purchase experience is quite satisfactory. Thank you for your advice!"

[0112] Through the ASR model, speech recognition is performed on the audio data of the human voice sentence 2 to obtain the recognition sentence S of the target audio data n which is "The screen of the computer is very big". Based on the semantic vectorization model, the semantic similarity between S n and {S n_predict} is compared. In response to the fact that there are 7 or more of the 10 predicted sentences with a semantic similarity greater than the set similarity threshold to the recognition sentence, the recognition sentence S n is output as the target sentence corresponding to the human voice sentence 2.

[0113] In response to the fact that there are not 7 or more of the 10 predicted sentences with a semantic similarity greater than the set similarity threshold to the recognition sentence, the phoneme sequence of the human voice sentence 2 is determined as

dian nao de ping mu hen da

Summer vacation has started

I'm going to have a good time playing games this time

I bought a new computer

Summer vacation has started

I'm going to have a good time playing games this time

I bought a new computer

dian nao de ping mu hen da

[0114] Next, continue to describe the exemplary structure of the software modules included in the data processing device provided by the embodiments of the present application. In some embodiments, as Figure 7 shown, the data processing device may include: an obtaining module 701, configured to obtain audio data before the first moment in the current dialogue; a prediction module 702, configured to generate a prediction statement based on the audio data before the first moment by using a target model; the target model is used to predict the subsequent information of the audio data before the first moment; an identification module 703, configured to identify the target audio data after the first moment to obtain an identification statement; a determination module 704, configured to determine the target statement corresponding to the target audio data after the first moment based on the semantic similarity between the prediction statement and the identification statement.

[0115] In some embodiments, the prediction module 702 may further be configured to: obtain the speech recognition result corresponding to the audio data before the first moment; generate multiple prediction statements based on the speech recognition result and the set prompt word information by using the target model.

[0116] In some embodiments, the audio data before the first moment at least includes the continuous dialogue audio in the current dialogue; the continuous dialogue audio represents the dialogue audio located before the first moment in the time series; the prediction statement at least includes multiple text representations corresponding to the subsequent information of the audio data before the first moment.

[0117] In some embodiments, the identification module 703 may be configured to: perform speech recognition on the target audio data after the first moment based on a speech recognition model to obtain an identification statement of the target audio data; the speech recognition model is used to extract acoustic features from the target audio data and determine the corresponding text representation based on the acoustic features; the identification statement of the target audio data at least includes the text representation of the target audio data.

[0118] In some embodiments, the data processing device may further include a judgment module, and the judgment module may be configured to: based on a semantic vectorization model, convert a prediction statement into a first text vector; based on the semantic vectorization model, convert an identification statement of target audio data into a second text vector; and based on the first text vector and the second text vector, determine the semantic similarity between the prediction statement and the identification statement of the target audio data.

[0119] In some embodiments, the determination module 704 may be configured to: in response to the semantic similarity being greater than or equal to a set similarity threshold, determine the identification statement of the target audio data as the target statement.

[0120] In some embodiments, the determination module 704 may be configured to: in response to the semantic similarity being less than the set similarity threshold, determine the phoneme sequence of the target audio data; obtain the continuous dialogue text corresponding to the audio data before the first moment in the current dialogue; and based on the continuous dialogue text and the phoneme sequence, use a large language model to generate the target statement corresponding to the target audio data, where the large language model is used to generate the target statement corresponding to the phoneme sequence according to the context corresponding to the continuous dialogue text.

[0121] According to an embodiment of the present application, the present application also provides an electronic device and a non-transitory computer-readable storage medium.

[0122] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that may be used to implement the embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present application described herein and / or claimed.

[0123] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which may perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 may also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0124] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0125] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital data processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the data processing method by any other suitable means (e.g., by means of firmware).

[0126] The various embodiments of the systems and technologies described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0127] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0128] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0129] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, speech input, or tactile input).

[0130] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0131] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.

[0132] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved. There is no limitation herein.

[0133] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" can explicitly or implicitly include at least one of such features. In the description of this application, "a plurality" means two or more, unless otherwise specifically defined.

[0134] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A data processing method, comprising: Get the audio data before the first moment in the current conversation; Generate a prediction sentence based on the audio data before the first moment using the target model; The target model is used to predict the following information of the audio data before the first moment; Recognize the target audio data after the first moment to obtain a recognition sentence; Based on the semantic similarity between the predicted sentence and the recognized sentence, a target sentence corresponding to the target audio data after the first moment is determined.

2. The method according to claim 1, further comprising: Obtaining a speech recognition result corresponding to the audio data before the first moment; Based on the speech recognition result and the set prompt word information, a plurality of predicted sentences are generated using a target model.

3. The method according to claim 1, wherein the audio data before the first moment at least includes continuous conversation audio in the current conversation; the continuous conversation audio represents the conversation audio before the first moment in time sequence; The predicted sentence includes at least a plurality of text representations corresponding to the predicted context information of the audio data before the first moment.

4. The method according to claim 1, wherein the step of recognizing the target audio data after the first moment to obtain a recognition sentence comprises: Performing speech recognition on the target audio data after the first moment based on the speech recognition model to obtain a recognition sentence of the target audio data; The speech recognition model is used to extract acoustic features from the target audio data and determine corresponding text representations based on the acoustic features; The recognized sentence of the target audio data includes at least a textual representation of the target audio data.

5. The method according to claim 1, further comprising: Based on the semantic vectorization model, converting the predicted sentence into a first text vector; Based on the semantic vectorization model, converting the recognition sentence of the target audio data into a second text vector; Based on the first text vector and the second text vector, a semantic similarity between the predicted sentence and a recognized sentence of the target audio data is determined.

6. The method according to claim 1, wherein determining the target sentence corresponding to the target audio data after the first moment based on the semantic similarity between the predicted sentence and the recognized sentence comprises: In response to the semantic similarity being greater than or equal to a set similarity threshold, the recognized sentence of the target audio data is determined as the target sentence.

7. The method according to claim 1, wherein determining the target sentence corresponding to the target audio data after the first moment based on the semantic similarity between the predicted sentence and the recognized sentence comprises: In response to the semantic similarity being less than a set similarity threshold, determining a phoneme sequence of the target audio data; Obtaining a continuous conversation text corresponding to the audio data before the first moment in the current conversation; Based on the continuous dialogue text and the phoneme sequence, a target sentence corresponding to the target audio data is generated using a large language model; the large language model is used to generate a target sentence corresponding to the phoneme sequence according to the context corresponding to the continuous dialogue text.

8. A data processing device, comprising: An acquisition module, used to obtain audio data before the first moment in the current conversation; A prediction module, configured to generate a prediction statement based on the audio data before the first moment using a target model; the target model is configured to predict the following information of the audio data before the first moment; A recognition module, used to recognize the target audio data after the first moment to obtain a recognition sentence; A determination module is used to determine a target sentence corresponding to the target audio data after the first moment based on the semantic similarity between the predicted sentence and the recognized sentence.

9. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method according to any one of claims 1 to 7.