Speech processing method, electronic device, storage medium, program product, and system
By performing initial speech processing on the device and sending it to the server for secondary processing in conjunction with historical information when the confidence level is low, the problems of long response time of large language models and lack of historical information in the cloud are solved, achieving more accurate and efficient speech response.
Patent Information
- Application Number
- PCT/CN2025/088105
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-05
- Filing Date
- 2025-04-09
- Publication Date
- 2026-01-08
AI Technical Summary
Large language models have slow inference speed in voice interaction systems, resulting in excessively long response times. Furthermore, in continuous interaction scenarios, the lack of historical information when the cloud intervenes leads to response errors.
The target speech is initially processed on the device side to obtain recognition results and confidence levels. When the confidence level is low, the speech and feature information are sent to the server side for further processing, and historical processing information is provided as an aid. The server side combines the historical information to generate an accurate response.
It improves the accuracy and efficiency of voice response, reduces data transmission volume, shortens response time, and avoids errors caused by a lack of historical information in the cloud.
Smart Images

Figure CN2025088105_08012026_PF_FP_ABST
Abstract
Description
Speech processing method, electronic device, storage medium, program product and system
[0001] The present application claims priority to Chinese Patent Application No. 2024109041354, the contents of which are incorporated herein by reference.
TECHNICAL FIELD
[0002] The present application relates to the field of speech processing, in particular to a speech processing method, an electronic device, a storage medium, a program product and a system.
BACKGROUND
[0003] In the current speech interaction system, a large language model is an important part of realizing language understanding and generation. However, the inference speed of the large language model is slow, which leads to a long response time of speech interaction. In order to solve this problem, the large language model is usually placed in the cloud to improve its inference efficiency, or local processing and cloud processing are used together to realize the acceleration of speech interaction. When the local processing finds that the response to the target speech is unreasonable, the target speech is handed over to the cloud for response processing. However, if the cloud needs to intervene in the processing in the middle of continuous interaction, the cloud only relies on the target speech for response processing, and due to the lack of historical information, speech intent recognition errors may occur, resulting in response errors.
SUMMARY
[0004] The main purpose of the present application is to provide a speech processing method, an electronic device, a storage medium, a program product and a system, which can improve the accuracy of speech response.
[0005] To solve the above technical problems, the first technical solution adopted by the present application is to provide a speech processing method. The method is applied to a device end, and the device end is in communication connection with a server end. The method comprises: obtaining a target speech; identifying the target speech, and obtaining a first response result and a confidence degree of the first response result according to the identification result; in response to the confidence degree of the first response result being less than a preset threshold, sending at least one of the target speech and feature information corresponding to the target speech and historical processing information to the server end for processing to obtain a second response result of the target speech; wherein the historical processing information comprises historical identification results and historical response results corresponding to historical speeches; and responding to the target speech according to the second response result.
[0006] To solve the above technical problems, a second technical solution adopted by the present application is to provide a voice processing method. The method is applied to a server end, and the server end is in communication connection with a device end. The method comprises receiving at least one of a target voice and feature information corresponding to the target voice from the device end and historical processing information; wherein the historical processing information comprises a historical recognition result corresponding to a historical voice and a historical response result in a voice processing process of the device end; processing the at least one of the target voice and the feature information corresponding to the target voice and the historical processing information to obtain a second response result of the target voice; and sending the second response result to the device end to enable the device end to respond to the target voice based on the second response result.
[0007] To solve the above technical problems, a third technical solution adopted by the present application is to provide an electronic device. The electronic device comprises a memory and a processor, the memory is used to store program data, and the program data can be executed by the processor to implement the method as described in the first technical solution or the second technical solution.
[0008] To solve the above technical problems, a fourth technical solution adopted by the present application is to provide a computer readable storage medium. The computer readable storage medium stores program data, and the program data can be executed by a processor to implement the method as described in the first technical solution and / or the second technical solution.
[0009] To solve the above technical problems, a fifth technical solution adopted by the present application is to provide a computer program product. The computer program product stores program data, and the program data can be executed by a processor to implement the method as described in the first technical solution and / or the second technical solution.
[0010] To solve the above technical problems, a sixth technical solution adopted by the present application is to provide a voice processing system, which comprises a device end and a server end, and the device end is in communication connection with the server end to implement the method as described in the first technical solution and / or the second technical solution.
[0011] The present application has the beneficial effects that after obtaining the target voice, the target voice is processed at the device end to obtain a recognition result of the target voice by the device end, a first response result responding to the recognition result, and a confidence degree of the first response result. When the confidence degree of the first response result is less than a preset threshold, it is indicated that the response of the processing system of the device end to the target voice can be inaccurate, and therefore the target voice and / or the feature information corresponding to the target voice are sent to the server end for second response processing. In order to avoid the response error of the processing system of the server end due to the lack of historical information, the device end also sends the historical processing information to the server end as auxiliary information for processing the target voice by the server end, and the historical processing information can make the response of the server end to the target voice more accurate. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor based on these drawings. Among them:
[0013] Fig. 1 is a schematic diagram of a voice processing flow;
[0014] Fig. 2 is a schematic diagram of a first embodiment of the voice processing method of the present application;
[0015] Fig. 3 is a schematic diagram of a second embodiment of the voice processing method of the present application;
[0016] Fig. 4 is a schematic diagram of a third embodiment of the voice processing method of the present application;
[0017] Fig. 5 is a schematic diagram of a fourth embodiment of the voice processing method of the present application;
[0018] Fig. 6 is a schematic diagram of a specific embodiment of the voice processing method of the present application;
[0019] Fig. 7 is a schematic diagram of the structure of an embodiment of the electronic device of the present application;
[0020] Fig. 8 is a schematic diagram of the structure of an embodiment of the computer readable storage medium of the present application;
[0021] Fig. 9 is a schematic diagram of the structure of an embodiment of the computer program product of the present application;
[0022] Fig. 10 is a schematic diagram of the structure of an embodiment of the voice processing system of the present application.
DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0024] The terms "first", "second", and the like in the specification and claims of this application are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. Moreover, the terms "comprising", "having", "including", and the like, are used in their open-ended, non-limiting sense, and do not exclude additional steps, elements, features, or the like. The use of the negative "not" in the claims is intended to be a specific limitation, and does not include the meaning of "not comprising" or "not including."
[0025] Reference herein to "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment can be included in at least one embodiment of the application. The appearances of the phrase "in an embodiment" in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily all directed to the same embodiment, or to a single alternative embodiment. It is expressly understood that any of the embodiments described herein can be combined with any of the other embodiments unless specifically noted otherwise.
[0026] In the general voice interaction scenario, the larger the language model used, the longer the corresponding response time. Common large language models generally have a relatively obvious delay in voice response when used. However, in voice interaction, if the response interval exceeds a certain time, for example, 0.2 seconds, the dialogue process embodied is not very natural. Therefore, in order to reduce the response time as much as possible in voice interaction, the inference speed of the large language model can be accelerated, a classifier and a traditional NLP link can be used for shunting, and an offline processing system can be set on a local device to respond to voice using the offline processing system, and when the offline processing system cannot meet the processing demand, a cloud voice processing system is used to process the voice. As shown in FIG. 1, FIG. 1 is a voice processing flow diagram.
[0027] In order to solve the problem of response error caused by the lack of historical information in the cloud in the middle of continuous interaction, the following embodiments are proposed in the application.
[0028] Referring to FIG. 2, FIG. 2 is a flowchart of a first embodiment of the voice processing method. The method is applied to a device end, which is in communication connection with a server end. The device end is provided with a local voice processing system. The local voice processing system can include an ASR processing module, an NLP processing module, a decision maker module, and a cache module. The voice is processed by the ASR processing module and the NLP processing module to obtain a response result to the voice and a confidence degree corresponding to the response result. The decision maker module is used to judge the confidence degree of the response result to determine whether the current voice needs the intervention processing of the server end. The cache module is used to cache data that needs to be uploaded to the server end for processing. The data stored in the cache module can include voice data and / or processed voice feature information data. The method includes but is not limited to the following steps.
[0029] S11: Obtain a target voice.
[0030] The target voice is the sound collected by the sound collection system of the device end that needs to be processed. The judgment of the device end on the target voice is often realized through a voice wake-up word. When the device end recognizes the voice wake-up word, the sound collected in the subsequent preset time period will be regarded as the sound that needs to be processed, and the sound needs to be responded.
[0031] S12: Identify the target voice, and obtain a first response result and a confidence degree of the first response result according to the identification result.
[0032] S13: In response to the confidence degree of the first response result being less than a preset threshold, send at least one of the target voice and target voice corresponding feature information and historical processing information to the server end for processing to obtain a second response result of the target voice.
[0033] When the confidence degree of the first response result is less than the preset threshold, it indicates that the voice processing system of the device end may not cover the target voice, and the response result of the device end to the target voice may not be accurate, so the intervention of the server end is needed. Because the voice processing system of the server end is networked, its coverage range for voice is larger than that of the offline voice processing system of the device end. The data sent to the server end can be the target voice, can be the feature information processed from the target voice, or can be both. Since the feature information is the information obtained by extracting the features of the target voice, the data amount of the feature information is reduced compared with that of the target voice data, so sending only the feature information can reduce the data transmission amount between the device end and the server end, reduce the data transmission time, and thus reduce the voice response time.
[0034] And, in order to avoid the response error caused by the lack of historical information of the processing system of the server side, the device side also sends historical processing information to the server side as auxiliary information for the server side to process the target voice, so as to help the server to process the target voice. The historical processing information includes historical recognition results corresponding to historical voices and historical response results.
[0035] S14: responding to the target voice according to the second response result.
[0036] After the server side processes the target voice, the second response result is obtained, and the server side sends the second response result to the device side. The device side generates corresponding execution instructions according to the second response result. Alternatively, the server side generates corresponding execution instructions according to the second response result, and then sends the execution instructions to the device side.
[0037] Responding to the target voice according to the second response result includes performing corresponding operations according to the execution instructions corresponding to the second response result, and further includes responding to the voice according to the second response result.
[0038] In this embodiment, after obtaining the target voice, the device side processes the target voice, and obtains the recognition result of the target voice by the device side, the first response result responding to the recognition result, and the confidence of the first response result. When the confidence of the first response result is less than a preset threshold, it indicates that the response of the local processing system of the device side to the target voice may be inaccurate, and therefore the target voice and / or the feature information corresponding to the target voice are sent to the server side for second response processing. In order to avoid the response error caused by the lack of historical information of the processing system of the server side, the device side also sends historical processing information to the server side as auxiliary information for the server side to process the target voice, so as to help the server to process the target voice. The historical processing information includes historical recognition results corresponding to historical voices and historical response results.
[0039] In an embodiment, the voice processing flow of the device side is started in response to a voice wake-up word, and the historical processing information includes historical recognition results and historical response results in a time period from the current time to the response to the last voice wake-up word. All voice processing processes from the last wake-up word to the current time are regarded as a complete scene, and the processing information in this process is regarded as the historical processing information of the current processing flow.
[0040] In an embodiment, the historical processing information includes historical recognition results and historical response results of all voice processing processes in a preset number of times or a preset time period before the target voice processing process. The processing information in a time period or at least one voice processing process before the current voice processing process is regarded as the historical processing information of the current voice processing process.
[0041] In an embodiment, the two can be combined, and the historical processing information can include historical recognition results and historical response results of all voice processing processes before a preset number of times or a preset time period before the target voice processing process within a time period from the current time to the time when the last voice wake-up word is responded to.
[0042] In an embodiment, in response to the confidence of the first response result being greater than or equal to a preset threshold, the target voice is responded to according to the first response result. After obtaining the first response result at the device end, a corresponding execution instruction is generated according to the first response result. Responding to the target voice according to the first response result includes performing a corresponding operation according to the execution instruction corresponding to the first response result. Further, it can also include voice reply according to the first response result.
[0043] Referring to FIG. 3, FIG. 3 is a flowchart of a second embodiment of the voice processing method of the present application. The method is applied to the server end, which is in communication connection with the device end. The server end of the method is provided with a server voice processing system. The server voice processing system can include a voice recognition module and a large language model module. Further, it can also include a classification module and a natural language processing module. After the voice passes through the voice recognition module, it then passes through the large language model module or the natural language processing module to obtain a response result. The classification module is used to judge whether to pass through the large language model module or the natural language processing module. The method includes but is not limited to the following steps.
[0044] S21: receiving at least one of a target voice and feature information corresponding to the target voice from the device end and historical processing information.
[0045] The target voice is the sound collected by the sound collection system of the device end that needs to be processed. The judgment of the device end on the target voice is often realized through a voice wake-up word. When the device end recognizes the voice wake-up word, the sound collected within a preset time period thereafter will be regarded as the sound that needs to be processed, and the sound needs to be responded to.
[0046] The server end can also receive feature information corresponding to the target voice. The feature information is information obtained by feature extraction on the target voice, and the data amount thereof is reduced compared to the data amount of the target voice. When the device end only uploads the feature information corresponding to the target voice, compared to uploading the voice data, the data transmission amount from the device end to the server end can be reduced, thereby further reducing the response time of voice interaction.
[0047] For example, the speech recognition module can include an ASR processing module. Taking Fbank features acceptable to ASR as an example, the data quantity of 1s speech originally sampled at a sampling rate of 16K is 16000*32bit. According to the general setting of 20ms frame shift and 50ms frame, if 80-dimensional Fbank features are used, the data of 1s is compressed to 50 frames*80*32bit. If FP16 quantization is used, the data quantity is further compressed to 1 / 8 of the original data quantity. Or discrete speech tokens can also be used. Taking Encodec as an example, about 8 codebooks with a size of 1024 are used, and excellent compression efficiency is achieved. Then the data quantity of the encoder using Encodec is only 50 (frames)*8*10 (codable 1024) bits.
[0048] The historical processing information includes historical recognition results and historical response results of the device end in a speech processing process.
[0049] In an embodiment, the speech processing flow of the device end is started in response to a speech wake-up word, and the historical processing information includes historical recognition results and historical response results in a time period from the current time to responding to a previous speech wake-up word. All speech processing processes from the previous wake-up word to the current time are regarded as a complete scene, and the processing information in this process is regarded as the historical processing information of the current processing flow.
[0050] In an embodiment, the historical processing information includes historical recognition results and historical response results of all speech processing processes in a preset number of times or a preset time period before the target speech processing process. The processing information in a period of time or at least one speech processing process before the current speech processing process is regarded as the historical processing information of the current speech processing process.
[0051] In an embodiment, the above two can be combined, and the historical processing information can include historical recognition results and historical response results of all speech processing processes in a preset number of times or a preset time period before the target speech processing process in a time period from the current time to responding to a previous speech wake-up word.
[0052] S22: processing at least one of the target speech and the feature information corresponding to the target speech and the historical processing information to obtain a second response result of the target speech.
[0053] S23: sending the second response result to the device end to make the device end respond to the target speech based on the second response result.
[0054] The server side obtains a second response result according to the received target voice and / or the feature information corresponding to the target voice, and further in combination with historical processing information, and sends the generated second response result to the device side. Or, the server side generates corresponding execution instructions according to the second response result, and sends the second response result and the corresponding execution instructions to the device side.
[0055] Referring to FIG. 4, FIG. 4 is a flowchart of a third embodiment of the voice processing method of the present application. The present method is a further extension of step S22, which includes but is not limited to the following steps.
[0056] S31: input at least one of the target voice and the feature information corresponding to the target voice into a voice recognition module to obtain a recognition result.
[0057] S32: input the recognition result and historical processing information into a large language model module to obtain a second response result of the target voice.
[0058] When the server side includes a voice recognition module (which can include an ASR processing module) and a large language model module, the target voice or the feature information corresponding to the target voice can obtain a corresponding response result after sequentially passing through the voice recognition module and the large language model module.
[0059] The voice recognition model processes the target voice or the feature information corresponding to the target voice to obtain a recognition result. If the data uploaded to the server side includes both feature information and voice data, the recognition result is obtained through the feature information, and the recognition result obtained from the target voice can also be confirmed again, so as to obtain the final output recognition result. The large language model receives the recognition result input from the voice recognition module, and processes it with historical processing information as prompt information to obtain a second response result.
[0060] Referring to FIG. 5, FIG. 5 is a flowchart of a fourth embodiment of the voice processing method of the present application. The present method is a further extension of the above-mentioned embodiments, which includes but is not limited to the following steps.
[0061] S41: input the recognition result into a classification module.
[0062] When the server side further includes a classification module and a natural language processing module (which can include an NLP processing module), after obtaining the recognition result, it is input into the classification module for judgment, to judge the intent type of the recognition result, so as to judge whether to input into the natural language processing module or the large language model module according to the intent type. The classification module and the natural language processing module are used to distribute the voice processing, so as to reduce the number of data processing in the large language model module, thereby reducing the response time.
[0063] S42: in response to the type of the recognition result being a single intent, inputting the recognition result to a natural language processing module to obtain a second response result of the target voice in combination with historical processing information; and in response to the type of the recognition result being a multi-intent, inputting the recognition result to a large language model module to obtain a second response result of the target voice in combination with historical processing information.
[0064] If the response result is a single intent, it is processed by the natural language processing module, and if it is a multi-intent, it is processed by the large language model module in combination with historical processing information.
[0065] If processed by the natural language processing module, the recognition result and the historical processing information are input into the natural language processing module to obtain the second response result. The historical processing information will serve as the context information of the recognition result to assist the natural language processing module in processing it.
[0066] The voice processing method of the present application will be described in more detail below with reference to a specific embodiment. Referring to FIG. 6, FIG. 6 is a flowchart of a specific embodiment of the voice processing method of the present application.
[0067] As can be seen, the local loop of the device end will dominate most of the voice interaction, and only when the decision maker determines that the cloud end needs to be involved will the local voice data be uploaded. The cloud voice processing system is only a secondary voice processing system, and when the local decision maker determines that secondary intervention is needed, the data in the cache will be uploaded.
[0068] If the confidence of the current response result is low, the current target voice or the corresponding feature information will be recorded in the cache for uploading to the cloud for a second processing. The device end also needs to transmit the local ASR processing result (the recognition result described above), the NLP result (the response result described above), and even the local instruction execution result (the execution instruction generated based on the response result described above) to the cloud, and the cloud combines these information. For traditional NLP, the context information is composed (this route is not shown), and for the large model link, the prompt information is composed.
[0069] As shown in FIG. 7, FIG. 7 is a structural schematic diagram of an embodiment of the electronic device of the present application.
[0070] The electronic device includes a processor 110 and a memory 120.
[0071] The processor 110 controls the operation of the electronic device, and can also be referred to as a CPU (Central Processing Unit). The processor 110 can be an integrated circuit chip having a processing capability of a signal sequence. The processor 110 can also be a general-purpose processor, a digital signal sequence processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like.
[0072] The memory 120 stores instructions and program data required for the operation of the processor 110. The memory 120 can include a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic or optical disk, or other media that can store program instructions, or can be a server that stores program instructions and sends the stored program instructions to other devices for execution, or can also execute the stored program instructions.
[0073] The processor 110 is used to respond to the results to implement any embodiment and possible combination of the speech processing method applied to the device end in the preceding application, and to implement the method provided by any embodiment and possible combination of the speech processing method applied to the server end in the preceding application.
[0074] As shown in FIG. 8, FIG. 8 is a structural schematic diagram of an embodiment of the computer readable storage medium of the present application.
[0075] An embodiment of the readable storage medium of the present application includes a memory 210, which stores program data that, when executed, implements the method provided by any embodiment and possible combination of the speech processing method.
[0076] The memory 210 can include a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic or optical disk, or other media that can store program instructions, or can be a server that stores program instructions and sends the stored program instructions to other devices for execution, or can also execute the stored program instructions.
[0077] As shown in FIG. 9, FIG. 9 is a structural schematic diagram of an embodiment of the computer program product of the present application.
[0078] An embodiment of the computer program product provided in the present application comprises a memory 310, and the memory 310 stores program data, and the program data is executed to implement the method provided by any one of the embodiments of the voice processing method and possible combinations.
[0079] The memory 310 can comprise a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and the like, which can store program instructions, or can be a server storing the program instructions, which can send the stored program instructions to other devices for running, or can run the stored program instructions.
[0080] As shown in FIG. 10, FIG. 10 is a structural schematic diagram of an embodiment of the voice processing system provided in the present application.
[0081] An embodiment of the voice processing system provided in the present application comprises a device end 410 and a server end 420. The device end 410 can comprise an electronic device implementing the method applied to the device end in the embodiment of the electronic device provided in the present application, and the server end 420 can comprise an electronic device implementing the method applied to the server end in the embodiment of the electronic device provided in the present application. The device end 410 and the server end 420 are in communication connection to implement the method provided by any one of the aforementioned embodiments of the voice processing method and possible combinations.
[0082] In summary, after obtaining the target voice, the target voice is processed at the device end to obtain the recognition result of the target voice at the device end, the first response result responding to the recognition result, and the confidence degree of the first response result. When the confidence degree of the first response result is less than a preset threshold, it is indicated that the response of the processing system at the device end to the target voice can be inaccurate, and therefore the target voice and / or the feature information corresponding to the target voice are sent to the server end for second response processing. In order to avoid the response error of the processing system at the server end due to the lack of historical information, the device end further sends the historical processing information to the server end as auxiliary information for processing the target voice at the server end, and the historical processing information can make the response of the server end to the target voice more accurate.
[0083] In the several embodiments provided in the present application, it should be understood that the disclosed method and device can be implemented in other ways. For example, the device embodiments described above are only schematic. The division of the modules or units is only a logical function division. There can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0084] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e. may be located in one place, or may be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.
[0085] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0086] The integrated unit in the above other embodiments, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical scheme of the present application, essentially or the part that contributes to the prior art, or all or part of the technical scheme can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0087] The above is only an embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A voice processing method, characterized by, Applied to a device end, the device end is in communication connection with a server end, and the method comprises: Acquiring a target voice; Recognizing the target voice, and obtaining a first response result and a confidence degree of the first response result according to a recognition result; In response to the confidence degree of the first response result being less than a preset threshold, sending at least one of the target voice and feature information corresponding to the target voice and historical processing information to the server end for processing to obtain a second response result of the target voice; wherein the historical processing information comprises historical recognition results and historical response results corresponding to historical voices; Responding to the target voice according to the second response result.
2. The method of claim 1, wherein, The voice processing flow of the device end is started in response to a voice wake-up word, and the historical processing information comprises the historical recognition results and the historical response results in a time period from the current to responding to a last voice wake-up word.
3. The method according to any one of claims 1-2, characterized in that, The historical processing information comprises the historical recognition results and the historical response results of all voice processing processes in a preset number of times or a preset time period before a target voice processing process.
4. The method according to any one of claims 1-3, characterized in that, The method further comprises: In response to the confidence degree of the first response result being greater than or equal to the preset threshold, responding to the target voice according to the first response result.
5. A voice processing method characterized by, Applied to a server end, the server end is in communication connection with a device end, and the method comprises: Receiving at least one of a target voice and feature information corresponding to the target voice and historical processing information from the device end; wherein the historical processing information comprises historical recognition results and historical response results corresponding to historical voices in a voice processing process of the device end; Processing at least one of the target voice and the feature information corresponding to the target voice and the historical processing information to obtain a second response result of the target voice; Sending the second response result to the device end to make the device end respond to the target voice based on the second response result.
6. The method of claim 5, wherein, The server end comprises a voice recognition module and a large language model module, and the processing of at least one of the target voice and the feature information corresponding to the target voice and the historical processing information to obtain the second response result of the target voice comprises: Inputting at least one of the target voice and the feature information corresponding to the target voice into the voice recognition module to obtain a recognition result; Inputting the recognition result and the historical processing information into the large language model module to obtain the second response result of the target voice.
7. The method of claim 6, wherein, The server end further comprises a classification module and a natural language processing module, and after obtaining the recognition result, the method comprises: Inputting the recognition result into the classification module; In response to the type of the recognition result being a single intent, inputting the recognition result into the natural voice processing module to obtain the second response result of the target voice in combination with the historical processing information; in response to the type of the recognition result being a multiple intent, inputting the recognition result into the large language model module to obtain the second response result of the target voice in combination with the historical processing information.
8. An electronic device, comprising: A computer program product comprising a memory for storing program data and a processor capable of executing the program data to implement the method of any of claims 1-4 or any of claims 5-7.
9. A computer-readable storage medium, characterized in that, A computer program product comprising a memory for storing program data and a processor capable of executing the program data to implement the method of any of claims 1-7.
10. A computer program product, characterised in that, A computer program product comprising a memory for storing program data and a processor capable of executing the program data to implement the method of any of claims 1-7.
11. A speech processing system characterized by A computer program product comprising a memory for storing program data and a processor capable of executing the program data to implement the method of any of claims 1-7.
Citation Information
Patent Citations
Voice recognition system and method used for mobile equipment
CN102543071A
Intelligent session method and device
CN112069830A
Speech recognition processing method and processing system thereof, vehicle and readable storage medium
CN115410578A
Voice processing method, electronic equipment, storage medium, program product and system
CN118865956A
Electronic device and controlling the electronic device
US20210166678A1