Voice response method and device
By acquiring and splicing voice data text during a voice call and combining it with the feature information of historical cached text, the problem of long speech segmentation errors under short speech capabilities is solved, achieving more accurate voice response.
Patent Information
- Application Number
- CN202511050021.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-09-16
AI Technical Summary
In the prior art, the ASR function that uses short speech capabilities is prone to speech segmentation errors when processing long speech, affecting the accuracy of speech response.
By acquiring voice data for text conversion, combining the feature information in the historical cached text, splicing and processing to obtain a complete text, and sending operation instructions based on the text.
The accuracy of voice responses is improved and the misjudgment of operation instructions caused by excessive segmentation during text conversion is reduced.
Smart Images

Figure CN120656454A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of electronic technology, and specifically relates to a voice response method and device thereof. Background Art
[0002] With the development of mobile communications, voice calls via electronic devices are becoming increasingly common. For example, users may dial customer service hotlines via mobile phones to seek assistance. However, to reduce labor costs, many companies involved in voice calls often use voice robots as the first access point during voice calls. This cumbersome process requires users to repeatedly interact with the robot to access the desired service flow or transfer to a human operator, reducing the user experience.
[0003] In related technologies, AI technology is used to automatically replace users to interact with voice robots in order to reduce the tedious interactions between users and voice robots. This technology often uses automatic speech recognition (ASR). Currently, the ASR function in the industry has two main recognition capabilities: short speech capability and long speech capability. Among them, the short speech capability supports up to 60 seconds of voice input and has strong real-time performance. Therefore, it is widely used in scenarios of voice interaction with robots. However, in this scenario, due to the use of the ASR function with short speech capability, a longer robot voice may be incorrectly segmented into multiple text segments, causing ASR recognition errors, which ultimately affects the accuracy of the voice response. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a voice response method and apparatus thereof, which can improve the accuracy of voice response.
[0005] In a first aspect, an embodiment of the present application provides a voice response method, which is executed by a first electronic device, and the method includes: during a voice call, obtaining first voice data, and converting the first voice data into text to obtain a first text; when there is a target cache text whose characteristic information meets the first condition in the historical cache text, splicing the first text and the target cache text to obtain a second text, and the historical cache text is the text corresponding to the voice data received earlier than the first voice data during the voice call; based on the above second text, sending corresponding operation instructions to the second electronic device.
[0006] In a second aspect, an embodiment of the present application provides a device for voice response, which includes a receiving module, a recognition module, a splicing module, and a sending module. The receiving module is used to obtain first voice data during a voice call; the recognition module is used to convert the first voice data into text to obtain a first text; the splicing module is used to splice the first text with the target cache text to obtain a second text when there is a target cache text whose characteristic information meets the first condition in the historical cache text, and the historical cache text is the text corresponding to the voice data received earlier than the first voice data during the voice call; the sending module is used to send corresponding operation instructions to the second electronic device based on the second text.
[0007] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
[0008] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0009] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method described in the first aspect.
[0010] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the method described in the first aspect.
[0011] In an embodiment of the present application, during a voice call, a first electronic device first obtains first voice data and converts the first voice data into text to obtain a first text. Then, when there is a target cache text whose characteristic information satisfies a first condition in the historical cache text, the first text is spliced with the target cache text to obtain a second text. The historical cache text is the text corresponding to the voice data received earlier than the first voice data during the voice call. Finally, based on the second text, a corresponding operation instruction is sent to the second electronic device. In this solution, the first electronic device can select the target cache text that meets the first condition from the historical cache text based on the characteristic information to splice with the first text, so that the semantics of the spliced second text are more complete, so as to facilitate the recognition of the user intention expressed by the voice data; and the first electronic device can send the corresponding operation instruction to the second electronic device based on the spliced second text, that is, the operation performed by the second electronic device is determined based on the context information of the first voice data. Therefore, the second electronic device can make an accurate voice response, reduce the misjudgment of the operation instruction caused by excessive sentence segmentation during the text conversion process during the voice response, and improve the accuracy of the voice response. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is one of the flowcharts of the voice response method provided in some embodiments of the present application;
[0013] Figure 2 This is one of the function activation diagrams described in the voice response method provided in some embodiments of the present application;
[0014] Figure 3 This is one of the interface content schematic diagrams of the voice response method provided in some embodiments of the present application;
[0015] Figure 4 This is one of the schematic diagrams of time threshold determination of the voice response method provided in some embodiments of the present application;
[0016] Figure 5 This is a second flow chart of the voice response method provided in some embodiments of the present application;
[0017] Figure 6 is a schematic diagram of a voice response device provided in some embodiments of the present application;
[0018] Figure 7 is a schematic diagram of an electronic device provided by some embodiments of the present application;
[0019] Figure 8 This is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application. DETAILED DESCRIPTION
[0020] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0021] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0022] The terms "at least one" and "at least one of" in the specification and claims of this application refer to any one, any two, or a combination of more than two of the objects included. For example, at least one of a, b, and c can be represented by: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two" means two or more, and its meaning is similar to "at least one".
[0023] The voice response method provided in the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0024] For example, consider a scenario where a user dials a customer service line via their mobile phone. When a user dials a customer service line through their phone's calling app and speaks to a customer service robot, if the user's phone's ASR function is enabled, the phone automatically captures the voice broadcast by the customer service robot and uploads it to the server. The server then uses the ASR technology in its short voice function to perform semantic segmentation and command intent recognition on the broadcasted voice in real time. Based on the command intent recognition results, the server sends control commands to the phone to control the phone to perform relevant operations, such as automatically pressing the phone's * key or automatically controlling the phone to play a message in response to the customer service robot. However, if the voice broadcast by the customer service robot is long and continuous and exceeds the maximum voice input duration of the short voice function, the server's semantic segmentation using the ASR technology in the short voice function may incorrectly segment a long speech segment. Consequently, when the server subsequently uses reasoning to recognize the command intent, the server will fail to recognize the command intent due to a lack of context, ultimately resulting in an incorrect voice response.
[0025] In an embodiment of the present application, during a voice call, the first electronic device first obtains the first voice data and converts the first voice data into text to obtain the first text. Then, when there is a target cache text whose characteristic information satisfies the first condition in the historical cache text, the first text and the target cache text are spliced together to obtain a second text. The historical cache text is the text corresponding to the voice data received earlier than the first voice data during the voice call. Finally, based on the second text, the corresponding operation instruction is sent to the second electronic device. In this solution, the first electronic device can select the target cache text that meets the first condition from the historical cache text based on the characteristic information to splice it with the first text, and send the corresponding operation instruction to the second electronic device based on the spliced second text. Therefore, the second electronic device can make an accurate voice response based on the operation instruction, reducing the misjudgment of the operation instruction caused by excessive sentence segmentation during the text conversion process during the voice response, and improving the accuracy of the voice response.
[0026] The execution subject of the voice response method provided in this application may be a voice response device. For example, the voice response device may be an electronic device such as a mobile phone, a tablet computer, a laptop computer, a PDA, a wearable device, an in-vehicle electronic device, a server, or the like. Alternatively, the voice response device may be a functional module or a processing module in the electronic device, which is not limited in this application. In some embodiments of this application, the voice response method provided in the embodiments of this application is described by taking an electronic device as the execution subject to perform the voice response method as an example.
[0027] The present invention provides a method for voice response. Figure 1 FIG. 1 shows a flow chart of a voice response method provided by an embodiment of the present application. Figure 1As shown, the voice response method provided in the embodiment of the present application is executed by the first electronic device, and the method may include the following steps 201 to 203.
[0028] Step 201: During a voice call, a first electronic device obtains first voice data and converts the first voice data into text to obtain a first text.
[0029] In some embodiments of the present application, the above-mentioned voice call can be an online voice call or video call conducted through the Internet, such as a video conference, or a voice call or video call conducted through a mobile network, such as a customer service call, or other types of calls, which are not limited in this application.
[0030] In some embodiments of the present application, the first voice data is voice data collected by the second electronic device during a voice call and sent to the first electronic device. For example, the first voice data may be data obtained by the second electronic device recording voice data sent by the user to another electronic device during a voice call via the second electronic device, or the user's own voice data using the second electronic device, or other types of voice data.
[0031] In some embodiments of the present application, during a voice call, the first electronic device may receive voice data in real time and then obtain first voice data from the received voice data.
[0032] In some embodiments of the present application, the first electronic device segments the received voice data and obtains the first voice data from the multiple sub-voice data obtained after the segmentation.
[0033] In some embodiments of the present application, the above-mentioned voice data includes multiple segments of voice data. During a voice call, the first electronic device receives the multiple segments of voice data in real time, and divides the voice data in the received multiple segments of voice data where the silence duration between two adjacent segments exceeds a second preset duration to obtain multiple sub-voice data.
[0034] In some embodiments of the present application, the first voice data is a sub-voice data corresponding to the current moment among the multiple sub-voice data.
[0035] In some embodiments of the present application, the second preset duration may be a default duration of the electronic device system or a user-set duration. For example, the second preset duration may be 800ms. Of course, the second preset duration may also be other durations, which are not limited in the present application.
[0036] In some embodiments of the present application, the above-mentioned multiple voice data are voice data recorded in real time by the second electronic device and sent to the first electronic device.
[0037] In some embodiments of the present application, the second electronic device activates a recording function in response to the first input, and sends the recorded voice data to the first electronic device in real time.
[0038] In some embodiments of the present application, the first input is an input for selecting the user's intention during a voice call by the second electronic device.
[0039] For example, taking the scenario where a user dials a customer service number via a mobile phone as an example, Figure 2 In (a), the call app on the mobile phone can provide a "find manual customer service" button, and the user can click the button (i.e., the first input) to enable the automatic search for manual customer service function. Further, optionally, Figure 2 In (b), users can also select the purpose of transferring to a human operator (e.g., checking phone bills, changing plans, applying for broadband, etc.). Once the automatic call-to-human feature is enabled, the mobile app enters an automated process, using the customer service representative's voice as input and initiating the recording process. This allows for real-time monitoring, collection, and uploading of the customer service representative's audio content.
[0040] In some embodiments of the present application, the above-mentioned multiple segments of voice data are sorted in order of reception time.
[0041] In some embodiments of the present application, the two adjacent segments of voice data refer to two adjacent segments of voice data containing voice broadcast content among multiple segments of voice data.
[0042] In some embodiments of the present application, when the silence duration between two adjacent segments of voice data exceeds a second preset duration, the first electronic device generates a first identifier and a second identifier, and sends the first identifier to the second electronic device. The first identifier is used to instruct the second electronic device to display the first text, and the second identifier is used to indicate that the first text has ended, that is, the second identifier is used to indicate that the voice corresponding to the first text is a complete sentence, that is, the new text received after the first text belongs to two different sentences from the first text.
[0043] In some embodiments of the present application, when the first electronic device includes an ASR server and a server, the ASR server generates the second identifier, and the server receives the second identifier from the ASR server. It is understandable that since the second identifier is used to indicate that the speech corresponding to the first text is a complete sentence, the server can determine that the speech corresponding to the first text has ended by receiving the second identifier.
[0044] In some embodiments of the present application, the first electronic device can determine whether to segment the received multiple segments of voice data by judging whether the silence period between two adjacent segments of voice data exceeds a second preset period, and use the voice data at the current moment as the first voice data among the multiple sub-voice data obtained by segmentation to carry out the subsequent voice response process.
[0045] In this way, the first electronic device can reasonably segment the received voice data in real time, which helps to perform model reasoning based on the complete sentence when making a subsequent voice response, thereby improving the accuracy of the voice response.
[0046] In some embodiments of the present application, the first electronic device may be one electronic device, two electronic devices, or multiple electronic devices. For example, the first electronic device may include an ASR server and a server. The ASR server obtains the first voice data and converts it into a first text, and then sends the first text to the server for subsequent processing. In addition, the ASR server may also generate a first identifier and a second identifier, and send the first identifier to the second electronic device and the second identifier to the server.
[0047] In some embodiments of the present application, the first electronic device performs speech recognition through at least one of the following speech recognition methods when performing text conversion: ASR technology, a model based on a hidden Markov model, a recurrent neural network model based on deep learning, a long short-term memory network model, a transcription attention model, etc.
[0048] Exemplarily, taking scenario 1 as an example, the first electronic device includes an ASR server and a server, and the second electronic device is the mobile phone in scenario 1. After the ASR server receives the real-time recording data (i.e., multiple segments of voice data in this application) from the mobile phone, it first pre-processes the recording data, such as noise reduction, echo cancellation, etc., and then recognizes the recording data. The recognition process specifically includes: extracting voice features, converting voice features into phonemes or character probabilities through acoustic models, performing error correction and context understanding through language models, and converting probability sequences into text through decoders. Then, every time the ASR server recognizes text, it immediately uses the previously transcribed text in this round of recording data in combination with the currently transcribed text to issue a pre-request to the server, and sends it to the mobile phone at the same time. For example, for the same round of recording data, the ASR server transcribes it multiple times and sends the pre-request text to the server as follows:
[0049] First time: Please ask;
[0050] Second time: May I ask you;
[0051] Third time: May I ask if you;
[0052] Fourth time: May I ask if you still have any?
[0053] Fifth time: Do you have anything else?
[0054] Sixth: Do you have any other requests?
[0055] When the ASR server receives no customer service voice broadcast in the recording data after the sixth transcription, it will enable the silence detection function to detect the silence duration after the recording data. When the silence duration exceeds 800ms, the ASR server will use the silence duration position as the sentence break position to split the recording data. At this time, the system determines that the recording segment has ended and sends the end mark of this round of recording transcription to the mobile phone, and sends the pre-request end mark to the server. Figure 3 As shown, after receiving the recording transcription end mark, the mobile phone puts the entire content of the text obtained in this round into a dialogue bubble for display.
[0056] Step 202: When there is a target cache text whose characteristic information satisfies the first condition in the historical cache text, the first electronic device concatenates the first text with the target cache text to obtain a second text.
[0057] In some embodiments of the present application, the historical cached text is text corresponding to voice data received earlier than the first voice data during a voice call.
[0058] In some embodiments of the present application, the above-mentioned text corresponds to a corresponding voice data reception time, and each target cache sub-text also corresponds to a voice data reception time. The first electronic device obtains the above-mentioned second text after merging the first text and the target cache text end to end in the order of the voice data reception time corresponding to the first text and the target cache text.
[0059] In some embodiments of the present application, when there is no target cache text whose characteristic information meets the first condition in the historical cache text, the first electronic device directly uses the first text as the second text.
[0060] In some embodiments of the present application, the above-mentioned feature information is used to characterize the target intent information and voice response information of the historical cached text.
[0061] In some embodiments of the present application, the historical cached text includes one or more cached subtexts, and the voice response method provided by the present application further includes the following steps 204a and 204b:
[0062] Step 204a: Acquire feature information of each cached subtext from the cache area of the first electronic device, where the feature information includes target intent information and voice response information.
[0063] Step 204b: Based on the characteristic information of each cached subtext, determine a cached subtext whose characteristic information meets the first condition from the multiple cached subtexts to obtain the target cached text.
[0064] In some embodiments of the present application, the target intent information is used to characterize whether the semantic information of the cached sub-text is related to the target intent, and the voice response information is used to characterize whether the second electronic device responds to the semantic information corresponding to the cached sub-text.
[0065] In some embodiments of the present application, an intent identifier corresponding to each cached sub-text is stored in a cache area of the first electronic device, and the above-mentioned target intent information is indicated by the intent identifier, which includes a first intent identifier and a second intent identifier. The first intent identifier is used to indicate that the semantic information of the cached sub-text is related to the target intent, and the second intent identifier is used to indicate that the semantic information of the cached sub-text is not related to the target intent.
[0066] In some embodiments of the present application, the target intent includes at least one of the following: an intent determined based on a user input operation on the second electronic device; an intent to transfer to a human customer service representative. For example, the user intent may include at least one of the following: checking phone bills, checking insurance policy anomalies, checking data usage, or other user intents, which are not limited in the present embodiment.
[0067] In some embodiments of the present application, a user may select an intention to make a voice call on a second electronic device, and the second electronic device sends the intention information to the first electronic device based on the user's selection.
[0068] In some embodiments of the present application, one or more intent controls are displayed on the second electronic device, each intent control corresponds to a user intent, and the second electronic device determines the target intent in response to the user's input operation on at least one intent control.
[0069] In some embodiments of the present application, after determining the target intention, the second electronic device sends the target intention to the first electronic device.
[0070] In some embodiments of the present application, the second electronic device displays one or more intent controls while also displaying a control for finding human customer service. The second electronic device can respond to the user's input to the control for finding human customer service, collect the other party's voice data during the voice call, and send it to the first electronic device.
[0071] It is understandable that the intention corresponding to one or more intention controls already displayed on the second electronic device may not match the user's actual intention (i.e., the target intention). Therefore, the user can automatically enter the automatic search process for manual customer service by inputting the control for finding manual customer service.
[0072] Exemplarily, before caching each cached sub-text, the first electronic device first calls the large language model, performs semantic analysis on the cached sub-text, and determines whether the cached sub-text contains information related to the user's selected intention or information related to transfer to manual customer service, and based on the judgment result, outputs and caches the target intent information corresponding to the cached sub-text. For example, if the cached sub-text contains information related to the user's intention or has been connected to manual customer service, the first intent identifier is output; otherwise, the second intent identifier is output.
[0073] In some embodiments of the present application, a voice response identifier corresponding to each cached subtext is stored in a cache area of the first electronic device. The voice response information is indicated by the voice response identifier, which includes a first voice response identifier and a second voice response identifier. Before caching the cached subtext, if the first electronic device sends an operation instruction to the second electronic device for responding to the semantic information of the cached subtext, the first electronic device will set the voice response identifier corresponding to the cached subtext to the first voice response identifier; otherwise, the first electronic device will set the voice response identifier corresponding to the cached subtext to the second voice response identifier.
[0074] In some embodiments of the present application, the above-mentioned cached sub-text includes at least one of the following: one or more texts obtained by segmenting voice data received earlier than the first voice data; a text obtained by splicing multiple texts obtained by segmenting voice data received earlier than the first voice data.
[0075] In some embodiments of the present application, the cache area of the first electronic device includes a first cache and a second cache. The first cache includes the above-mentioned multiple cache subtexts and feature information corresponding to each cache subtext, and the second cache includes a pre-request cache.
[0076] In some embodiments of the present application, the pre-request cache is used to store in real time the semantic information corresponding to multiple pre-request texts of the cached sub-text corresponding to the current moment.
[0077] For example, in conjunction with Figure 3 The first electronic device performs large-model inference on the texts pre-requested multiple times in the same round of voice (i.e., the first voice data mentioned above) and caches the inference results; when the server receives the pre-request information for the first time in each round, for example, the server receives the pre-request text "Excuse me" for the first time, it initializes the pre-request cache and identifies that the round of voice has started by generating and storing an initialization identifier.
[0078] In some embodiments of the present application, the first electronic device can select a target cache text whose target intent information and voice response information meet a first condition from multiple cache sub-texts based on the target intent information and voice response information corresponding to each cache sub-text, so that subsequent voice response can be performed based on the second text obtained by splicing the target cache sub-text with the first text.
[0079] In this way, the probability of speech response errors caused by over-segmentation of speech data can be reduced.
[0080] Step 203: The first electronic device sends a corresponding operation instruction to the second electronic device based on the second text.
[0081] In some embodiments of the present application, the first electronic device performs semantic recognition on the second text and sends corresponding operation instructions to the second electronic device based on the voice information obtained by the recognition.
[0082] In some embodiments of the present application, after the first electronic device performs semantic recognition on the second text, it determines whether the second text is related to the target intent, and / or determines whether the second electronic device has been connected to manual service. When the first electronic device determines that the second text is related to the target intent, the intention identifier of the second text is set to the first intention identifier; when it is determined that the second text is not related to the target intent, the intention identifier of the second text is set to the second intention identifier.
[0083] In some embodiments of the present application, when the first electronic device determines that the second electronic device has connected to manual service, the first electronic device can send an indication message to the second electronic device, and the second electronic device can directly end the relevant process of the voice response method after receiving the indication message.
[0084] It can be understood that the first electronic device executes the voice response method in order to achieve the target intention or connect to manual customer service. When it is determined that the second electronic device has connected to manual customer service, the second electronic device does not need to continue to send the voice data to the first electronic device for processing. That is, the second electronic device does not need to automatically respond to the voice data with the assistance of the first electronic device, but the user directly conducts a manual conversation with the customer service.
[0085] In some embodiments of the present application, the above step 203 may include the following steps 203a and 203b:
[0086] Step 203a: When the silence duration after the first voice data is greater than or equal to the first preset duration, a first operation instruction is sent to the second electronic device, where the first operation instruction is used to instruct the second electronic device to respond to the semantic information corresponding to the second text.
[0087] Step 203b: When the silence duration is less than the first preset duration, a second operation instruction is sent to the second electronic device, where the second operation instruction is used to instruct the second electronic device not to perform an operation.
[0088] In some embodiments of the present application, when the silence duration after the first voice data reaches the above-mentioned first preset duration, the first electronic device obtains the second identifier.
[0089] In some embodiments of the present application, the first preset duration is the sum of the second preset duration and the time threshold.
[0090] In some embodiments of the present application, the above-mentioned time threshold may be a default value of the electronic device system or a user-set value. For example, the above-mentioned time threshold may be 500ms. Of course, the above-mentioned time threshold may also be other durations, which are not limited in the embodiments of the present application.
[0091] It should be noted that by setting the above-mentioned time threshold, it is possible to prevent the first electronic device from performing semantic analysis on the second text, resulting in incomplete semantics, but not receiving new voice data containing the voice broadcast content, causing the server to continue waiting. In addition, the customer service voice broadcast content will first broadcast the reason, and then broadcast a voice prompting the user to make an operation selection. Moreover, the pause time between the above two different content segments exceeds the second preset time length. Therefore, the server needs to splice the above two ASR segments together for joint processing when performing subsequent semantic reasoning to ensure semantic integrity.
[0092] In some embodiments of the present application, the first operating instruction is used to instruct the second electronic device to respond to the voice broadcast content, for example, instructing the second electronic device to press a specific digital button or instructing the second electronic device to automatically broadcast a voice message in response to the voice data corresponding to the second text. Of course, the first operating instruction can also instruct the second electronic device to perform other operations, which are not limited in the embodiments of the present application.
[0093] In some embodiments of the present application, the first electronic device determines the first operation instruction based on the target intention and semantic information of the second text.
[0094] In some embodiments of the present application, the first operation instruction is used to assist the second electronic device in responding to the monitored voice to connect to the human customer service corresponding to the target intent. It is understandable that the instructions for responding to the same monitored voice are usually diverse, but these instructions may not all correspond to the target intent. Since the first electronic device obtains the first operation instruction for responding to the monitored voice based on the target intent, the related operations subsequently performed by the second electronic device according to the first operation instruction can be adapted to the target intent.
[0095] For example, taking the scenario given above as an example, when the ASR server determines that the silence duration after the first voice data reaches 800ms (i.e., the first preset duration), the ASR server generates a second identifier and sends the second identifier to the server. When the server receives the second identifier, Figure 4 As shown in (b) in Figure 4 As shown in (a), if the server detects a new voice segment ASR2 within a time threshold of 500ms after the silence detection time after ASR1, it sends an instruction to the mobile phone to instruct the mobile phone not to perform any operation. If the server does not detect a new voice input within a time threshold of 500ms after the silence detection time after ASR1 (i.e., the voice input corresponding to the above second text), it obtains a voice response instruction based on the target intent and the semantic information of the above second text, and sends the voice response instruction to the mobile phone.
[0096] In some embodiments of the present application, if the first electronic device sends a first operation instruction to the second electronic device, the first electronic device uses the first voice response identifier as the voice response information of the second text and caches it; if the first electronic device sends a second operation instruction to the second electronic device, the first electronic device uses the second voice response identifier as the voice response information of the second text and caches it.
[0097] In some embodiments of the present application, the first electronic device sends an operation instruction, i.e., the above-mentioned first operation instruction, only when the silence period after the first voice data exceeds a first preset period, to instruct the electronic device to respond to the above-mentioned second text. Otherwise, the first electronic device sends an instruction to the second electronic device to instruct the second electronic device not to perform any operation.
[0098] In this way, the impact of excessive segmentation of voice data on the accuracy of voice response during the voice response process can be avoided.
[0099] In some embodiments of the present application, the above step 202 includes the following step 202a.
[0100] Step 202a: When the semantic information of the first cached subtext is irrelevant to the target intention and the second electronic device does not respond to the semantic information corresponding to the first cached subtext, the first text is concatenated with the first cached subtext to obtain a second text.
[0101] In some embodiments of the present application, the first cache subtext is one or more of the multiple cache subtexts.
[0102] In some embodiments of the present application, the first electronic device determines whether the semantic information of the first cached subtext is related to the target intent based on the intent identifier corresponding to the first cached subtext in the cache area.
[0103] In some embodiments of the present application, the first electronic device determines whether the second electronic device responds to the semantic information corresponding to the first cached subtext based on a voice response identifier corresponding to the first cached subtext in the cache area.
[0104] In some embodiments of the present application, the first electronic device performs head-to-tail splicing according to the chronological order of each text in the first cached sub-text and the voice monitoring time corresponding to the first text to obtain the second text.
[0105] For example, taking the first electronic device including an ASR server and a server as an example, after obtaining the first text, the server judges the multiple cached sub-texts in the cache area of the server one by one in chronological order, selects the first cached sub-text whose semantic information is irrelevant to the target intention and the second electronic device does not respond to the semantic information corresponding to the first cached sub-text, and splices it with the first text in chronological order.
[0106] In some embodiments of the present application, the first electronic device can concatenate the text in the historical cache text, in which the voice information is irrelevant to the target intention and is responded to by the second electronic device, with the first cache subtext and the first text to obtain the second text.
[0107] In this way, the first electronic device can subsequently send corresponding operation instructions to the second electronic device in combination with the above-mentioned first cached subtext. Since the obtained operation instructions take into account the contextual semantic information of the first voice data, the accuracy of the voice response is improved.
[0108] In some embodiments of the present application, before executing the above step 202a, the above step 202 further includes the following step 202b.
[0109] Step 202b: Determine whether there is a text in the plurality of cached subtexts whose similarity to the first text is greater than a first threshold. If not, execute step 202a. If so, execute subsequent operations by determining whether the second identifier of the first text is received.
[0110] In some embodiments of the present application, when there is a text in multiple cached sub-texts whose similarity with the first text is greater than a first threshold and the first electronic device obtains the second identifier of the first text, if the second electronic device does not respond to the semantic information corresponding to the text whose similarity with the first text is greater than the first threshold, the first electronic device sends an operation instruction to the second electronic device to respond to the semantic information of the first text.
[0111] It can be understood that after the first electronic device receives the first text, if there is a text in multiple cached sub-texts whose similarity with the first text is greater than a preset threshold, if the second electronic device does not respond to the text whose similarity with the first text is greater than the preset threshold, it can be considered that the first voice data corresponding to the first text is a repeated broadcast voice that has not been responded to. At this time, the first electronic device can send an operation instruction to the second electronic device to respond to the repeated broadcast voice in a timely manner.
[0112] Further optionally, in some embodiments of the present application, after the first electronic device sends an operation instruction to the second electronic device to respond to the semantic information of the first text, the first electronic device discards the first text, or the first electronic device deletes the texts in multiple cached sub-texts whose similarity with the first text is greater than a first threshold, and caches the first text.
[0113] Further optionally, in some embodiments of the present application, the electronic device caches feature information corresponding to the first text while caching the first text.
[0114] In some embodiments of the present application, the first electronic device acquiring the second identifier of the first text may be that the first electronic device generates the second identifier of the first text, or the first electronic device receives the second identifier of the first text.
[0115] In some embodiments of the present application, the second identifier is used to indicate that the speech corresponding to the first text is a complete sentence, that is, the new text received after the first text and the first text belong to two different sentences.
[0116] In some embodiments of the present application, when there is a text in multiple cached sub-texts whose similarity with the first text is greater than a first threshold and the first electronic device has not received the second identifier of the first text, the first electronic device caches the first text and caches the feature information corresponding to the text whose similarity with the first text is greater than the first threshold as the feature information of the first text.
[0117] In some embodiments of the present application, when the silence duration after the first voice data exceeds a second preset duration, the second electronic device sends a second identifier to the first electronic device.
[0118] Exemplarily, taking the first electronic device including an ASR server and a server as an example, the ASR server converts the first voice data into text and sends it to the server. After receiving the above-mentioned first text, the server first combines the bge-m3 vector model and the bge-reranker-v2-m3 rearrangement model to determine the similarity between the first text and each of the multiple cached sub-texts in the server cache area. Assuming that the maximum similarity is 1, when there is a cached sub-text with a similarity greater than 0.9 to the first text, if the second identifier is not obtained, the feature information corresponding to the cached sub-text is used as the feature information of the first text. Otherwise, if the server does not respond to the voice information of the cached sub-text with a similarity greater than 0.9 to the first text, the server directly caches the first text and the feature information corresponding to the first text, and deletes the cached sub-text with a similarity greater than 0.9 to the first text.
[0119] In some embodiments of the present application, the voice response method provided by the present application further includes the following step 205:
[0120] Step 205: Update the target cache text in the historical cache text to the second text, and cache the feature information of the second text.
[0121] In some embodiments of the present application, the first electronic device obtains feature information of the second text by performing semantic analysis on the second text and sending an instruction to the second electronic device based on the second text.
[0122] In some embodiments of the present application, the first electronic device deletes the original characteristic information of the target cached text while caching the characteristic information of the second text.
[0123] In some embodiments of the present application, the first electronic device can update the target cache text and its corresponding feature information in the historical cache text based on the second text after using the second text. In this way, the electronic device can process the first voice data obtained in the next round in a timely manner based on the latest data information in the cache.
[0124] In this way, the real-time performance and reliability of voice response are improved.
[0125] In the voice response method of the present application, during a voice call, a first electronic device first obtains first voice data and converts the first voice data into text to obtain a first text. Then, when there is a target cache text whose characteristic information meets the first condition in the historical cache text, the first text and the target cache text are spliced together to obtain a second text. The historical cache text is the text corresponding to the voice data received earlier than the first voice data during the voice call. Finally, based on the second text, a corresponding operation instruction is sent to the second electronic device. In this solution, the first electronic device can select the target cache text that meets the first condition from the historical cache text based on the characteristic information to splice it with the first text, and send the corresponding operation instruction to the second electronic device based on the spliced second text. Therefore, the second electronic device can make an accurate voice response based on the operation instruction, reducing the misjudgment of the operation instruction caused by excessive sentence segmentation during the text conversion process during the voice response process, and improving the accuracy of the voice response.
[0126] In some embodiments of the present application, the customer service response method provided by the present application is exemplified by taking the first electronic device including an ASR server and a server and the second electronic device being a mobile client as an example. Figure 5 As described above, after the user starts to dial the customer service phone number, the customer service response method of the present application may include the following steps 101 to 108.
[0127] Step 101: The mobile client enables a voice response function based on the user's input intention information;
[0128] Step 102: The mobile client uploads the current recording data of the customer service voice to the ASR server in real time;
[0129] Step 103: The ASR server converts the current voice data uploaded by the mobile client into the current ASR text in real time and sends a pre-request to the server;
[0130] Step 104: The server determines whether to splice the current ASR text based on the historical ASR cache in the cache area, and splices or not splices the current ASR text based on the determination result to obtain a processed ASR text. The server also determines whether there is a historical ASR similar to the current ASR text in the cache area. If so, the process proceeds to step 107; if not, the process proceeds to step 105.
[0131] Step 105: The server uses a large model algorithm to perform semantic recognition on the processed ASR text and determines whether to transfer the call to a human operator. If the call is not transferred to a human operator, the server obtains a voice response result based on the semantic recognition result. If the call is transferred to a human operator, the server sends an indication message to the mobile phone to indicate that the call has been connected to a human operator.
[0132] Step 106: The server determines whether the silence duration after the current recording data exceeds the preset duration of 1300ms. If it is determined to be longer than 1300ms, the server proceeds to step 107. Otherwise, the server returns to step 102 and sends an operation instruction to the mobile client that does not perform any operation.
[0133] Step 107: Upon receiving the pre-request end identifier, the server sends an operation instruction corresponding to the voice response result to the mobile client, and caches the processed ASR text and its feature information;
[0134] Step 108: The mobile client performs corresponding operations according to the operation instructions or instruction information sent by the server. If the mobile client receives an instruction information indicating that the manual customer service has been connected, the process ends. Otherwise, after executing the response operation, it returns to step 101.
[0135] In this way, since the above-mentioned voice response method can splice the current ASR text based on historical semantic judgment, and judge whether to continue splicing the next round of voice data based on the spliced text combined with the preset time length, and send the corresponding operation instruction or indication information to the second electronic device based on the above-mentioned judgment result, the second electronic device can make an accurate voice response based on the operation instruction, reducing the misjudgment of the operation instruction caused by excessive sentence segmentation in the text conversion process during the voice response process, and improving the accuracy of the voice response.
[0136] It should be noted that the above-mentioned method embodiments, or various possible implementation methods in each method embodiment, can be executed separately, or, under the premise that there is no contradiction, can also be executed in combination with each other. The specific implementation can be determined according to actual usage requirements, and the embodiments of this application do not limit this.
[0137] The voice response method provided in the embodiment of the present application can be executed by a voice response device. In the embodiment of the present application, the voice response device provided in the embodiment of the present application is described by taking the voice response device executing the voice response method as an example.
[0138] Attachment Figure 6 FIG. 1 shows a possible structural diagram of a voice response device involved in an embodiment of the present application. Figure 6 As shown, the voice response device 800 may include: an acquisition module 801, a recognition module 802, a splicing module 803 and a sending module 804;
[0139] The acquisition module 801 is configured to acquire first voice data during a voice call;
[0140] The recognition module 802 is configured to perform text conversion on the first voice data acquired by the acquisition module 801 to obtain a first text;
[0141] The splicing module 803 is configured to, if there is a target cached text whose characteristic information satisfies the first condition in the historical cached text, splice the first text obtained by the recognition module 802 with the target cached text to obtain a second text, where the historical cached text is the text corresponding to the voice data received earlier than the first voice data during the voice call;
[0142] The sending module 804 is configured to send a corresponding operation instruction to a second electronic device based on the second text identified by the identifying module 802 .
[0143] In some embodiments of the present application, the historical cached text includes multiple cached subtexts, and the splicing module 803 is further used to: obtain characteristic information of each cached subtext from the cache area of the first electronic device, the characteristic information including target intent information and voice response information; based on the characteristic information of each cached subtext, determine the cached subtext whose characteristic information meets the first condition from the multiple cached subtexts to obtain the target cached text. It should be noted that the target intent information is used to indicate whether the semantic information of the cached subtext is related to the target intent, and the voice response information is used to indicate whether the second electronic device has responded to the semantic information corresponding to the cached subtext.
[0144] In some embodiments of the present application, the above-mentioned splicing module 803 is also used to: when the semantic information of the first cache sub-text is irrelevant to the target intention and the above-mentioned second electronic device does not respond to the semantic information corresponding to the first cache sub-text, splice the above-mentioned first text with the above-mentioned first cache sub-text to obtain a second text, and the first cache sub-text is one or more of the multiple cache sub-texts.
[0145] In some embodiments of the present application, the sending module 804 is specifically used to: when the silence period after the first voice data is greater than or equal to the preset period, send a first operation instruction to the second electronic device, and the above-mentioned first operation instruction is used to instruct the above-mentioned second electronic device to respond to the semantic information corresponding to the second text; when the above-mentioned silence period is less than the preset period, send a second operation instruction to the second electronic device, and the above-mentioned second operation instruction is used to instruct the above-mentioned second electronic device not to perform the operation.
[0146] In some embodiments of the present application, the voice response device provided by the present application further includes a cache module 805, which is used to update the above-mentioned target cache text in the historical cache text to the above-mentioned second text and cache feature information of the second text.
[0147] In some embodiments of the present application, the above-mentioned acquisition module 801 is also used to: receive multiple segments of voice data during a voice call; split the voice data in the multiple segments of voice data received during the voice call, in which the silence duration between two adjacent segments exceeds a second preset duration, to obtain multiple sub-voice data, wherein the first voice data is a sub-voice data corresponding to the current moment.
[0148] In the voice response device provided in this application, during a voice call, the voice response device's acquisition module 801 first acquires first voice data. Then, the recognition module 802 performs text conversion on the first voice data to obtain a first text. Then, if there is a target cached text whose characteristic information satisfies a first condition in the historical cached text, the splicing module 803 splices and processes the first text with the target cached text to obtain a second text. The historical cached text corresponds to the voice data received earlier than the first voice data during the voice call. Finally, the sending module 804 sends a corresponding operation instruction to the second electronic device based on the second text. In this solution, the voice response device can select the target cached text that satisfies the first condition from the historical cached text based on the characteristic information and splice it with the first text. The second text then sends the corresponding operation instruction to the second electronic device based on the spliced second text. Therefore, the second electronic device can accurately respond to the operation instruction based on the operation instruction, reducing the misjudgment of the operation instruction caused by excessive punctuation during the text conversion process during the voice response, and improving the accuracy of the voice response.
[0149] The voice response device in the embodiments of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiments of the present application do not specifically limit it.
[0150] The voice response device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0151] The voice response device provided in the embodiment of the present application can achieve Figures 1 to 5 To avoid repetition, the various processes implemented in the method embodiment are not described here.
[0152] Alternatively, as Figure 7 As shown, an embodiment of the present application further provides an electronic device 900, including a processor 901 and a memory 902, wherein the memory 902 stores a program or instruction that can be run on the processor 901, and when the program or instruction is executed by the processor 901, the various steps of the above-mentioned voice response method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, they are not described here.
[0153] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0154] Figure 8 A schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.
[0155] The electronic device 100 includes but is not limited to components such as a radio frequency unit 101 , a network module 102 , an audio output unit 103 , an input unit 104 , a sensor 105 , a display unit 106 , a user input unit 107 , an interface unit 108 , a memory 109 , and a processor 110 .
[0156] Those skilled in the art will understand that the electronic device 100 may also include a power source (such as a battery) to power each component, and the power source may be logically connected to the processor 110 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. Figure 8 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.
[0157] The radio frequency unit 101 is configured to obtain first voice data during a voice call;
[0158] The processor 110 is configured to perform text conversion on the first voice data acquired by the radio frequency unit 101 to obtain a first text;
[0159] The processor 110 is further configured to, if there is a target cached text whose characteristic information satisfies the first condition in the historical cached text, concatenate the first text obtained by the processor 110 with the target cached text to obtain a second text, where the historical cached text is text corresponding to voice data received earlier than the first voice data during the voice call;
[0160] The radio frequency unit 101 is further configured to send a corresponding operation instruction to the second electronic device based on the second text recognized by the processor 110 .
[0161] In some embodiments of the present application, the historical cached text includes multiple cached subtexts, and the processor 110 is further configured to: obtain characteristic information of each cached subtext from the cache area of the first electronic device, the characteristic information including target intent information and voice response information; and based on the characteristic information of each cached subtext, determine a cached subtext from the multiple cached subtexts whose characteristic information satisfies a first condition, thereby obtaining the target cached text. It should be noted that the target intent information is used to indicate whether the semantic information of the cached subtext is relevant to the target intent, and the voice response information is used to indicate whether the second electronic device has responded to the semantic information corresponding to the cached subtext.
[0162] In some embodiments of the present application, the processor 110 is further used to: when the semantic information of the first cache subtext is irrelevant to the target intention and the second electronic device does not respond to the semantic information corresponding to the first cache subtext, splice the first text with the first cache subtext to obtain a second text, where the first cache subtext is one or more of the multiple cache subtexts.
[0163] In some embodiments of the present application, the radio frequency unit 101 is specifically used to: when the silence duration of the first voice data hand is greater than or equal to a preset duration, send a first operation instruction to the second electronic device, and the above-mentioned first operation instruction is used to instruct the above-mentioned second electronic device to respond to the semantic information corresponding to the second text; when the above-mentioned silence duration is less than the preset duration, send a second operation instruction to the second electronic device, and the above-mentioned second operation instruction is used to instruct the above-mentioned second electronic device not to perform the operation.
[0164] In some embodiments of the present application, the voice response device provided by the present application further includes a memory 109, which is used to update the above-mentioned target cache text in the historical cache text to the above-mentioned second text and cache feature information of the second text.
[0165] In some embodiments of the present application, the above-mentioned radio frequency unit 101 is also used to: receive multiple segments of voice data during a voice call; split the voice data received during the voice call, in which the silence duration between two adjacent segments exceeds a second preset duration, to obtain multiple sub-voice data, and the first voice data is a sub-voice data corresponding to the current moment.
[0166] In the voice response device provided in this application, during a voice call, the RF unit 101 of the voice response device first obtains first voice data. Then, the processor 110 performs text conversion on the first voice data to obtain a first text. Then, if a target cached text with characteristic information that meets a first condition exists in the historical cached text, the processor 110 concatenates the first text with the target cached text to obtain a second text. The historical cached text corresponds to voice data received earlier than the first voice data during the voice call. Finally, the RF unit 101 sends a corresponding operation instruction to a second electronic device based on the second text. In this solution, the voice response device can select the target cached text that meets the first condition from the historical cached text based on the characteristic information to concatenate with the first text, and send the corresponding operation instruction to the second electronic device based on the concatenated second text. Therefore, the second electronic device can accurately respond to the operation instruction based on the operation instruction, reducing the misjudgment of the operation instruction caused by excessive punctuation during the text conversion process during the voice response, and improving the accuracy of the voice response.
[0167] It should be understood that in an embodiment of the present application, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042, and the graphics processor 1041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 107 includes a touch panel 1071 and at least one of other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an operating stick, which will not be repeated here.
[0168] The memory 109 can be used to store software programs and various data. The memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 109 may include a volatile memory or a non-volatile memory, or the memory x09 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 109 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0169] Processor 110 may include one or more processing units. Optionally, processor 110 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 110.
[0170] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned voice response method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0171] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0172] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned voice response method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0173] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0174] An embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-mentioned voice response method embodiment and can achieve the same technical effects. To avoid repetition, it will not be described here.
[0175] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0176] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0177] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A voice response method, characterized in that: The method is performed by a first electronic device and includes: During a voice call, obtaining first voice data, and performing text conversion on the first voice data to obtain a first text; If there is a target cache text whose characteristic information satisfies the first condition in the historical cache text, concatenating the first text with the target cache text to obtain a second text, where the historical cache text is the text corresponding to the voice data received earlier than the first voice data during the voice call; Based on the second text, a corresponding operation instruction is sent to the second electronic device.
2. The method according to claim 1, characterized in that The historical cached text includes one or more cached subtexts; the method further includes: Acquire feature information of each cached subtext from a cache area of the first electronic device, the feature information including target intent information and voice response information; Based on the characteristic information of each cached subtext, determining a cached subtext whose characteristic information meets a first condition from the multiple cached subtexts to obtain the target cached text; The target intent information is used to indicate whether the semantic information of the cached subtext is related to the target intent, and the voice response information is used to indicate whether the second electronic device responds to the semantic information corresponding to the cached subtext.
3. The method according to claim 2, characterized in that When there is a target cache text whose characteristic information satisfies the first condition in the historical cache text, the first text and the target cache text are concatenated to obtain a second text, including: If the semantic information of the first cached subtext is irrelevant to the target intent and the second electronic device does not respond to the semantic information corresponding to the first cached subtext, concatenate the first text with the first cached subtext to obtain a second text; The first cache subtext is one or more of the multiple cache subtexts.
4. The method according to claim 1, wherein The sending a corresponding operation instruction to the second electronic device based on the second text includes: When the silence duration after the first voice data is greater than or equal to a first preset duration, sending a first operation instruction to the second electronic device, the first operation instruction being used to instruct the second electronic device to respond to the semantic information corresponding to the second text; In a case where the silence duration is less than the first preset duration, a second operation instruction is sent to the second electronic device, where the second operation instruction is used to instruct the second electronic device not to perform an operation.
5. The method according to claim 1, wherein The method further comprises: During the voice call, receiving multiple segments of voice data; Splitting the voice data, in the multiple voice data segments received during the voice call, wherein the silence duration between two voice data segments exceeds a second preset duration, to obtain a plurality of sub-voice data; The first voice data is a sub-voice data corresponding to the current moment.
6. A voice response device, characterized in that: The device includes an acquisition module, an identification module, a splicing module and a sending module; The acquisition module is used to acquire the first voice data during the voice call; The recognition module is configured to perform text conversion on the first voice data acquired by the acquisition module to obtain a first text; The splicing module is configured to, if there is a target cached text in the historical cached text whose characteristic information satisfies the first condition, splice the first text obtained by the recognition module with the target cached text to obtain a second text, wherein the historical cached text is the text corresponding to the voice data received earlier than the first voice data during the voice call; The sending module is configured to send a corresponding operation instruction to a second electronic device based on the second text recognized by the recognition module.
7. The device according to claim 6, characterized in that The historical cached text includes a plurality of cached subtexts; the splicing module is further configured to: Acquire feature information of each cached subtext from a cache area of the first electronic device, the feature information including target intent information and voice response information; Based on the characteristic information of each cached subtext, determining a cached subtext whose characteristic information meets a first condition from the multiple cached subtexts to obtain the target cached text; The target intent information is used to indicate whether the semantic information of the cached subtext is related to the target intent, and the voice response information is used to indicate whether the second electronic device responds to the semantic information corresponding to the cached subtext.
8. The device according to claim 7, characterized in that The splicing module is further specifically used for: If the semantic information of the first cached subtext is irrelevant to the target intent and the second electronic device does not respond to the semantic information corresponding to the first cached subtext, concatenating the first text with the first cached subtext to obtain a second text; The first cache subtext is one or more of the multiple cache subtexts.
9. The device according to claim 6, characterized in that The sending module is specifically used to: When the silence duration after the first voice data is greater than or equal to a preset duration, sending a first operation instruction to the second electronic device, the first operation instruction being used to instruct the second electronic device to respond to the semantic information corresponding to the second text; When the silence duration is less than the preset duration, a second operation instruction is sent to the second electronic device, where the second operation instruction is used to instruct the second electronic device not to perform an operation.
10. The method according to claim 6, characterized in that The acquisition module is further used to: During the voice call, receiving multiple segments of voice data; Splitting the voice data, in the plurality of voice data segments received during the voice call, wherein the silence duration between two adjacent voice data segments exceeds a second preset duration, to obtain a plurality of sub-voice data; The first voice data is a sub-voice data corresponding to the current moment.