Voice response processing method and device, storage medium, and program product

By recording the moment the user's voice input ends and outputting a reassuring voice message when the terminal device detects the end of the user's voice input, the problem of excessively long waiting time for users is solved. This enables continuous feedback and connection error prompts when no response is received, thus improving the user experience.

WO2026056305A1PCT designated stage Publication Date: 2026-03-19BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

The problem stems from users waiting too long for feedback from the large model service device on their terminal devices.

Method used

The system records the receiving time when the user's voice input ends, sends the voice information to the large model service device, and outputs the soothing voice corresponding to the soothing information if no response information is received within a preset time. This includes a soothing information generation request and the loop output of the soothing voice until a response information is received.

Benefits of technology

By outputting reassuring voice messages and information, user waiting time is reduced, user experience is improved, and continuous feedback and connection error alerts are provided even when no feedback is received for an extended period.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025093217_19032026_PF_FP_ABST
    Figure CN2025093217_19032026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to the technical field of artificial intelligence, and provide a voice response processing method and device, a storage medium, and a program product. The method is applied to an electronic device, and comprises: when it is detected that an input voice of a user ends, recording reception time of the input voice; sending voice information corresponding to the input voice to a large model service device, so that the large model service device generates response information corresponding to the voice information; and in response to not receiving, within a first preset time period after the reception time, the response information sent by the large model service device, outputting a soothing voice corresponding to soothing information, wherein the soothing information is sent by the large model service device.
Need to check novelty before this filing date? Find Prior Art

Description

Voice response processing method and device, storage medium and program product

[0001] The present application claims priority from the Chinese patent application No. 202411296767.3 filed on September 14, 2024, and entitled "Voice response processing method and device, storage medium and program product", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] Embodiments of the present disclosure relate to the field of artificial intelligence technology, and particularly relate to a voice response processing method, device, storage medium and program product. BACKGROUND

[0003] With the continuous development of artificial intelligence technology, the current terminal device can receive a question input by a user, and answer the question input by the user by interacting with a server and using a large model in the server.

[0004] Currently, when the terminal device answers the question by using the large model in the server, the user can only wait for the terminal device to return the result.

[0005] However, the inventors have found that in the process of the current terminal device answering the question, at least the following technical problem exists: the user waits for a long time for the feedback of the terminal device. SUMMARY

[0006] Embodiments of the present disclosure provide a voice response processing method, device, storage medium and program product to overcome the problem that the user waits for a long time for the feedback of the terminal device.

[0007] In a first aspect, embodiments of the present disclosure provide an information display method, comprising: recording a receiving time of an input voice when detecting that the input voice of a user ends; sending voice information corresponding to the input voice to a large model service device, so that the large model service device generates response information corresponding to the voice information; and in response to that no response information sent by the large model service device is received within a first preset time period after the receiving time, outputting a pacification voice corresponding to pacification information, wherein the pacification information is sent by the large model service device.

[0008] In a second aspect, the embodiments of the present disclosure provide a voice response processing device, comprising: a time recording unit configured to record a receiving time of an input voice when detecting that the input voice of a user ends; an information sending unit configured to send voice information corresponding to the input voice to a large model service device, so that the large model service device generates response information corresponding to the voice information; and a voice output unit configured to output pacification voice corresponding to pacification information in response to that no response information sent by the large model service device is received within a first preset time length after the receiving time, wherein the pacification information is sent by the large model service device.

[0009] In a third aspect, the embodiments of the present disclosure provide an electronic device, comprising: a processor and a memory; the memory stores computer execution instructions; and the processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the voice response processing method according to the first aspect and various possible designs of the first aspect.

[0010] In a fourth aspect, the embodiments of the present disclosure provide a computer readable storage medium, which stores computer execution instructions, and when a processor executes the computer execution instructions, the voice response processing method according to the first aspect and various possible designs of the first aspect is implemented.

[0011] In a fifth aspect, the embodiments of the present disclosure provide a computer program product, which comprises a computer program, and when the computer program is executed by a processor, the voice response processing method according to the first aspect and various possible designs of the first aspect is implemented. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art without creative labor.

[0013] FIG. 1 is a schematic diagram of an application scenario of the voice response processing method provided by the embodiments of the present disclosure;

[0014] FIG. 2 is a flowchart of the voice response processing method provided by the embodiments of the present disclosure;

[0015] FIG. 3 is a schematic diagram of the overall flow of the voice response processing method provided by the embodiments of the present disclosure;

[0016] FIG. 4 is a schematic diagram of the structure of the voice response processing device provided by the embodiments of the present disclosure;

[0017] FIG. 5 is a schematic diagram of the structure of the electronic device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION

[0018] The technical solutions and advantages of the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.

[0019] With the development of artificial intelligence technology, terminal devices can now receive user questions and interact with large models of cloud servers to analyze questions and generate answers.

[0020] However, after the terminal device sends the user's question to the cloud server, it can only wait for the server to return the answer, causing the user to wait for feedback for too long.

[0021] To solve the above technical problems, the inventors propose the following technical concept: in the case that no response returned by the large model service device is received after a period of time after sending voice information to the large model service device, outputting a pacification voice corresponding to pacification information sent by the large model service device.

[0022] Referring to FIG. 1, FIG. 1 is a schematic diagram of an application scenario of a voice response processing method provided by an embodiment of the present disclosure. As shown in FIG. 1, the application scenario of the voice response processing method includes an electronic device 101 and a large model service device 102.

[0023] The electronic device 101 can include a mobile phone, a headset, a computer, a server, a tablet, a personal digital assistant (PDA), a notebook, and the like, which can input voice, transmit data, and receive data.

[0024] The large model service device 102 can include a single server or a cluster composed of multiple servers to achieve, and in possible cases, a computer or a notebook with stronger computing power can be used as an alternative.

[0025] The electronic device 101 and the large model service device 102 can be connected in a wired or wireless manner.

[0026] The electronic device 101 receives the input speech of the user, and sends speech information corresponding to the input speech to the large model service device 102. The large model service device 102 generates response information according to the speech information, and sends the response information to the electronic device 101. The electronic device 101 outputs speech corresponding to the response information. The large model service device 102 is also configured to generate soothing information when the generation of the response information takes too long or when the electronic device 101 sends a request for generating soothing information, and send the soothing information to the electronic device 101. The electronic device 101 is also configured to output soothing speech corresponding to the soothing information.

[0027] Referring to FIG. 2, FIG. 2 is a flowchart of a speech response processing method provided by an embodiment of the present disclosure. The method of the present embodiment can be applied to the electronic device of FIG. 1. The speech response processing method includes the following steps.

[0028] S201: When detecting that the input speech of the user ends, recording a receiving time of the input speech.

[0029] In this step, the receiving time of the input speech can be obtained by acquiring a timestamp after a period of time in which the input speech of the user is not detected. Alternatively, the receiving time of the input speech can be obtained by acquiring a timestamp after receiving an end signal sent by a headset or a button of the electronic device, and determining that the input speech of the user ends.

[0030] Specifically, when the input speech of the user is not detected, it can be determined that the input speech ends when the volume of the speech received within a preset end determination duration is less than a preset volume threshold. The end signal can include a hand release, a click, a double click, etc.

[0031] S202: Sending speech information corresponding to the input speech to a large model service device, so that the large model service device generates response information corresponding to the speech information.

[0032] In this step, the speech information corresponding to the input speech can include an audio file, or text data obtained by converting the speech into text. The speech information can be sent to the large model service device in the form of a file, a message, or a data packet. The large model service device can input the speech information into a large model directly to obtain response information output by the large model, or convert the speech information into text content and input the text content into the large model to obtain response information output by the large model.

[0033] S203: In response to that no response information sent by the large model service device is received within a first preset time duration after the receiving time, outputting soothing speech corresponding to soothing information, wherein the soothing information is sent by the large model service device.

[0034] In this step, outputting the pacification voice corresponding to the pacification information includes: if the pacification information is audio, directly playing the pacification voice; if the pacification information is text, converting the pacification information into voice and then playing.

[0035] As can be seen from the description of the above embodiments, the embodiments of the present disclosure reduce the waiting time of the user by recording the receiving time of the input voice when detecting the end of the input voice of the user, sending the voice information corresponding to the input voice to the large model service device, generating the response information by the large model service device, and outputting the pacification voice in the case that the response information is not received within the first preset time period after the receiving time.

[0036] In a possible implementation, in the step S203, in response to the fact that the response information sent by the large model service device is not received within the first preset time period after the receiving time, the pacification voice corresponding to the pacification information is outputted, including:

[0037] S2031: If the response information sent by the large model service device is not received within the first preset time period after the receiving time, a pacification information generation request is sent to the large model service device to make the large model service device generate the pacification information corresponding to the voice information.

[0038] In this step, the pacification information generation request can be preset or can correspond to the voice information. The sending mode of the pacification information can include sending the pacification information to the large model service device in the format of a message or a data packet. The large model service device can input the content of the pacification information generation request and the voice information into a pacification information generation model to obtain the corresponding pacification information.

[0039] S2032: Receiving the pacification information sent by the large model service device.

[0040] In this step, the pacification information can be received in the form of a data packet or a message.

[0041] The pacification information can be text or audio data.

[0042] S2033: Outputting the pacification voice corresponding to the pacification information.

[0043] In this step, if the pacification information is text, the text is converted into voice and then outputted; if the pacification information is audio data, the audio data is directly played to output the pacification voice.

[0044] As can be known from the description of the above embodiments, the embodiments of the present disclosure actively send a pacification information generation request to the large model service device in the case that no reply information is received within a certain time after the receiving time, so that the large model service device generates pacification information, and the pacification information is output after being received, thereby completing the generation and output of the pacification information and reducing the time for the user to wait for feedback.

[0045] In a possible implementation, after the pacification voice corresponding to the pacification information is output in the step S203, the method further includes:

[0046] S2031: If no feedback information and reply information sent by the large model service device are received within a second preset time length after the pacification voice is output, the pacification sound effect is output at a preset time interval, wherein the feedback information is used to confirm that the voice information has been received by the large model service device.

[0047] In this step, the preset time interval can be a fixed time interval, or a time interval corresponding to the output times of the pacification sound effect. For example, if no reply information is received within 1 second after the pacification voice is output, the pacification sound effect is output, if no reply information is received within 1 second again, the pacification sound effect is output again, and if no reply information is received within 2 seconds again, the pacification sound effect is output again.

[0048] The second preset time length is, for example, 0.5 seconds, 1 second, 1.5 seconds, etc. The second preset time length and the preset time interval can be preset by the staff according to experimental data or experience parameters.

[0049] As can be known from the description of the above embodiments, the embodiments of the present disclosure output the pacification sound effect at a preset time interval in the case that no reply information is received, so as to avoid that the user cannot receive feedback for a long time and reduce the waiting time of the user.

[0050] In a possible implementation, after the pacification sound effect is output at the preset time interval in the step S2031, the method further includes:

[0051] S220: Count the output times of the pacification sound effect.

[0052] In this step, the output times variable can be increased by n each time the pacification sound effect is output.

[0053] The n is a positive integer.

[0054] S221: If the output times are greater than a preset output times threshold, output a timeout pacification voice.

[0055] In this step, the output times threshold and the timeout pacification voice can be preset by the staff.

[0056] The timeout appeasing voice is, for example, "server connection timeout, please check the network settings", "server connection timeout, please try again after a period of time", and the like.

[0057] As can be known from the description of the above embodiments, the output times of the technical appeasing voice are determined, and in a case where the output times exceed a threshold value, it is determined that the connection with the large model service device is abnormal, a timeout appeasing voice is output to prompt, and the user is prompted to check the network state or to retry.

[0058] In a possible implementation, after the appeasing voice is output at the preset time interval in step S2031, the method further includes:

[0059] S230: If the feedback information and the response information sent by the large model service device are not received within a third preset time length after the appeasing voice is output, a timeout appeasing voice is output, where the third preset time length is greater than the second preset time length.

[0060] In this step, the third preset time length can also be set by the staff in advance, for example, 5 seconds, 6 seconds, 7.5 seconds, and the like.

[0061] Other contents of this step are similar to step S221 described above, and will not be described here again.

[0062] As can be known from the description of the above embodiments, whether to output a timeout appeasing voice is determined by taking whether the time length exceeds a third preset time length as a judgment standard, and the connection abnormality is prompted when there is no response from the large model service device for a long time.

[0063] In a possible implementation, after the appeasing voice corresponding to the appeasing information is output in step S203, the method further includes:

[0064] S240: If the feedback information sent by the large model service device according to the voice information is received within a second preset time length after the appeasing voice is output, but the response information sent by the large model service device is not received, the appeasing voice is output at a preset time interval until the response information sent by the large model service device is received, and the response voice corresponding to the response information is output, where the feedback information is used to confirm that the voice information has been received by the large model service device.

[0065] In this step, the response information can be text or audio data, if the response information is text, the text is converted into audio to obtain the response voice, and then the response voice is output, if the response information is audio, the response information is directly output.

[0066] As can be seen from the description of the above embodiments, the embodiments of the present disclosure output the pacifying sound effect until the response information is received, and output the corresponding voice information, so as to continuously feedback in the case where the feedback information and the response information are not received, and output the response voice in time when the response information is received.

[0067] In a possible implementation, after the pacifying voice corresponding to the pacifying information is output in step S203, the method further includes:

[0068] S250: If the feedback information sent by the large model service device according to the voice information is received within a second preset time length after the pacifying voice is output, but the response information is not received, the pacifying sound effect is output at a preset time interval, wherein the feedback information is used for the large model service device to confirm that the voice information has been received.

[0069] In this step, the feedback information can be received by receiving a message, a data packet, etc.

[0070] In the disclosure, the pacifying sound effect can be pre-stored.

[0071] As can be seen from the description of the above embodiments, the embodiments of the present disclosure output the pacifying sound effect until the response information is received, and output the corresponding voice information, so as to continuously feedback in the case where the feedback information and the response information are not received, and output the response voice in time when the response information is received.

[0072] In a possible implementation, after the pacifying sound effect is output at the preset time interval in step S250, the method further includes:

[0073] S251A: Count the output times of the pacifying sound effect.

[0074] S252A: If the output times are greater than or equal to a preset number of times, and the response information is not received, a delayed feedback pacifying generation request is sent to the large model service device, so that the large model service device generates delayed feedback pacifying information according to the delayed feedback pacifying generation request and the voice information.

[0075] In this step, the counting method of the output times can include that each time the pacifying sound effect is output, the output times variable is added by n, wherein n is a positive integer. The delayed feedback pacifying generation request can be pre-set, or can be generated according to the voice information or the identifier of the voice information. The identifier of the voice information can be generated when the voice information is sent to the large model service device. The large model service device inputs the delayed feedback pacifying generation request and the voice information into the large model to obtain the delayed feedback pacifying information generated by the large model.

[0076] For example, in the case where the number of output times is greater than 2 and no response information is received, a delayed feedback pacification generation request is sent to the large model service device; for another example, in the case where the number of output times is greater than 3 and no response information is received, a delayed feedback pacification generation request is sent to the large model service device.

[0077] S253A: receiving the delayed feedback pacification information sent by the large model service device.

[0078] In this step, the delayed feedback pacification information can also be received by receiving messages, data packets, etc.

[0079] S254A: outputting the delayed feedback voice corresponding to the delayed feedback pacification information.

[0080] In this step, if the delayed feedback pacification information is audio, the delayed feedback pacification information is played; if the delayed feedback pacification information is text, the text is converted into audio and played.

[0081] As can be seen from the description of the above embodiments, the embodiments of the present disclosure count the number of output times of pacification sound effects, in the case where the number of output times exceeds a certain number, a delayed feedback pacification generation request is sent to the large model service device, and the delayed feedback pacification information returned by the server is received, the delayed feedback voice corresponding to the delayed feedback pacification information is output, so that in the case where the large model service device can be connected, the matched delayed feedback pacification information is generated by the large model service device, the targeted feedback is obtained, and the waiting time of the user is reduced.

[0082] In a possible implementation, after the pacification sound effect is output at the preset time interval in the above step S250, the method further includes:

[0083] S251B: if the response information is still not received within the third preset time length after the pacification voice is output, a delayed feedback pacification generation request is sent to the large model service device, so that the large model service device generates the delayed feedback pacification information according to the delayed feedback pacification generation request and the voice information, wherein the third preset time length is greater than the second preset time length.

[0084] In this step, for example, if no response information is received within 5 seconds of outputting the pacification voice, a delayed feedback pacification generation request is sent to the large model service device; for another example, if no response information is received within 7 seconds of outputting the pacification voice, a delayed feedback pacification generation request is sent to the large model service device.

[0085] The third preset time length can be set by the staff according to experimental data or experience parameters in advance, for example, 5, 5.5, 6, 7, etc.

[0086] S252B: receiving the delayed feedback pacification information sent by the large model service device.

[0087] This step is similar to the above step S253A, and will not be described here.

[0088] S253B: Output the delayed feedback soothing information corresponding to the delayed feedback voice.

[0089] This step is similar to the above step S254A, and will not be described here.

[0090] As can be seen from the description of the above embodiments, the disclosed embodiments can actively request the large model service device to generate matching delayed feedback soothing information in the case of no response information for a long time, obtain and output targeted feedback, and reduce the waiting time of the user, by sending a delayed feedback soothing generation request to the large model service device in the case of no response information for a certain time after outputting the soothing voice, receiving the delayed feedback soothing information returned by the large model service device, and outputting the delayed feedback soothing information.

[0091] In a possible implementation, after the delayed feedback soothing information corresponding to the delayed feedback voice is output in the above step S254A or S253B, the method further includes:

[0092] S260: Generate a feedback task corresponding to the delayed feedback soothing information.

[0093] In this step, the feedback task can be generated by using a preset task type and the identifier or field of the delayed feedback soothing information, or the identifier or field of the delayed feedback soothing information can be directly used as the feedback task.

[0094] S261: Add the feedback task to a task queue.

[0095] In this step, the feedback task can be written to the end of the task queue.

[0096] S262: If the first feedback task corresponding to the response information in the task queue sent by the large model service device is received, output the response voice corresponding to the response information.

[0097] In this step, the task queue can be searched according to the identifier or field in the response information to determine whether the response information is the response information corresponding to the first feedback task, and if so, the corresponding response voice is output.

[0098] As can be seen from the description of the above embodiments, the disclosed embodiments can output the response voice corresponding to each response information in sequence in the case of multiple input voices for question and answer and no timely response information for the input voice, by generating a feedback task corresponding to the delayed feedback soothing information and writing the feedback task to the task queue.

[0099] In a possible implementation, before the step S201, the method further includes:

[0100] S270: in response to receiving the interaction start instruction sent by the earphone, outputting a wake-up sound and receiving the input voice of the user.

[0101] In this step, the interaction start instruction can include long pressing, double clicking, etc., or can be a preset instruction sent by the earphone. The wake-up sound can be an audio preset by the staff.

[0102] From the description of the above embodiments, it can be known that the embodiments of the present disclosure can output a wake-up sound, receive an input voice, and prompt the user that the voice input can be performed by receiving an interaction start instruction sent by an earphone.

[0103] In a possible implementation, after the step S202, the method further includes:

[0104] S280: if the response information includes operation information, performing an operation process corresponding to the operation information to obtain an operation result.

[0105] In this step, the operation information can be an operation instruction or an operation process. The operation instruction or the operation process can be input into a preset program to perform the operation process and obtain an operation result. The operation result can include operation success, operation end, operation completion, etc.

[0106] S281: outputting an end sound effect corresponding to the operation result.

[0107] In this step, the corresponding sound effect can be output according to the type of the operation result.

[0108] For example, a neutral sound effect is output when the operation is completed, a success sound effect is output when the operation is successful, and a failure sound effect is output when the operation fails.

[0109] From the description of the above embodiments, it can be known that the embodiments of the present disclosure can output a wake-up sound, receive an input voice, and prompt the user that the voice input can be performed by receiving an interaction start instruction sent by an earphone.

[0110] In a possible implementation, after the step S202, the method further includes:

[0111] S290: if an interrupt instruction input by the user is received in the process of outputting the calming voice, the output of the calming voice is interrupted.

[0112] In this step, the interrupt instruction can be directly input by the user through the screen, the key of the electronic device, or through an external device such as a headset, or can be a voice input by the user with an interrupt intention. The interrupt instruction is, for example, a single click, a slide, and the like.

[0113] If the user inputs the interrupt instruction in the form of a voice, the electronic device can send the interrupt instruction to the large model service device, the large model service device generates a corresponding to-be-sent instruction using the interrupt instruction, and in response to receiving the to-be-sent instruction sent by the large model service device, the output of the calming voice is interrupted.

[0114] S291: Output the interrupt sound effect.

[0115] In this step, the interrupt sound effect can be pre-set by the staff.

[0116] As can be known from the description of the above embodiments, by receiving the interrupt instruction input by the user in the process of outputting the calming voice, the output of the calming voice is interrupted and the interrupt sound effect is output, so that the user can stop the voice output when the user wants to stop the voice output.

[0117] Referring to FIG. 3, FIG. 3 is a schematic diagram of the overall flow of the voice response processing method provided by the embodiments of the present disclosure. As shown in FIG. 3, the overall flow of the voice response processing method includes:

[0118] In response to receiving the interaction start instruction, a wake-up sound is output, and the input voice of the user is received.

[0119] A microphone switching sound effect is output when the input voice ends, and a receiving time is recorded.

[0120] If no response information sent by the large model service device is received within a first preset time period after the receiving time, a calming voice corresponding to calming information is output, wherein the calming information can be generated by the large model service device after sending a calming information generation request to the large model service device, or can be generated by the large model service device in the case where no response information is generated within a certain time.

[0121] If no response information is received within a second preset time period after the calming voice is output, the calming sound effect is output in a loop.

[0122] If the number of times of outputting the calming sound effect reaches an output frequency threshold, or no calming information is received within a third preset time period after the calming voice is output, a timeout feedback is output. The timeout feedback includes a timeout calming voice that no feedback information and response information sent by the large model service device are received, and can also include a delayed feedback voice that feedback information sent by the large model service device is received but no response information is received.

[0123] If an interrupt instruction is received during output of the appeasing voice, the output of the appeasing voice is interrupted, and an interrupt sound effect is output.

[0124] In response to completion of output of the appeasing voice, a microphone switching sound effect is output.

[0125] If the response information is received, a response voice corresponding to the response information is output.

[0126] Referring to FIG. 4, FIG. 4 is a structural schematic diagram of a voice response processing device provided by an embodiment of the present disclosure. As shown in FIG. 4, the voice response processing device 400 includes a time recording unit 401, an information sending unit 402, and a voice output unit 403.

[0127] The time recording unit 401 is configured to record a receiving time of the input voice when it is detected that the input voice of the user ends.

[0128] The information sending unit 402 is configured to send voice information corresponding to the input voice to a large model service device, so that the large model service device generates response information corresponding to the voice information.

[0129] The voice output unit 403 is configured to, in response to the fact that no response information sent by the large model service device is received within a first preset time length after the receiving time, output an appeasing voice corresponding to appeasing information, wherein the appeasing information is sent by the large model service device.

[0130] The device provided by the embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects, which will not be described here again in the embodiment.

[0131] In a possible implementation manner, the voice output unit 403 is specifically configured to, if no response information sent by the large model service device is received within the first preset time length after the receiving time, send an appeasing information generation request to the large model service device, so that the large model service device generates the appeasing information corresponding to the voice information; receive the appeasing information sent by the large model service device; and output the appeasing voice corresponding to the appeasing information.

[0132] The device provided by the embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects, which will not be described here again in the embodiment.

[0133] In a possible implementation manner, the voice response processing device 400 further includes a sound effect output unit 404.

[0134] The sound effect output unit 404 is configured to output the soothing sound effect at preset time intervals if the feedback information and the response information sent by the large model service device are not received within a second preset time length after the soothing voice is output, where the feedback information is used for the large model service device to confirm that the voice information has been received.

[0135] The device provided in the embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects. Details are not described herein again.

[0136] In a possible implementation, the voice output unit 403 is further configured to count the number of times of outputting the soothing sound effect, and output the timeout soothing voice if the number of times of outputting is greater than a preset number of times threshold.

[0137] The device provided in the embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects. Details are not described herein again.

[0138] In a possible implementation, the voice output unit 403 is further configured to output the timeout soothing voice if the feedback information and the response information sent by the large model service device are still not received within a third preset time length after the soothing voice is output, where the third preset time length is greater than the second preset time length.

[0139] The device provided in the embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects. Details are not described herein again.

[0140] In a possible implementation, the sound effect output unit 404 is further configured to output the soothing sound effect at preset time intervals if the feedback information sent by the large model service device according to the voice information is received but the response information sent by the large model service device is not received within a second preset time length after the soothing voice is output, and output the response voice corresponding to the response information if the response information sent by the large model service device is received, where the feedback information is used for the large model service device to confirm that the voice information has been received.

[0141] The device provided in the embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects. Details are not described herein again.

[0142] In a possible implementation, the sound effect output unit 404 is further configured to output the soothing sound effect at preset time intervals if the feedback information sent by the large model service device according to the voice information is received but the response information is not received within a second preset time length after the soothing voice is output, where the feedback information is used for the large model service device to confirm that the voice information has been received.

[0143] The device provided by the embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects, which will not be described here again in the embodiment.

[0144] In a possible implementation, the voice output unit 403 is further configured to count the number of times of outputting the pacification sound effect, and if the number of times is greater than or equal to a preset number and no response information is received, send a delayed feedback pacification generation request to the large model service device, so that the large model service device generates delayed feedback pacification information according to the delayed feedback pacification generation request and the voice information, receive the delayed feedback pacification information sent by the large model service device, and output the delayed feedback voice corresponding to the delayed feedback pacification information.

[0145] The device provided by the embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects, which will not be described here again in the embodiment.

[0146] In a possible implementation, the voice output unit 403 is further configured to, if no response information is received within a third preset time length after outputting the pacification voice, send a delayed feedback pacification generation request to the large model service device, so that the large model service device generates delayed feedback pacification information according to the delayed feedback pacification generation request and the voice information, wherein the third preset time length is greater than the second preset time length, receive the delayed feedback pacification information sent by the large model service device, and output the delayed feedback voice corresponding to the delayed feedback pacification information.

[0147] The device provided by the embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects, which will not be described here again in the embodiment.

[0148] In a possible implementation, the voice output unit 403 is further configured to generate a feedback task corresponding to the delayed feedback pacification information, add the feedback task to a task queue, and if response information corresponding to a first feedback task in the task queue is received from the large model service device, output response voice corresponding to the response information.

[0149] The device provided by the embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects, which will not be described here again in the embodiment.

[0150] In a possible implementation, the voice output unit 403 is further configured to, in response to receiving an interaction start instruction sent by the earphone, output a wake-up sound and receive input voice of a user.

[0151] The device provided by the embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects, which will not be described here again in the embodiment.

[0152] In a possible implementation, the sound effect output unit 404 is further configured to, if the response information comprises operation information, execute an operation process corresponding to the operation information to obtain an operation result; and output an ending sound effect corresponding to the operation result.

[0153] The device provided in this embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects, which will not be described here again in this embodiment.

[0154] In a possible implementation, the sound effect output unit 404 is further configured to, if an interrupt instruction input by the user is received during output of the calming voice, interrupt the output of the calming voice; and output an interrupt sound effect.

[0155] The device provided in this embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects, which will not be described here again in this embodiment.

[0156] Referring to FIG. 5, a structural schematic diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure is shown, which can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, headsets, notebook computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. The electronic device shown in FIG. 5 is only an example, and should not limit the functions and use range of the embodiments of the present disclosure.

[0157] As shown in FIG. 5, the electronic device 500 can include a processing device (such as a central processor, a graphics processor, etc.) 501, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 502 or loaded into a random access memory (RAM) 503 from a storage device 508. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0158] In general, the following devices can be connected to the I / O interface 505: input devices 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 507 including, for example, a Liquid Crystal Display (LCD), a speaker, a vibrator, and the like; storage devices 508 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 509. The communication devices 509 can allow the electronic device 500 to communicate wirelessly or wired with other devices to exchange data. While FIG. 5 illustrates the electronic device 500 with various devices, it is understood that all of the illustrated devices are not required to be implemented or possessed. More or less devices can alternatively be implemented or possessed.

[0159] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 509, or installed from the storage devices 508, or installed from the ROM 502. When the computer program is executed by the processing devices 501, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0160] It should be noted that the computer-readable medium in the above disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.

[0161] The computer-readable medium described above can be contained in the electronic device described above; or can exist separately and not be assembled into the electronic device.

[0162] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0163] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0164] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0165] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the names of the units do not constitute a limitation on the units themselves. For example, the first obtaining unit can also be described as a unit for obtaining at least two Internet protocol addresses. The functions described above can be performed at least in part by one or more hardware logic components. For example, non-limiting examples of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.

[0166] The present disclosure further provides a computer readable storage medium, wherein computer execution instructions are stored, and when a processor executes the computer execution instructions, a technical solution of the voice response processing method in any of the above embodiments is realized, the implementation principle and beneficial effects of which are similar to those of the voice response processing method, and can be referred to the implementation principle and beneficial effects of the voice response processing method, which will not be described herein again.

[0167] The present disclosure further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the information display method of the first aspect and various possible designs of the first aspect

[0168] The above description is merely preferred embodiments of the present disclosure and a description of principles of applied technologies. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0169] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.

[0170] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A voice response processing method, wherein the method is applied to an electronic device, and the method comprises: recording a receiving time of an input voice of a user when it is detected that the input voice ends; sending voice information corresponding to the input voice to a large model service device, so that the large model service device generates response information corresponding to the voice information; in response to no response information sent by the large model service device being received within a first preset time period after the receiving time, outputting a pacification voice corresponding to pacification information, wherein the pacification information is sent by the large model service device.

2. The method of claim 1, wherein the response to no response information sent by the large model service device being received within a first preset time period after the receiving time, outputting a pacification voice corresponding to pacification information, comprises: if no response information sent by the large model service device is received within a first preset time period after the receiving time, sending a pacification information generation request to the large model service device, so that the large model service device generates pacification information corresponding to the voice information; receiving the pacification information sent by the large model service device; outputting a pacification voice corresponding to the pacification information.

3. The method of claim 1, wherein after the outputting a pacification voice corresponding to pacification information, further comprising: if no feedback information and the response information sent by the large model service device are received within a second preset time period after the outputting the pacification voice, outputting a pacification sound effect at a preset time interval, wherein the feedback information is used for the large model service device to confirm that the voice information has been received.

4. The method of claim 3, wherein after the outputting a pacification sound effect at a preset time interval, further comprising: counting the number of times of outputting the pacification sound effect; if the number of times of outputting is greater than a preset number of times of outputting threshold, outputting an overtime pacification voice.

5. The method of claim 3, wherein after the outputting a pacification sound effect at a preset time interval, further comprising: if no feedback information and the response information sent by the large model service device are still not received within a third preset time period after the outputting the pacification voice, outputting an overtime pacification voice, wherein the third preset time period is greater than the second preset time period.

6. The method of claim 1, wherein after the outputting a pacification voice corresponding to pacification information, further comprising: if the feedback information sent by the large model service device according to the voice information is received within a second preset time period after the outputting the pacification voice, but the response information sent by the large model service device is not received, outputting a pacification sound effect at a preset time interval until the response information sent by the large model service device is received, and outputting a response voice corresponding to the response information, wherein the feedback information is used for the large model service device to confirm that the voice information has been received.

7. The method of claim 1, wherein after the outputting a pacification voice corresponding to pacification information, further comprising: If the feedback information sent by the large model service device according to the voice information is received within a second preset time period after the soothing voice is output, but the response information is not received, a soothing sound effect is output at a preset time interval.

8. The method of claim 7, wherein after the soothing sound effect is output at the preset time interval, further comprising: counting the number of times the soothing sound effect is output; if the number of times the soothing sound effect is output is greater than or equal to a preset number of times and the response information is not received, sending a delayed feedback soothing generation request to the large model service device to cause the large model service device to generate delayed feedback soothing information according to the delayed feedback soothing generation request and the voice information; receiving the delayed feedback soothing information sent by the large model service device; outputting delayed feedback voice corresponding to the delayed feedback soothing information.

9. The method of claim 7, wherein after the soothing sound effect is output at the preset time interval, further comprising: if the response information is not received within a third preset time period after the soothing voice is output, sending a delayed feedback soothing generation request to the large model service device to cause the large model service device to generate delayed feedback soothing information according to the delayed feedback soothing generation request and the voice information, wherein the third preset time period is greater than the second preset time period; receiving the delayed feedback soothing information sent by the large model service device; outputting delayed feedback voice corresponding to the delayed feedback soothing information.

10. The method of claim 7 or 8, wherein after the delayed feedback voice corresponding to the delayed feedback soothing information is output, further comprising: generating a feedback task corresponding to the delayed feedback soothing information; adding the feedback task to a task queue; if the response information corresponding to the first feedback task in the task queue sent by the large model service device is received, outputting response voice corresponding to the response information.

11. The method of any one of claims 1 to 9, wherein before the time when the input voice of the user is detected to end, the method further comprises: in response to receiving an interaction start instruction sent by the earphone, outputting a wake-up sound and receiving the input voice of the user.

12. The method of any one of claims 1 to 9, wherein after the voice information corresponding to the input voice is sent to the large model service device, the method further comprises: if the response information includes operation information, performing an operation process corresponding to the operation information to obtain an operation result; outputting an end sound effect corresponding to the operation result.

13. The method of any one of claims 1 to 9, wherein after the voice information corresponding to the input voice is sent to the large model service device, the method further comprises: if an interrupt instruction input by the user is received during the output of the soothing voice, interrupting the output of the soothing voice; outputting an interrupt sound effect.

14. A voice response processing device, comprising: A time recording unit configured to record a receiving time of the input voice when detecting that the input voice of the user ends; An information sending unit configured to send voice information corresponding to the input voice to a large model service device, so that the large model service device generates response information corresponding to the voice information; A voice output unit configured to output a calming voice corresponding to calming information sent by the large model service device in response to no response information sent by the large model service device being received within a first preset time length after the receiving time.

15. An electronic device comprising: A processor and a memory; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the voice response processing method according to any one of claims 1 to 13.

16. A computer readable storage medium, the computer readable storage medium storing computer execution instructions, when the processor executes the computer execution instructions, the voice response processing method according to any one of claims 1 to 13 is realized.

17. A computer program product, comprising a computer program, when the computer program is executed by the processor, the voice response processing method according to any one of claims 1 to 13 is realized.

Citation Information

Patent Citations

  • Reply method and system for voice interaction

    CN112289317A

  • Voice interaction method and device, equipment and computer storage medium

    CN112382290A

  • Offline and online semantic recognition arbitration method, electronic equipment and storage medium

    CN112992145A

  • Data processing method and device and electronic equipment

    CN117612533A

  • Method and device for improving large language model reasoning concurrency and computing equipment

    CN118153693A