Information processing method and device, electronic equipment and computer readable storage medium
Patent Information
- Application Number
- CN202510821918.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-06-18
AI Technical Summary
[0003]但是,目前的语音交互系统还存在一些缺陷,例如,基于用户语音场景的不确定性,用户输入的语音完整性、语音真实性等有效性还有待提升,不准确的用户语音信息会导致无效、或错误的识别进而产生不达预期的交互反馈,影响用户体验
[0052] In a seventh aspect, embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in embodiments of this application.
Smart Images

Figure CN120726995B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to an information processing method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] With the rapid development of artificial intelligence technology, voice interaction systems have been widely used in smart homes, in-vehicle devices, customer service and other scenarios, enabling users to quickly control smart home devices or conduct convenient conversations through voice interaction.
[0003] However, current voice interaction systems still have some shortcomings. For example, due to the uncertainty of user voice scenarios, the effectiveness of user input voice completeness and authenticity needs to be improved. Inaccurate user voice information can lead to invalid or incorrect recognition, resulting in unsatisfactory interactive feedback and affecting user experience. Summary of the Invention
[0004] This application provides an information processing method, apparatus, electronic device, and computer-readable storage medium, which can improve the accuracy of intent recognition, the accuracy of interaction, and the user experience.
[0005] In a first aspect, embodiments of this application provide an information processing method applied to a server, the method comprising:
[0006] Obtain the current pending information sent by the terminal device for the target user;
[0007] The current information to be processed is matched with the skill domain according to the preset identification strategy to obtain the matching result. The preset identification strategy is optimized based on the historical feedback information of the target user.
[0008] The intent recognition result of the target user is determined based on the matching result;
[0009] The control command corresponding to the intent recognition result is executed to obtain the execution result.
[0010] Secondly, embodiments of this application also provide an information processing apparatus applied to a server, the apparatus comprising:
[0011] The acquisition module is used to acquire the current pending information for the target user sent by the terminal device;
[0012] The identification module is used to perform skill domain matching processing on the current information to be processed according to a preset identification strategy to obtain a matching result. The preset identification strategy is optimized based on the historical feedback information of the target user.
[0013] The determining module is used to determine the intent recognition result of the target user based on the matching result;
[0014] The execution module is used to execute the control instructions corresponding to the intent recognition result and obtain the execution result.
[0015] Optionally, in some embodiments of this application, determining the intent recognition result of the target user based on the matching result includes:
[0016] If the matching result includes no match with any of the skill domains, then no user intent will be taken as the intent recognition result.
[0017] If the matching result includes a partial match with any skill domain, then the following information to be processed corresponding to the current information to be processed is obtained, and the current information to be processed and the following information to be processed are fused to obtain the target information to be processed. Furthermore, the target information to be processed is subjected to intent recognition to obtain the user intent, and the user intent is used as the intent recognition result.
[0018] Optionally, in some embodiments of this application, after executing the control instruction corresponding to the intent recognition result and obtaining the execution result, the method further includes:
[0019] Generate voice prompts for the execution result, which are output by the terminal device to provide feedback on the execution result;
[0020] If the terminal device is currently outputting first audio information, then output control information for the voice prompt information and the first audio information is generated according to a preset output strategy;
[0021] The output control information is sent to the terminal device, wherein the terminal device controls the output of the voice prompt information and the first audio information according to the output control information;
[0022] The preset output strategy includes a fixed output strategy and a dynamic output strategy for different application scenarios. The dynamic output strategy is obtained by optimizing the initial output strategy based on the historical feedback information.
[0023] Optionally, in some embodiments of this application, the step of generating output control information for the voice prompt information and the first audio information according to a preset output strategy if the terminal device is currently outputting first audio information includes:
[0024] If the terminal device is currently outputting first audio information, then the first output information of the first audio information is determined according to the preset output strategy, and the second output information of the first audio information is determined based on the first output information;
[0025] The first output information and the second output information are used as the output control information;
[0026] The first output information includes at least one of uninterruptible, interruptible and recoverable, or interruptible and unrecoverable, and the second output information includes at least one of direct output or subsequent output.
[0027] Optionally, in some embodiments of this application, the preset identification strategy includes a first identification strategy;
[0028] The step of performing skill domain matching processing on the current information to be processed according to a preset recognition strategy to obtain a matching result includes:
[0029] The current information to be processed is input into the intelligent agent, and the intelligent agent performs skill domain matching processing on the current information to be processed according to the first recognition strategy to obtain the matching result.
[0030] Before performing skill domain matching processing on the current information to be processed according to a preset recognition strategy to obtain a matching result, the method further includes:
[0031] The current information to be processed is subjected to skill domain matching processing using a second identification strategy to obtain preliminary identification results;
[0032] Based on the preliminary identification results, the step of performing skill domain matching processing on the current information to be processed according to the preset identification strategy to obtain the matching results is controlled.
[0033] Optionally, in some embodiments of this application, the step of inputting the current information to be processed into the intelligent agent, and then performing skill domain matching processing on the current information to be processed according to the first recognition strategy to obtain a matching result, includes:
[0034] The current information to be processed is input into the intelligent agent, which then clarifies the current information to be processed based on the target user's historical feedback information and user characteristic information to obtain the processed information.
[0035] The processed information is subjected to skill domain matching according to a preset recognition strategy to obtain the matching result;
[0036] The user characteristic information includes at least one of user preference information or user commonly used phrases information;
[0037] The clarification process includes at least one of completion processing or error correction processing.
[0038] Thirdly, embodiments of this application provide an information processing method applied to a terminal device, the method comprising:
[0039] If the target user's current pending information is collected, the current pending information is sent to the server;
[0040] Receive the execution feedback result returned by the server for the current pending information, the execution feedback result being generated based on the server's execution result for the current pending information;
[0041] Output the execution feedback result;
[0042] The server performs skill domain matching on the current information to be processed according to a preset identification strategy to obtain a matching result, determines the intent recognition result of the target user based on the matching result, and executes the control command corresponding to the intent recognition result to obtain the execution result. The preset identification strategy is obtained by the server based on the historical feedback information of the target user.
[0043] Fourthly, embodiments of this application provide an information processing apparatus applied to a terminal device, the apparatus comprising:
[0044] The sending module is used to send the current pending information of the target user to the server if it collects the current pending information of the target user.
[0045] The receiving module is configured to receive the execution feedback result returned by the server for the current pending information, wherein the execution feedback result is generated based on the execution result of the server on the current pending information;
[0046] The output module is used to output the execution feedback results;
[0047] The server performs skill domain matching on the current information to be processed according to a preset identification strategy to obtain a matching result, determines the intent recognition result of the target user based on the matching result, and executes the control command corresponding to the intent recognition result to obtain the execution result. The preset identification strategy is obtained by the server based on the historical feedback information of the target user.
[0048] Optionally, in some embodiments of this application, the current information to be processed includes voice information to be processed, and the method further includes:
[0049] When the voice information to be processed is detected, if the terminal device is currently outputting second audio information, the output volume of the second audio information is reduced, and the output volume of the second audio information is restored when the voice information to be processed is received.
[0050] Fifthly, embodiments of this application also provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the information processing method described above.
[0051] Sixthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the information processing method described above.
[0052] In a seventh aspect, embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in embodiments of this application.
[0053] In summary, the server in this application embodiment obtains the current pending information for the target user sent by the terminal device, performs skill domain matching processing on the current pending information according to a preset identification strategy to obtain a matching result. The preset identification strategy is optimized based on the historical feedback information of the target user. The server determines the intention recognition result of the target user based on the matching result, executes the control command corresponding to the intention recognition result, and obtains the execution result.
[0054] Specifically, by performing skill domain matching on the current information to be processed before analyzing the user's intent, and determining the target user's intent recognition result based on the matching result, the accuracy of user intent analysis is improved, thereby improving the accuracy of interaction processing based on intent recognition.
[0055] Among these measures, by optimizing the skill domain matching process based on users' historical feedback information, the matching degree between intent recognition results and user expectations is improved, thereby further enhancing the accuracy of user intent analysis and improving the user experience. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a schematic diagram illustrating a scenario where a server executes the information processing method according to an embodiment of this application;
[0058] Figure 2 This is a flowchart illustrating the information processing method provided in an embodiment of this application;
[0059] Figure 3 This is a schematic diagram of data flow in the information processing method provided in the embodiments of this application;
[0060] Figure 4 This is another flowchart illustrating the information processing method provided in the embodiments of this application;
[0061] Figure 5 This is a schematic diagram of the structure of the information processing device provided in the embodiments of this application;
[0062] Figure 6 This is another structural schematic diagram of the information processing device provided in the embodiments of this application;
[0063] Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0064] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0065] This application provides an information processing method, apparatus, electronic device, and computer-readable storage medium. Specifically, this application provides an information processing apparatus suitable for electronic devices, which include terminal devices or servers. The terminal devices include, but are not limited to, mobile phones, tablets, laptops, smart TVs, smart speakers, and other devices with voice assistants. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The server can be directly or indirectly connected via wired or wireless communication.
[0066] For example, please see Figure 1 , Figure 1This is a schematic diagram illustrating a scenario where a server executes the information processing method according to an embodiment of this application. The process by which a terminal device executes the information processing method can be understood by referring to the execution process of the server. Specifically, the specific execution process of the server executing the information processing method is as follows:
[0067] Server 101 obtains the current pending information for the target user sent by terminal device 102, performs skill domain matching processing on the current pending information according to a preset recognition strategy, and obtains a matching result. The preset recognition strategy is optimized based on the target user's historical feedback information. Based on the matching result, the server determines the target user's intent recognition result, executes the control command corresponding to the intent recognition result, and obtains an execution result.
[0068] For example, a user issues a voice command to control a terminal device. The terminal device collects the voice command to obtain the current information to be processed. Then, the terminal device sends the current information to the server. The server performs skill domain matching processing on the current information to be processed according to a preset recognition strategy to obtain a matching result. Based on the matching result, the server determines the target user's intent recognition result and executes the control operation corresponding to the intent recognition result to obtain the execution result.
[0069] Subsequently, the server generates an execution feedback result for the execution outcome and sends it to the terminal device, which then outputs the execution feedback result. This execution feedback result is used to declare or represent the execution result. For example, it can explain the current execution status through voice or text. For instance, if the current pending information includes "turn on the fresh air system," then a "fresh air system is on" prompt will be output.
[0070] In summary, the embodiments of this application improve the accuracy of user intent analysis by performing skill domain matching processing on the current information to be processed before analyzing the user's intent based on the current information to be processed, and determining the intent recognition result of the target user based on the matching result, thereby improving the accuracy of interaction processing based on intent recognition.
[0071] Among these measures, by optimizing the skill domain matching process based on users' historical feedback information, the matching degree between intent recognition results and user expectations is improved, thereby further enhancing the accuracy of user intent analysis and improving the user experience.
[0072] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the priority of the embodiments.
[0073] Please see Figure 2 , Figure 2This is a flowchart illustrating an information processing method provided in an embodiment of this application. Although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown in the flowchart. Specifically, the executing entity of this information processing method includes a server, and the specific flow of the information processing method is as follows:
[0074] 201. Obtain the current pending information for the target user sent by the terminal device.
[0075] It is understandable that the currently pending information is the information currently received and used as the basis for analyzing user intent. For example, the currently pending information may be instruction information for device control or popular science information.
[0076] The information to be processed may include information based on text stream or audio stream format. For example, in this embodiment of the application, an audio stream is used as an example. The user issues an instruction to the terminal device through voice, and the terminal device collects the user's voice to obtain the information to be processed.
[0077] The terminal device can be a device with audio acquisition capabilities and the ability to connect to the server for data transmission. For example, the terminal device includes devices with voice assistant functions such as mobile phones, tablets, smart TVs, or smart speakers.
[0078] It should be noted that, in this embodiment, the terminal device also supports full-duplex communication, which enables simultaneous voice input and output without interference. This technology allows users and voice assistants to listen and speak simultaneously during conversations and supports continuous dialogue without the need for repeated wake words. For example, if a human voice is detected while the terminal device is currently outputting audio, the user's voice can be captured while the audio output continues. That is, the user's voice input is captured simultaneously with audio output.
[0079] In this embodiment, taking a device control scenario as an example, for instance, during video playback, the device receives a user's voice instruction to play related music. Similarly, this full-duplex communication can be applied in various scenarios, such as weather inquiries and casual conversations, enabling real-time bidirectional processing of user voice input while simultaneously outputting audio.
[0080] 202. Perform skill domain matching processing on the current information to be processed according to the preset identification strategy to obtain the matching result. The preset identification strategy is optimized based on the historical feedback information of the target user.
[0081] In this context, a skill domain refers to a category of tasks or functions that can be processed. Each skill domain corresponds to a set of predefined intents and operational logic, used to fulfill a specific type of user request. For example, a skill domain is the various functions that can be executed in response to different user voice commands.
[0082] It should be noted that skill domain matching processing refers to matching the currently pending information with various skill domains. The purpose is to determine the user instruction corresponding to the currently pending information, that is, to determine the task or function to be executed. The matching result refers to the result of matching the currently pending information with each skill domain. In this embodiment, the matching result includes no match with any skill domain, partial match with one of the skill domains, etc. Specifically, no match with any skill domain means that the currently pending information does not match any skill domain. For example, the current skill domain does not contain the control operation corresponding to the currently pending information. Partial match with one of the skill domains means that the currently pending information partially matches one of the skill domains, but not completely. For example, the currently pending information is incomplete, which makes it impossible to accurately determine the operation logic to be executed.
[0083] The preset recognition strategy refers to the pre-set strategy used for skill domain matching of the information to be processed. That is, the specific strategy for matching the current information to be processed with each skill domain affects the matching result. For example, the matching result can be obtained by performing skill domain matching processing through an agent based on a large language model, or by extracting keywords from the current information to be processed and judging the matching result by matching the keywords with each skill domain.
[0084] In this embodiment, the preset identification strategy is optimized based on the target user's historical feedback information from previous stages, so that the matching results of the skill domain matching process better meet the user's expectations. For example, if the user is dissatisfied with the result of the operation after the previous skill domain matching process and execution of the corresponding operation logic, the preset identification strategy is adjusted, thereby affecting or changing the current matching result.
[0085] For example, if the predefined skill domain is for "encyclopedia of XX", and the user's input voice can only recognize "YY", then it's a rejection scenario. If the user's input voice can recognize "XX", then it's a non-recognition scenario. However, if the user requests "XX", "encyclopedia of XX", or "introduction to XX" again, it indicates that the matching result of the previous skill domain matching process does not match the user's expectations. In this case, the skill domain matching strategy is adjusted so that the skill domain "encyclopedia of XX" or "introduction to XX" can be matched by "XX".
[0086] 203. Determine the intent recognition result of the target user based on the matching result.
[0087] Since each skill domain corresponds to a user's intent, matching a skill domain indicates that the user's intent has been clarified, thus yielding an intent recognition result. If no skill domain is matched, it means the user's intent is still unclear, requiring the user to provide further instructions or combine more instructions to determine the intent.
[0088] Among these methods, determining the target user's intent based on the matching results improves the accuracy of intent recognition.
[0089] 204. Execute the control command corresponding to the intent recognition result to obtain the execution result.
[0090] The control command instructs the server to perform the action. This control command corresponds to the control domain. For example, if the matched control domain is "turn on the fresh air", then the control command corresponding to the user intent recognition result is also "turn on the fresh air".
[0091] Specifically, by executing control commands corresponding to the intent recognition results, feedback to user instructions is provided, enabling interaction between the server and the user.
[0092] The execution result refers to the result of executing the control command. For example, executing the control command "turn on fresh air" will result in the air conditioner turning on fresh air.
[0093] Correspondingly, in some scenarios, the server will also feed back the execution result to the terminal device so that the terminal device can inform the user of the current device execution status. For example, the server can generate a voice message "Air conditioner and fresh air have been turned on" and send the voice message to the terminal device. After receiving the voice message, the terminal device can play the voice message.
[0094] In summary, the embodiments of this application improve the accuracy of user intent analysis by performing skill domain matching processing on the current information to be processed before analyzing the user's intent based on the current information to be processed, and determining the intent recognition result of the target user based on the matching result, thereby improving the accuracy of interaction processing based on intent recognition.
[0095] Among these measures, by optimizing the skill domain matching process based on users' historical feedback information, the matching degree between intent recognition results and user expectations is improved, thereby further enhancing the accuracy of user intent analysis and improving the user experience.
[0096] It is understandable that if the current pending information does not match any of the skill domains, it indicates that even if the current pending information is responded to, the response will be incorrect due to the difficulty in matching the accurate skill domain. Therefore, the response to the current pending information can be actively rejected, i.e., rejection. Alternatively, if the current pending information is incomplete and cannot accurately identify the user's intent, or if the necessary parameters are missing when determining the operation logic, the user's complete expression can be obtained by combining the following methods, thereby obtaining the accurate user intent and operation logic. That is, optionally, in some embodiments of this application, the step "determining the intent recognition result of the target user based on the matching result" includes:
[0097] If the matching result includes no match with any of the skill domains, then no user intent will be taken as the intent recognition result.
[0098] If the matching result includes a partial match with any skill domain, then the following information to be processed corresponding to the current information to be processed is obtained, and the current information to be processed and the following information to be processed are fused to obtain the target information to be processed. Furthermore, the target information to be processed is subjected to intent recognition to obtain the user intent, and the user intent is used as the intent recognition result.
[0099] If the information to be processed does not match any of the skill domains, a rejection strategy is adopted, that is, the response to the information to be processed is refused, and the absence of user intent is taken as the intent recognition result. No user intent means that no control command needs to be executed. For example, if meaningless voice information (such as a cough) is collected, it indicates that the user has no actual valid intent and does not need to respond to the voice information.
[0100] If the current information to be processed partially matches one of the skill domains, it indicates that the current information to be processed cannot accurately match a specific and clear skill domain. In this case, the decision-making-not-stop strategy is adopted, that is, to continue collecting the user's voice, that is, to fuse the following information to be processed with the current information to be processed to obtain the complete target processing information, and then determine the skill domain corresponding to the target information to be processed and perform intent recognition, etc.
[0101] The "to be processed following information" refers to the information collected after the "to be processed currently information." For example, if a user speaks slowly, the first half of the voice is the "to be processed currently information," and the second half is the "to be processed following information." Alternatively, if the user's initial input is a clear statement of intent but lacks the necessary parameters for execution, the information the user supplements through full-duplex communication is the "to be processed following information." For instance, if the user's initial statement is "book a flight," the intent is recognized as opening a booking application and needing to book a ticket. However, due to the lack of destination and time—essential parameters for execution—the process cannot proceed. In this case, the system can either ask the user a question or wait for their subsequent input to specify the destination and time. This re-input of destination and time is the "to be processed following information."
[0102] It is understandable that after the current information to be processed is matched with each skill domain, it may also include the case of matching with a certain skill domain. That is, the corresponding skill domain can be clearly identified. In this case, the intent of the user can be directly determined by intent recognition of the current information to be processed. For example, the intent corresponding to the skill domain can be used as the intent recognition result.
[0103] It should be noted that, in this embodiment, when performing skill domain matching, the current information to be processed is first analyzed to determine the corresponding user intent. This user intent is then matched against various skill domains. If it matches one skill domain, it indicates a match; if it doesn't match any of them, it indicates no match at any skill domain; if the consistency with one or more skill domains is low or incomplete, it indicates a partial match with any skill domain. Accordingly, in cases of consistency, the user intent is directly used as the intent recognition result; in cases of inconsistency, it is considered a rejection, indicating no user intent; and in cases of low or incomplete consistency, the user intent is determined by combining this information with the information to be processed below.
[0104] In summary, by performing skill domain matching on the current information to be processed, the system can determine the effective user intent recognition result, avoiding invalid or erroneous responses caused by invalid user intents, i.e., avoiding the execution of incorrect instructions based on invalid user intents. Furthermore, by optimizing the skill domain matching processing strategy based on user feedback, intent recognition becomes more accurate and better meets user expectations.
[0105] Skill domain matching processing also helps overcome the low recognition accuracy problem caused by echo cancellation. Echo cancellation (AEC) is an audio signal processing technique designed to reduce or eliminate echo phenomena in audio communication. Echoes typically occur when sound from a speaker is picked up by a microphone and retransmitted to a distant location, causing the user at the remote end to hear their own voice. Echo cancellation technology ensures the clarity and quality of audio communication by identifying and removing echo components from the input signal. However, when echo cancellation is ineffective, the currently processed information transmitted to the remote end (such as the server in this embodiment) may also contain echoes. Skill domain matching processing can further identify these echoes. For example, voice prompts often begin with "Already for you." Therefore, if the currently processed information contains "Already for you," it indicates that the information contains an echo, meaning it is not the user's actual voice. In this case, skill domain matching processing identifies it as a rejection scenario, thereby improving speech recognition accuracy.
[0106] It's understandable that when outputting content (such as text, audio, image cards, etc.) on a terminal device, if there is already content being output (such as text, audio, image cards, etc.), two pieces of content may be output simultaneously. This simultaneous output can interfere with each other, affecting the user experience. For example, due to the characteristics of full-duplex communication, even when the terminal device is currently outputting audio, it can still receive new voice instructions from the user. Based on these new voice instructions, a new task will be generated. For instance, if the new task is to output a piece of audio, two audio files may need to be output simultaneously, which will also affect the user's auditory experience.
[0107] Therefore, in this application embodiment, a corresponding output control strategy is designed for scenarios with simultaneous output, in order to improve the user experience when two pieces of content are output simultaneously. That is, optionally, in some embodiments of this application, after the step "execute the control instruction corresponding to the intent recognition result to obtain the execution result", the method further includes:
[0108] Generate voice prompts for the execution result, which are output by the terminal device to provide feedback on the execution result;
[0109] If the terminal device is currently outputting first audio information, then output control information for the voice prompt information and the first audio information is generated according to a preset output strategy;
[0110] The output control information is sent to the terminal device, wherein the terminal device controls the output of the voice prompt information and the first audio information according to the output control information;
[0111] The preset output strategy includes a fixed output strategy and a dynamic output strategy for different application scenarios. The dynamic output strategy is obtained by optimizing the initial output strategy based on the historical feedback information.
[0112] Among them, voice prompts are feedback information on the execution result. For example, if the execution result is that the air conditioner has turned on the fresh air, the voice prompt will include "The air conditioner has turned on the fresh air", reminding the user through voice broadcast.
[0113] The first audio information is the audio currently being output by the terminal device. For example, when operating an air conditioner to turn on the fresh air, based on the full-duplex characteristic, the terminal device is also simultaneously outputting music or video works, and the audio or video works being output is the first audio information.
[0114] The preset output strategy is a control strategy for the output of voice prompts and the first audio information. The output control information is determined based on the preset output strategy and is used to control the output of the voice prompts and the first audio information. For example, the output control information includes whether both are output simultaneously or whether the first audio information is interrupted.
[0115] In this application embodiment, the preset output strategy includes output strategies for different application scenarios, such as output strategies for danger warning scenarios (e.g., earthquake early warning), output strategies for life reminder scenarios (e.g., alarm clock), output strategies for video playback scenarios, and output strategies for audio playback scenarios.
[0116] Correspondingly, to ensure stability in certain scenarios, a fixed output strategy can be set for certain scenarios, meaning that the corresponding output strategy in these scenarios remains unchanged; and to ensure flexibility in certain scenarios and meet the changing needs of users, a dynamic output strategy can be set for certain scenarios, meaning that the corresponding output strategy in these scenarios is dynamically adjusted based on the user's historical feedback information in order to meet the changing needs of users in a timely manner.
[0117] For example, danger alert scenarios, daily life alert scenarios, and video playback scenarios can be set to fixed output strategies, corresponding to uninterruptible, interruptible, and interruptible output strategies, respectively. Meanwhile, audio playback scenarios and device-executed feedback prompt scenarios can be set to dynamic output strategies, adjusting the default output strategy based on user feedback to meet the dynamic changes in user needs.
[0118] The dynamic output strategy can be adjusted based on the device's response and feedback. For example, if it's the user's first time using the device, a default strategy will be used: the terminal will interrupt and discard the old audio stream when a new one is received. If the user requests the previously interrupted audio stream again while a new one is being played, it means the user disagrees with the interruption and is requesting it again. In this case, the user's behavior will be recorded, and the output strategy for the previously interrupted audio stream will be adjusted; for example, the output strategy for that old audio stream will be changed to uninterruptible.
[0119] Optionally, in this embodiment, the output strategy of the first audio information can be determined by a preset output strategy and based on the scenario corresponding to the first audio information, thereby determining the output strategy of the voice prompt information. That is, optionally, in some embodiments of this application, the step "if the terminal device is currently outputting the first audio information, then generate output control information for the voice prompt information and the first audio information according to the preset output strategy" includes:
[0120] If the terminal device is currently outputting first audio information, then the first output information of the first audio information is determined according to the preset output strategy, and the second output information of the first audio information is determined based on the first output information;
[0121] The first output information and the second output information are used as the output control information;
[0122] The first output information includes at least one of uninterruptible, interruptible and recoverable, or interruptible and unrecoverable, and the second output information includes at least one of direct output or subsequent output.
[0123] For example, the application scenario corresponding to the first audio information is identified, and the output strategy for that application scenario is determined according to a preset output strategy. Then, the output strategy for the voice prompt information is determined based on this output strategy. For instance, if the application scenario corresponding to the first audio information is a video playback scenario, then according to the preset output strategy for video playback scenarios (interruptible), the first output information of the first audio information is determined to be interrupted, and the second output information of the voice prompt information is determined to be directly output. As another example, if the first video information is an earthquake alert, its corresponding output strategy is non-interruptible. Therefore, the broadcast of the earthquake alert information is maintained, while the broadcast of the voice prompt information is paused or the volume of the voice prompt information is reduced. In some options, the voice prompt information can also be converted into a text prompt and displayed outside the earthquake alert information card pop-up window.
[0124] In this embodiment of the application, the application scenarios of the voice prompt information and the first audio information can be identified respectively, and the respective output strategies can be determined according to the two application scenarios.
[0125] In this embodiment, the output priority for each application scenario can be set. For example, if multiple user intentions are recognized based on the user's voice, there may be situations where two or more voice prompts or audio messages are output simultaneously (e.g., the user's intention is to play a video and play the theme song from that video). In this case, the output order is controlled based on the output priority of each content to be output (multiple voice prompts or multiple audio messages corresponding to the intention). For example, if there is both a voice prompt for the device execution result and a request to play a movie, the movie can be played first, and then the voice prompt can be broadcast.
[0126] In some scenarios, based on output priority, output control information can be combined to control the output of multiple content items to be output. For example, if a film or television work has a higher priority, it will be output first. When it is determined that a voice prompt message should be output, the first output information of the film or television work and the second output information of the voice prompt message are determined. The output of the film or television work and the voice prompt message are controlled according to the first and second output information. For example, the film or television work can be played first, and then the voice prompt message can be played at a low volume without interrupting the playback of the film or television work.
[0127] Furthermore, in this embodiment, corresponding recovery mechanisms can be configured for interruptible and non-interruptible output strategies to determine whether a recovery control strategy is needed after an interruption. For example, if the first audio information is a currently playing video, and it is determined based on the output control information that the first audio information needs to be interrupted, then the playback of the first audio information is interrupted until the voice prompt information is finished playing. After the voice prompt information is finished playing, the first audio information determines whether it can resume playback based on the recovery mechanism. For example, if the recovery strategy for the video is pre-configured as recovery, then the resume playback function is executed after the voice prompt information is finished playing, that is, the playback of subsequent video works continues.
[0128] In this embodiment, an agent based on a large language model can be used to perform skill domain matching processing on the current information to be processed. Specifically, optionally, in some embodiments of this application, the preset recognition strategy includes a first recognition strategy, and the step "performing skill domain matching processing on the current information to be processed according to the preset recognition strategy to obtain a matching result" includes:
[0129] The current information to be processed is input into the intelligent agent, and the intelligent agent performs skill domain matching processing on the current information to be processed according to the first recognition strategy to obtain the matching result.
[0130] Before the step "perform skill domain matching processing on the current information to be processed according to the preset recognition strategy to obtain the matching result", the method further includes:
[0131] The current information to be processed is subjected to skill domain matching processing using a second identification strategy to obtain preliminary identification results;
[0132] Based on the preliminary identification results, the step of performing skill domain matching processing on the current information to be processed according to the preset identification strategy to obtain the matching results is controlled.
[0133] In this embodiment, the recognition strategy for skill domain matching of the intelligent agent is modified based on user feedback information. For example, the parameters or weights of the intelligent agent during skill domain matching are changed. For instance, labeled sample data is generated based on user feedback information, and a new intelligent agent is trained again based on this sample data.
[0134] The preliminary identification result is a matching result obtained based on the second identification strategy. This preliminary identification result includes no match with any skill domain, partial match with any skill domain, or match with a specific skill domain. The second identification strategy is also used to match the current information to be processed with skill domains. It should be noted that the difference between the second identification strategy and the first identification strategy is that the first identification strategy is generated and executed by the intelligent agent, while the second identification strategy is processed by the central control or other processing units. For example, the second identification strategy can obtain the preliminary identification result through keyword matching, regular expressions, etc. For instance, if the current information to be processed is a single word, a meaningless pause word, is not a command or question to the voice assistant, is mis-entered non-natural language text, or is a script broadcast by the voice assistant in the previous round that was mistakenly uploaded to the cloud due to terminal echo cancellation failure, then the preliminary identification result is no match with any skill domain. If the current information to be processed is part of any skill domain, then the preliminary identification result is partial match with any skill domain. If the current information to be processed is consistent with the description of any skill domain, then the preliminary identification result is a match with a specific skill domain.
[0135] After obtaining the preliminary identification results, the invocation and processing of the intelligent agent are controlled based on these results. For example, please refer to [link to relevant documentation]. Figure 3 , Figure 3This is a schematic diagram of data flow in the information processing method provided in this application embodiment, which includes a user side, a terminal device, and a server. The server includes an IO cloud, a central control service, an ASR module, an LLM module, and a TTS module. The ASR module is used for audio-to-text conversion, and the TTS module is used for text-to-audio conversion. Specifically, the data flow process includes:
[0136] 301. User-side speech generation;
[0137] 302. The voice is collected through the terminal device to obtain the current information to be processed, which is in the form of an audio stream, and the current information to be processed is sent to the central control service of the server;
[0138] 303. Call the ASR service through the central control service, i.e., start the ASR module;
[0139] 304. Convert the audio stream into a text stream through the ASR service, that is, obtain the current pending information in text format, and return the current pending information in text format to the central control service.
[0140] 305. The central control service performs skill domain matching on the current information to be processed in the text format according to the second recognition strategy to obtain preliminary recognition results;
[0141] 306. The central control service controls the invocation of the LLM module based on the preliminary identification results;
[0142] 307. If the LLM module is called, the skill domain matching process is performed by the LLM module according to the first recognition strategy to obtain the matching result;
[0143] For example, if the initial identification result is a rejection (no match with any skill domain), then the LLM module does not need to be called; if it is an indeterminate case (partially matches one or more skill domains but not completely), then the comprehensive information is obtained by combining the following text, and the LLM module is called to perform skill domain matching processing on the comprehensive information through the first identification strategy; if it is a normal case (matches a certain skill domain), then the current information to be processed is sent to the LLM module, so that the LLM module can perform skill domain matching processing on the current information to be processed through the first identification strategy.
[0144] 308. Determine the intent recognition result based on the matching result;
[0145] 309. Obtain the execution result by executing the control command corresponding to the intent recognition result through IO cloud;
[0146] 310. Display the execution results on the terminal device;
[0147] For example, the air conditioner switches to fresh air mode or the audio system plays music based on the user's voice commands.
[0148] 311. The central control server generates feedback characters by combining the execution results and the corresponding skill domain;
[0149] 312. The feedback characters are converted into speech through the TTS module to obtain voice prompt information, which is then distributed to the terminal devices in sequence through the central control service;
[0150] 313. The terminal device outputs the voice prompt information and responds to the user's feedback on the voice prompt information.
[0151] For example, if the user feedback is unsatisfactory, the current action of the terminal device is interrupted, such as stopping the operation of the fresh air mode or stopping the music playback.
[0152] In summary, before the agent in the LLM module performs skill domain matching according to the first recognition strategy, the central control service performs skill domain matching according to the second recognition strategy, which reduces the calls to the agent and avoids resource waste caused by frequent and invalid calls to the agent.
[0153] In this embodiment of the application, the intelligent agent can also clarify the information of skill domain matching based on the user's historical data to improve the accuracy of the matching results. That is, optionally, in some embodiments of this application, the step "inputting the current information to be processed into the intelligent agent, and performing skill domain matching processing on the current information to be processed by the intelligent agent according to the first recognition strategy to obtain the matching result" includes:
[0154] The current information to be processed is input into the intelligent agent, which then clarifies the current information to be processed based on the target user's historical feedback information and user characteristic information to obtain the processed information.
[0155] The processed information is subjected to skill domain matching according to a preset recognition strategy to obtain the matching result;
[0156] The user characteristic information includes at least one of user preference information or user commonly used phrases information;
[0157] The clarification process includes at least one of completion processing or error correction processing.
[0158] For example, based on users' historical feedback, preferences, and frequently used phrases, the information to be processed is completed and corrected to obtain the processed information. Then, skill domain matching is performed on the processed information to improve the accuracy of the matching results. For instance, an intelligent agent can complete "today's sky" to "today's weather" and correct "the wind speed is a little higher" to "the wind speed is a little higher," etc.
[0159] In this embodiment, the agent is based on a large language model and uses the SFT (Supervised Fine-Tuning) method to train a low-rank adaptation (LoRA) model in combination with business scenarios. This model can realize the intent understanding function of a sentence with multiple meanings, so as to perform the task of skill domain matching and the task of determining the user's intent recognition result.
[0160] In summary, the embodiments of this application improve the accuracy of user intent analysis by performing skill domain matching processing on the current information to be processed before analyzing the user's intent based on the current information to be processed, and determining the intent recognition result of the target user based on the matching result, thereby improving the accuracy of control processing based on intent recognition.
[0161] Among these measures, by optimizing the skill domain matching process based on users' historical feedback information, the matching degree between intent recognition results and user expectations is improved, thereby further enhancing the accuracy of user intent analysis and improving the user experience.
[0162] Among them, the output strategy is controlled by preset output strategy when multiple contents are output, thereby improving the user experience in the case of multiple outputs.
[0163] Accordingly, in order to more clearly understand the information processing method of the embodiments of this application, the information processing method of the embodiments of this application will be described from the perspective of the terminal device. For example, please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is another flowchart illustrating the information processing method provided in this application embodiment. Specifically, the information processing method includes:
[0164] 401. If the target user's current pending information is collected, the current pending information is sent to the server;
[0165] 402. Receive the execution feedback result returned by the server for the current pending information, wherein the execution feedback result is generated based on the execution result of the server on the current pending information;
[0166] 403. Output the execution feedback results;
[0167] The server performs skill domain matching on the current information to be processed according to a preset identification strategy to obtain a matching result, determines the intent recognition result of the target user based on the matching result, and executes the control command corresponding to the intent recognition result to obtain the execution result. The preset identification strategy is obtained by the server based on the historical feedback information of the target user.
[0168] The execution feedback results include voice prompts, which are used to inform the user of the current execution results via voice.
[0169] That is, the terminal device collects the user's voice information, then sends the voice information to the server, the server processes the voice information to obtain the execution result, and the server also generates execution feedback information (e.g., voice prompt information) based on the execution result, and the terminal device outputs the execution feedback information.
[0170] Since the terminal device supports full-duplex communication, it can simultaneously collect user voice and output audio. If the collected information to be processed is in voice format, the volume of the currently output audio is reduced to ensure the accuracy of the collected information, and the volume of the audio is restored after the collection is completed. Optionally, in some embodiments of this application, the information to be processed includes voice information to be processed. The method further includes:
[0171] When the voice information to be processed is detected, if the terminal device is currently outputting second audio information, the output volume of the second audio information is reduced, and the output volume of the second audio information is restored when the voice information to be processed is received.
[0172] For example, if the terminal device detects that the user has started a new round of voice requests during voice broadcast, the terminal device will automatically lower the volume until the user's voice request ends, and then adjust it back to the original volume.
[0173] To facilitate better implementation of the information processing method of this application, this application also provides an information processing apparatus based on the above-described information processing method. The meanings of the terms used are the same as in the information processing method described above, and specific implementation details can be found in the descriptions of the method embodiments.
[0174] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of the information processing device provided in the embodiment of this application, wherein the information processing device is applied to a server, and specifically as follows:
[0175] The acquisition module 501 is used to acquire the current pending information for the target user sent by the terminal device;
[0176] The identification module 502 is used to perform skill domain matching processing on the current information to be processed according to a preset identification strategy to obtain a matching result. The preset identification strategy is optimized based on the historical feedback information of the target user.
[0177] The determining module 503 is used to determine the intent recognition result of the target user based on the matching result;
[0178] The execution module 504 is used to execute the control instructions corresponding to the intent recognition result and obtain the execution result.
[0179] Optionally, in some embodiments of this application, determining the intent recognition result of the target user based on the matching result includes:
[0180] If the matching result includes no match with any of the skill domains, then no user intent will be taken as the intent recognition result.
[0181] If the matching result includes a partial match with any skill domain, then the following information to be processed corresponding to the current information to be processed is obtained, and the current information to be processed and the following information to be processed are fused to obtain the target information to be processed. Furthermore, the target information to be processed is subjected to intent recognition to obtain the user intent, and the user intent is used as the intent recognition result.
[0182] Optionally, in some embodiments of this application, after executing the control instruction corresponding to the intent recognition result and obtaining the execution result, the method further includes:
[0183] Generate voice prompts for the execution result, which are output by the terminal device to provide feedback on the execution result;
[0184] If the terminal device is currently outputting first audio information, then output control information for the voice prompt information and the first audio information is generated according to a preset output strategy;
[0185] The output control information is sent to the terminal device, wherein the terminal device controls the output of the voice prompt information and the first audio information according to the output control information;
[0186] The preset output strategy includes a fixed output strategy and a dynamic output strategy for different application scenarios. The dynamic output strategy is obtained by optimizing the initial output strategy based on the historical feedback information.
[0187] Optionally, in some embodiments of this application, the step of generating output control information for the voice prompt information and the first audio information according to a preset output strategy if the terminal device is currently outputting first audio information includes:
[0188] If the terminal device is currently outputting first audio information, then the first output information of the first audio information is determined according to the preset output strategy, and the second output information of the first audio information is determined based on the first output information;
[0189] The first output information and the second output information are used as the output control information;
[0190] The first output information includes at least one of uninterruptible, interruptible and recoverable, or interruptible and unrecoverable, and the second output information includes at least one of direct output or subsequent output.
[0191] Optionally, in some embodiments of this application, the preset identification strategy includes a first identification strategy;
[0192] The step of performing skill domain matching processing on the current information to be processed according to a preset recognition strategy to obtain a matching result includes:
[0193] The current information to be processed is input into the intelligent agent, and the intelligent agent performs skill domain matching processing on the current information to be processed according to the first recognition strategy to obtain the matching result.
[0194] Before performing skill domain matching processing on the current information to be processed according to a preset recognition strategy to obtain a matching result, the method further includes:
[0195] The current information to be processed is subjected to skill domain matching processing using a second identification strategy to obtain preliminary identification results;
[0196] Based on the preliminary identification results, the step of performing skill domain matching processing on the current information to be processed according to the preset identification strategy to obtain the matching results is controlled.
[0197] Optionally, in some embodiments of this application, the step of inputting the current information to be processed into the intelligent agent, and then performing skill domain matching processing on the current information to be processed according to the first recognition strategy to obtain a matching result, includes:
[0198] The current information to be processed is input into the intelligent agent, which then clarifies the current information to be processed based on the target user's historical feedback information and user characteristic information to obtain the processed information.
[0199] The processed information is subjected to skill domain matching according to a preset recognition strategy to obtain the matching result;
[0200] The user characteristic information includes at least one of user preference information or user commonly used phrases information;
[0201] The clarification process includes at least one of completion processing or error correction processing.
[0202] In this embodiment, the acquisition module 501 first acquires the current pending information for the target user sent by the terminal device. The identification module 502 performs skill domain matching processing on the current pending information according to a preset identification strategy to obtain a matching result. The preset identification strategy is optimized based on the target user's historical feedback information. The determination module 503 determines the target user's intent recognition result based on the matching result. The execution module 504 executes the control command corresponding to the intent recognition result to obtain an execution result.
[0203] In this embodiment, before analyzing the user's intent based on the current information to be processed, skill domain matching processing is performed on the current information to be processed, and the intent recognition result of the target user is determined based on the matching result, thereby improving the accuracy of user intent analysis and thus improving the accuracy of control processing based on intent recognition.
[0204] Among these measures, by optimizing the skill domain matching process based on users' historical feedback information, the matching degree between intent recognition results and user expectations is improved, thereby further enhancing the accuracy of user intent analysis and improving the user experience.
[0205] Please see Figure 6 , Figure 6 This is another structural schematic diagram of the information processing device provided in the embodiments of this application, wherein the information processing device is applied to a terminal device, specifically as follows:
[0206] The sending module 601 is used to send the current pending information of the target user to the server if the current pending information of the target user is collected.
[0207] The receiving module 602 is configured to receive the execution feedback result returned by the server for the current pending information, wherein the execution feedback result is generated based on the execution result of the server on the current pending information;
[0208] Output module 603 is used to output the execution feedback result;
[0209] The server performs skill domain matching on the current information to be processed according to a preset identification strategy to obtain a matching result, determines the intent recognition result of the target user based on the matching result, and executes the control command corresponding to the intent recognition result to obtain the execution result. The preset identification strategy is obtained by the server based on the historical feedback information of the target user.
[0210] Optionally, in some embodiments of this application, the current information to be processed includes voice information to be processed, and the method further includes:
[0211] When the voice information to be processed is detected, if the terminal device is currently outputting second audio information, the output volume of the second audio information is reduced, and the output volume of the second audio information is restored when the voice information to be processed is received.
[0212] In this embodiment, if the sending module 601 collects the current pending information of the target user, it sends the current pending information to the server. The receiving module 602 receives the execution feedback result returned by the server for the current pending information. The execution feedback result is generated based on the execution result of the server on the current pending information. The output module 603 outputs the execution feedback result.
[0213] In addition, this application also provides an electronic device, such as Figure 7 As shown, it illustrates the structural diagram of the electronic device involved in this application, specifically:
[0214] The electronic device may include components such as a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, a power supply 703, and an input unit 704. Those skilled in the art will understand that... Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0215] The processor 701 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 702, and by calling data stored in the memory 702, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. Optionally, the processor 701 may include one or more processing cores; preferably, the processor 701 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 701.
[0216] The memory 702 can be used to store software programs and modules. The processor 701 executes various functional applications and data processing by running the software programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 702 may also include a memory controller to provide the processor 701 with access to the memory 702.
[0217] The electronic device also includes a power supply 703 that supplies power to the various components. Preferably, the power supply 703 can be logically connected to the processor 701 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 703 may also include one or more DC or AC power supplies, recharging systems, power equipment debugging circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0218] The electronic device may also include an input unit 704, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0219] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 701 in the electronic device loads the executable files corresponding to the processes of one or more application programs into the memory 702 according to the following instructions, and the processor 701 runs the application programs stored in the memory 702, thereby implementing the steps in any of the information processing methods provided in the embodiments of this application.
[0220] In this embodiment of the application, the server obtains the current pending information for the target user sent by the terminal device, performs skill domain matching processing on the current pending information according to a preset identification strategy, and obtains a matching result. The preset identification strategy is optimized based on the historical feedback information of the target user. The server determines the intention recognition result of the target user according to the matching result, executes the control command corresponding to the intention recognition result, and obtains an execution result.
[0221] Specifically, by performing skill domain matching on the current information to be processed before analyzing the user's intent, and determining the target user's intent recognition result based on the matching result, the accuracy of user intent analysis is improved, thereby improving the accuracy of control processing based on intent recognition.
[0222] Among these measures, by optimizing the skill domain matching process based on users' historical feedback information, the matching degree between intent recognition results and user expectations is improved, thereby further enhancing the accuracy of user intent analysis and improving the user experience.
[0223] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0224] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0225] Therefore, this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps of any of the information processing methods provided in this application.
[0226] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0227] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0228] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the information processing methods provided in this application, the beneficial effects that any of the information processing methods provided in this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0229] The above provides a detailed description of an information processing method, apparatus, electronic device, and computer-readable storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. An information processing method characterized by comprising: Applied to a server, the method includes: Obtain the current pending information sent by the terminal device for the target user; The current information to be processed is matched with the skill domain according to the preset identification strategy to obtain the matching result. The preset identification strategy is optimized based on the historical feedback information of the target user. The intent recognition result of the target user is determined based on the matching result; Execute the control command corresponding to the intent recognition result to obtain the execution result; After executing the control command corresponding to the intent recognition result and obtaining the execution result, the method further includes: Generate voice prompts for the execution result, which are output by the terminal device to provide feedback on the execution result; If the terminal device is currently outputting first audio information, then output control information for the voice prompt information and the first audio information is generated according to a preset output strategy; The output control information is sent to the terminal device, wherein the terminal device controls the output of the voice prompt information and the first audio information according to the output control information; The preset output strategy includes a fixed output strategy and a dynamic output strategy for different application scenarios. The dynamic output strategy is obtained by optimizing the initial output strategy based on the historical feedback information.
2. The information processing method according to claim 1, characterized by, Determining the intent recognition result of the target user based on the matching result includes: If the matching result includes no match with any of the skill domains, then no user intent will be taken as the intent recognition result. If the matching result includes a partial match with any skill domain, then the following information to be processed corresponding to the current information to be processed is obtained, and the current information to be processed and the following information to be processed are fused to obtain the target information to be processed. Furthermore, the target information to be processed is subjected to intent recognition to obtain the user intent, and the user intent is used as the intent recognition result.
3. The information processing method according to claim 1, characterized by, If the terminal device is currently outputting first audio information, then output control information for the voice prompt information and the first audio information is generated according to a preset output strategy, including: If the terminal device is currently outputting first audio information, then the first output information of the first audio information is determined according to the preset output strategy, and the second output information of the first audio information is determined based on the first output information; The first output information and the second output information are used as the output control information; The first output information includes at least one of uninterruptible, interruptible and recoverable, or interruptible and unrecoverable, and the second output information includes at least one of direct output or subsequent output.
4. The information processing method according to claim 1, characterized in that, The preset identification strategy includes a first identification strategy; The step of performing skill domain matching processing on the current information to be processed according to a preset recognition strategy to obtain a matching result includes: The current information to be processed is input into the intelligent agent, and the intelligent agent performs skill domain matching processing on the current information to be processed according to the first recognition strategy to obtain the matching result. Before performing skill domain matching processing on the current information to be processed according to a preset recognition strategy to obtain a matching result, the method further includes: The current information to be processed is subjected to skill domain matching processing using a second identification strategy to obtain preliminary identification results; Based on the preliminary identification results, the step of performing skill domain matching processing on the current information to be processed according to the preset identification strategy to obtain the matching results is controlled.
5. The information processing method according to claim 4, characterized in that, The step of inputting the current information to be processed into the intelligent agent, and then performing skill domain matching processing on the current information to be processed by the intelligent agent according to the first recognition strategy to obtain a matching result includes: The current information to be processed is input into the intelligent agent, which then clarifies the current information to be processed based on the target user's historical feedback information and user characteristic information to obtain the processed information. The processed information is subjected to skill domain matching according to a preset recognition strategy to obtain the matching result; The user characteristic information includes at least one of user preference information or user commonly used phrases information; The clarification process includes at least one of completion processing or error correction processing.
6. An information processing method, characterized in that, Applied to a terminal device, the method includes: If the target user's current pending information is collected, the current pending information is sent to the server; Receive the execution feedback result returned by the server for the current pending information, the execution feedback result being generated based on the server's execution result for the current pending information; Output the execution feedback result; The server performs skill domain matching on the current information to be processed according to a preset identification strategy to obtain a matching result, determines the intent identification result of the target user based on the matching result, and executes the control command corresponding to the intent identification result to obtain the execution result. The preset identification strategy is obtained by the server based on the historical feedback information of the target user. The execution feedback result includes voice prompt information, which is output by the terminal device to provide feedback on the execution result; After generating the execution feedback result, the server also performs the following operations: If the terminal device is currently outputting first audio information, then output control information for the voice prompt information and the first audio information is generated according to a preset output strategy; The output control information is sent to the terminal device, wherein the terminal device controls the output of the voice prompt information and the first audio information according to the output control information; The preset output strategy includes a fixed output strategy and a dynamic output strategy for different application scenarios. The dynamic output strategy is obtained by optimizing the initial output strategy based on the historical feedback information. The output of the execution feedback result includes: The terminal device receives the output control information sent by the server; The terminal device controls the output of the voice prompt information and the first audio information according to the output control information.
7. The information processing method according to claim 6, characterized in that, The current information to be processed includes voice information to be processed, and the method further includes: When the voice information to be processed is detected, if the terminal device is currently outputting second audio information, the output volume of the second audio information is reduced, and the output volume of the second audio information is restored when the voice information to be processed is received.
8. An information processing device, characterized in that, Applied to a server, the device includes: The acquisition module is used to acquire the current pending information for the target user sent by the terminal device; The identification module is used to perform skill domain matching processing on the current information to be processed according to a preset identification strategy to obtain a matching result. The preset identification strategy is optimized based on the historical feedback information of the target user. The determining module is used to determine the intent recognition result of the target user based on the matching result; The execution module is used to execute the control instructions corresponding to the intent recognition result and obtain the execution result; After executing the control command corresponding to the intent recognition result and obtaining the execution result, the device further includes: Generate voice prompts for the execution result, which are output by the terminal device to provide feedback on the execution result; If the terminal device is currently outputting first audio information, then output control information for the voice prompt information and the first audio information is generated according to a preset output strategy; The output control information is sent to the terminal device, wherein the terminal device controls the output of the voice prompt information and the first audio information according to the output control information; The preset output strategy includes a fixed output strategy and a dynamic output strategy for different application scenarios. The dynamic output strategy is obtained by optimizing the initial output strategy based on the historical feedback information.
9. An information processing device, characterized in that, Applied to a terminal device, the device includes: The sending module is used to send the current pending information of the target user to the server if it collects the current pending information. The receiving module is configured to receive the execution feedback result returned by the server for the current pending information, wherein the execution feedback result is generated based on the execution result of the server on the current pending information; The output module is used to output the execution feedback results; The server performs skill domain matching on the current information to be processed according to a preset identification strategy to obtain a matching result, determines the intent identification result of the target user based on the matching result, and executes the control command corresponding to the intent identification result to obtain the execution result. The preset identification strategy is obtained by the server based on the historical feedback information of the target user. The execution feedback result includes voice prompt information, which is output by the terminal device to provide feedback on the execution result; After generating the execution feedback result, the server also performs the following operations: If the terminal device is currently outputting first audio information, then output control information for the voice prompt information and the first audio information is generated according to a preset output strategy; The output control information is sent to the terminal device, wherein the terminal device controls the output of the voice prompt information and the first audio information according to the output control information; The preset output strategy includes a fixed output strategy and a dynamic output strategy for different application scenarios. The dynamic output strategy is obtained by optimizing the initial output strategy based on the historical feedback information. The output of the execution feedback result includes: The terminal device receives the output control information sent by the server; The terminal device controls the output of the voice prompt information and the first audio information according to the output control information.
10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the information processing method as described in any one of claims 1-5, or to implement the steps of the information processing method as described in any one of claims 6-7.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the information processing method as described in any one of claims 1-7.
Citation Information
Patent Citations
Sound output control method and sound output control device
CN113516978A
Video playing method and device, electronic equipment, storage medium and program product
CN114125541A