Voice processing method, device, equipment, medium and program product
By displaying interactive dialogue flow information and related content information in partitions in the interactive dialogue interface, the problem of insufficient interactive dialogue information in the prior art is solved, and the dialogue quality and efficiency are improved.
Patent Information
- Application Number
- CN202510553886.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-12
AI Technical Summary
In existing interactive conversations, the form of only answers output is simpler and single. Each round of interactive conversations provides users with limited reference information, resulting in reduced conversation quality and efficiency.
Partition display is displayed in the interactive dialogue interface. The interactive dialogue area is used to present interactive dialogue flow information, and the content display area is used to display content information related to the dialogue flow information, including problem description and result information, as well as resources such as videos and images.
Through rich interactive dialogue content display, users can make decisions faster and more accurate, significantly improving the quality and efficiency of conversations.
Smart Images

Figure CN120475003A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, in particular to the field of artificial intelligence, and specifically to a speech processing method, a speech processing device, a computer device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Interactive dialogue refers to the process of completing voice tasks through two-way and dynamic information transmission between users and machines.
[0003] Currently, during interactive conversations between users and systems, the system only displays answers to user queries. However, this simple and single-minded interactive dialogue format, which only outputs answers, provides users with limited information for reference during each round of interactive conversation. This not only makes it difficult to make informed decisions based on this limited information, but also reduces the quality and efficiency of the interactive conversation. Summary of the Invention
[0004] The embodiments of the present application provide a speech processing method, apparatus, device, medium, and program product, which can enrich the presentation content during an interactive dialogue and improve the dialogue quality and efficiency of the interactive dialogue.
[0005] In one aspect, an embodiment of the present application provides a speech processing method, the method comprising:
[0006] Receive voice signals and display an interactive dialogue interface; the interactive dialogue interface includes an interactive dialogue area and a content display area;
[0007] Displaying interactive dialogue flow information triggered by voice signals in the interactive dialogue area; and
[0008] Content information related to the interactive dialog flow information is displayed in the content display area.
[0009] On the other hand, an embodiment of the present application provides a speech processing device, the device comprising:
[0010] A receiving unit, configured to receive voice signals and display an interactive dialogue interface; the interactive dialogue interface includes an interactive dialogue area and a content display area;
[0011] a processing unit, configured to display interactive dialogue flow information triggered by a voice signal in an interactive dialogue area; and
[0012] The processing unit is further configured to display content information related to the interactive dialogue flow information in the content display area.
[0013] In one implementation, the interactive dialogue flow information includes N rounds of interactive dialogue information, each round of interactive dialogue information includes question description information and question result information; N is a positive integer;
[0014] The content information related to the interactive dialog flow information includes at least one of the following:
[0015] One or more candidate question description information that matches the dialogue intent of the i-th round of interactive dialogue information; the i-th round of interactive dialogue information is displayed in the interactive dialogue area, where i is an integer and 1≤i≤N; and
[0016] Target content information related to the target dialogue information in the i-th round of interactive dialogue information; wherein the target content information includes at least one of the following: detailed information of the target dialogue information, source information of the target dialogue information, and commentary information of the target dialogue information; the target dialogue information is default or customized.
[0017] In one implementation, the interactive dialogue flow information includes N rounds of interactive dialogue information; the i-th round of interactive dialogue information includes first question description information and first question result information, and the first question result information includes a resource image corresponding to the video media resource; the resource image is defaulted to the target dialogue information in the i-th round of interactive dialogue information;
[0018] The processing unit is configured to display content information related to the interactive dialogue flow information in the content display area, specifically to:
[0019] In the content display area, target content information related to the resource image is displayed by default.
[0020] In one implementation, the target content information related to the resource image is a video media resource; and the processing unit is further configured to:
[0021] Play video media resources in the content display area;
[0022] During playback of the video media resource, receiving a pause operation for the video media resource;
[0023] In response to a pause operation on the video media resource, the video media resource is paused in the content presentation area.
[0024] In one implementation, the interactive dialogue flow information includes N rounds of interactive dialogue information; and the processing unit is configured to display content information related to the interactive dialogue flow information in the content display area, specifically configured to:
[0025] In response to an information viewing operation performed in the interactive dialogue area on target dialogue information in the i-th round of interactive dialogue information, displaying the target dialogue information in the interactive dialogue area as a selected state;
[0026] Target content information related to the target conversation information is displayed in the content display area.
[0027] In one implementation, the number of content information related to the interactive dialogue flow information is at least one; the content display area includes at least one content type option, each content type option corresponding to a type of content information; and the processing unit is further configured to:
[0028] In response to a selection operation on a target content type option, content information corresponding to the selected target content type option is displayed in the content display area; the target content type option is any one of the at least one content type option.
[0029] In one implementation, the processing unit is further configured to:
[0030] During the process of performing a dialogue viewing operation on the interactive dialogue information of the i-th round in the interactive dialogue area, maintaining the display of content information related to the interactive dialogue information of the i-th round in the content display area; the dialogue viewing operation includes at least one of the following: a page turning operation, a scrolling operation, and a dragging operation;
[0031] When a dialogue viewing operation is detected in the interactive dialogue area to switch from the i-th round of interactive dialogue information to the j-th round of interactive dialogue information, the content information related to the i-th round of interactive dialogue information will be updated and displayed in the content display area as the content information related to the j-th round of interactive dialogue information; j is a positive integer, j≠i, and 1≤j≤N.
[0032] In one implementation, the content display area includes dialogue options corresponding to each round of interactive dialogue information in the interactive dialogue flow information; the processing unit is further configured to:
[0033] In response to a triggering operation on a target dialogue option in the content display area, content information related to the target interactive dialogue information corresponding to the target dialogue option is displayed in the content display area; the target dialogue option is a dialogue option corresponding to any round of interactive dialogue information;
[0034] The target interactive dialogue information is displayed in the interactive dialogue area.
[0035] In one implementation, the number of target dialogue information in the i-th round of interactive dialogue information is at least two; and the processing unit is further configured to:
[0036] When the content positioning operation is performed on the target dialogue information designated in the interactive dialogue area, target content information related to the designated target dialogue information is displayed in the content display area.
[0037] In one implementation, the voice processing method is applied to a media resource application program that runs on a resource playback device; and a receiving unit, when receiving a voice signal, is specifically configured to:
[0038] Displaying the service interface of the resource playback device, the service interface includes voice prompt information in the prompt state; the voice prompt information in the prompt state is used to prompt that the resource playback device has a voice interaction function;
[0039] In response to a confirmation operation on the voice prompt information in the prompt state, the voice prompt information is converted from the prompt state to the wake-up state; the voice prompt information in the wake-up state is used to prompt the resource playback device to be in the voice collection state;
[0040] When the resource playback device starts to collect voice signals, the voice prompt information is converted from the awakening state to the collection state; the voice prompt information in the collection state is used to prompt the resource playback device that the voice signal is being collected;
[0041] When the resource playback device finishes collecting the voice signal, the voice prompt information is converted from the collection state to the understanding state; the voice prompt information in the understanding state is used to prompt the resource playback device to perform result search processing based on the voice signal.
[0042] In one implementation, the resource playback device is a device that transmits content based on the Internet; wherein,
[0043] The interactive dialogue area and the content display area may be displayed in the interactive dialogue interface in a manner including: horizontal arrangement display, vertical arrangement display, and mosaic display.
[0044] In one implementation, the result search process includes an intent recognition process and a result generation process, and the processing unit is further configured to:
[0045] Obtaining basic analysis data, where the basic analysis data includes one or more of scenario data, object data, permission data, and intent list data;
[0046] Call the fine-tuned generative model to perform intent recognition processing on the voice signal based on the basic analysis data to obtain the intent recognition result;
[0047] If the intention recognition result indicates that the speech task indicated by the speech signal is a dialogue task, result generation processing is performed based on the intention recognition result to generate question result information corresponding to the speech signal; the question result information and the question description information corresponding to the speech signal constitute a round of interactive dialogue information in the interactive dialogue flow information.
[0048] In one implementation, the processing unit is further configured to:
[0049] If the intention recognition result indicates that the speech task indicated by the speech signal is a direct task, the direct task is processed;
[0050] Among them, direct tasks include at least one of the following: control tasks, interface switching tasks and player control tasks.
[0051] In one implementation, the processing unit is configured to perform result generation processing based on the intention recognition result, and when generating question result information corresponding to the speech signal, specifically:
[0052] Obtaining at least one round of interactive dialogue flow information in the interactive dialogue flow information;
[0053] Performing intranet resource search processing based on the conversation intention of at least one round of interactive conversation flow information to obtain intranet resources; and
[0054] Performing external network resource search processing based on the conversation intention of at least one round of interactive conversation flow information to obtain external network resources;
[0055] The intranet resources and the extranet resources are reorganized to generate problem result information corresponding to the voice signal.
[0056] In another aspect, an embodiment of the present application provides a computer device, comprising:
[0057] a processor adapted to execute a computer program;
[0058] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned speech processing method is implemented.
[0059] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is suitable for being loaded by a processor and executing the above-mentioned speech processing method.
[0060] On the other hand, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned speech processing method.
[0061] In an embodiment of the present application, a computer device can receive a user's voice signal collected from a physical environment and display an interactive dialogue interface for presenting interactive dialogue information between the user and the computer device. In an embodiment of the present application, the interactive dialogue interface can display content in partitions, specifically, the interactive dialogue interface includes an interactive dialogue area and a content display area. Among them, the interactive dialogue flow information triggered by the voice signal is displayed in the interactive dialogue area. The interactive dialogue flow information includes each round of interactive dialogue information, which can help users more conveniently determine the results of the questions in the form of question description information-question result information, thereby improving the user's question-answering experience. At the same time, content information related to the interactive dialogue flow is displayed in the content display area. The content information is a further information supplement to the interactive dialogue flow information, thereby improving the dialogue quality of human-computer interaction; and the user can combine the interactive dialogue flow information in the interactive dialogue area and the content information displayed in the content display area to jointly make decisions, etc., which can reduce the number of dialogues (or dialogue rounds) to a certain extent, shorten the user's decision path, and effectively improve the dialogue efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0063] Figure 1a is a schematic diagram of an existing interactive dialogue interface;
[0064] Figure 1b is a schematic diagram of an interactive dialogue interface provided by an exemplary embodiment of the present application;
[0065] Figure 2a This is a schematic diagram of the architecture of an interactive dialogue system provided by an exemplary embodiment of the present application;
[0066] Figure 2b is a schematic diagram of the architecture of another interactive dialogue system provided by an exemplary embodiment of the present application;
[0067] Figure 3 This is a flowchart of a speech processing method provided by an exemplary embodiment of the present application;
[0068] Figure 4a is a schematic diagram of displaying an interactive dialogue interface on a display screen provided by an exemplary embodiment of the present application;
[0069] Figure 4b is another schematic diagram of displaying an interactive dialogue interface on a display screen provided by an exemplary embodiment of the present application;
[0070] Figure 4c This is a schematic diagram of a display method of an interactive dialogue area and a content display area provided by an exemplary embodiment of the present application;
[0071] Figure 4d This is a schematic diagram of another display method of the interactive dialogue area and the content display area provided by an exemplary embodiment of the present application;
[0072] Figure 5a This is a schematic diagram of an interactive dialogue message of a graphic type provided by an exemplary embodiment of the present application;
[0073] Figure 5b is a schematic diagram of a plain text type interactive dialogue message provided by an exemplary embodiment of the present application;
[0074] Figure 6 This is a schematic diagram of displaying candidate question description information in a content display area provided by an exemplary embodiment of the present application;
[0075] Figure 7 This is a schematic diagram of displaying content information related to target conversation information in a content display area provided by an exemplary embodiment of the present application;
[0076] Figure 8 This is a schematic diagram of previewing a video media resource in a content display area provided by an exemplary embodiment of the present application;
[0077] Figure 9a is a schematic diagram of an information viewing operation provided by an exemplary embodiment of the present application;
[0078] Figure 9b is a schematic diagram of another information viewing operation provided by an exemplary embodiment of the present application;
[0079] Figure 10 is a schematic diagram of a content type option provided by an exemplary embodiment of the present application;
[0080] Figure 11 This is a schematic diagram of a linkage display of an interactive dialogue area and a content display area provided by an exemplary embodiment of the present application;
[0081] Figure 12a is a schematic diagram of a dialogue option provided by an exemplary embodiment of the present application;
[0082] Figure 12bThis is a schematic diagram of an exemplary embodiment of the present application providing a dialog option and a content type option displayed simultaneously in a content display area;
[0083] Figure 13 is a flowchart of another speech processing method provided by an exemplary embodiment of the present application;
[0084] Figure 14 is a schematic diagram of voice prompt information in different states provided by an exemplary embodiment of the present application;
[0085] Figure 15a This is a schematic diagram of a direct-access task provided by an exemplary embodiment of the present application;
[0086] Figure 15b is a schematic diagram of another direct-access task provided by an exemplary embodiment of the present application;
[0087] Figure 15c This is a schematic diagram of another direct-access task provided by an exemplary embodiment of the present application;
[0088] Figure 16 This is a schematic diagram of the background logic flow of a voice processing method provided by an exemplary embodiment of the present application;
[0089] Figure 17 is a schematic diagram of training data construction provided by an exemplary embodiment of the present application;
[0090] Figure 18 This is a flowchart of a model training process provided by an exemplary embodiment of the present application;
[0091] Figure 19 This is a logical diagram of a film and television model service provided by an exemplary embodiment of the present application;
[0092] Figure 20 This is a schematic diagram of pausing a response during a streaming response to question result information in an interactive dialogue area, provided by an exemplary embodiment of the present application;
[0093] Figure 21 This is a schematic diagram of a streaming response content information in a content display area provided by an exemplary embodiment of the present application;
[0094] Figure 22 is a structural diagram of a speech processing device provided by an exemplary embodiment of the present application;
[0095] Figure 23 It is a structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0096] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0097] In the embodiments of the present application, a speech processing solution is proposed, specifically an AI speech processing solution for implementing interactive dialogue based on a large language model. To facilitate understanding, the following is a brief introduction to the technical terms and related concepts involved in the speech processing solution, including:
[0098] 1. Interactive dialogue.
[0099] Interactive dialogue, which can be understood as intelligent question answering (QA) or intelligent dialogue, belongs to the field of human-computer interaction and is an advanced form of information retrieval system. With the rapid development and application of artificial intelligence (AI), interactive dialogue technology has become a highly sought-after and promising area in the field of natural language processing (NLP). Interactive dialogue can be understood as a human-computer interaction process in which a user and a machine communicate through natural language, the machine dynamically understands the user's query, and completes a speech task (or speech goal) based on the context. Among them:
[0100] (1) Machine refers to applications, code programs and agents that have multiple functions such as semantic understanding, intention recognition and result generation. Among them, ① application refers to a computer program for completing one or more specific tasks; according to the classification of the operation mode of the application, the application may include but is not limited to: a client that needs to download an installation package and deploy the installation package in the terminal, and realizes intelligent question and answer by running the installation package; a small program that does not need to download the installation package but runs as a subroutine on the client; and a web (World Wide Web, Global Wide Area Network) application that is opened and run through the browser in the terminal; etc. ② Code program can be understood as a code fragment that can realize the above-mentioned multiple functions when loaded. ③ Agent refers to a physical robot or virtual robot that can transmit information, make decisions and perform actions to achieve goals with users in a physical environment (i.e., real-world environment). For example, an agent is an intelligent assistant deployed on a hardware device; an intelligent assistant can be called an intelligent assistant application, which refers to an intelligent software deployed on a hardware device that can interact with users through voice, text and images, and realize rich functions such as information query, life server and learning guidance.
[0101] Among them, there may be overlaps between the above-mentioned applications, code programs, and intelligent agents. For example, an intelligent agent - an intelligent assistant - is integrated into the application; in this way, the intelligent assistant can be evoked or called in the application to help the user realize an interactive dialogue in the application. Therefore, the embodiment of the present application does not limit the specific type of machine that actually interacts with the user in the interactive dialogue scenario; for the sake of convenience, the interactive dialogue between the user and the application, specifically the interactive dialogue with the intelligent assistant integrated in the application, is used as an example to realize human-computer interaction, which is specially explained here.
[0102] (2) The interactive dialogue process may include N rounds of interactive dialogue, where N is a positive integer. Each round of interactive dialogue in the N rounds of interactive dialogue corresponds to interactive dialogue information, and each round of interactive dialogue information includes: the question description information input by the user and the question result information generated by the machine based on the question description information. Among them: ① The question description information refers to the request, question or instruction input by the user to the machine; the question description information can be called the user query (or simply query), which is usually expressed in text, voice or other interactive forms (such as gestures, images); the query is the starting point for the machine to trigger a response or service, and is also the core processing object of the interactive dialogue. ② The question result information is the result information generated and returned to the user after the machine performs problem processing on the problem described in the problem description information; among them, the problem processing performed by the machine on the problem described in the problem description information may include but is not limited to: intention understanding → knowledge search → result generation and other processing operations. The machine will use accurate and concise natural language to output the question description information to achieve a response to the user query.
[0103] To optimize the user experience of intelligent question-and-answer (Q&A), a streaming response approach is often used to present interactive conversation information. This streaming response can be embodied in two key aspects: First, it allows for streaming presentation of N rounds of interactive conversations. Specifically, after a round of interactive conversation information has concluded, the user can continue to submit new queries. The machine then combines this new query with information from previous rounds of interactive conversations to generate the corresponding question result information. In other words, the N rounds of interactive conversation information include the question description and question result information for each round of continuous Q&A. Second, when processing a user query, rather than returning the question result information all at once after generating it, the machine returns the generated portion of the question result information in real time as it is generated. In other words, the machine feeds back the generated sub-information of the question result information to the user while continuing to generate other sub-information. This demonstrates that streaming responses enable multi-round interactive conversations, improve context understanding, reduce latency in question result output, and enhance the user experience. They are well-suited for interactive conversations with real-time updates or large amounts of data.
[0104] 2. Artificial intelligence.
[0105] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. Machine learning models within the AI field are network models derived through model training using machine learning. Machine learning is a multidisciplinary discipline that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. Machine learning typically encompasses techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, model-based learning, and deep learning (DL).
[0106] The embodiments of the present application mainly relate to large language models (LLMs) in the field of deep learning in machine learning technology. Large language models are general language processing models based on deep learning technology and massive parameter scales (usually billions to trillions); large language models have gradually become the core technology in the field of natural language processing (NLP), leading the technological innovation from traditional tasks to generative artificial intelligence. Large language models have achieved significant performance improvements in tasks such as text generation, question-answering systems, and machine translation through pre-training of large-scale parameters, demonstrating powerful generalization and context understanding capabilities. With the continuous expansion of the model scale of large language models, from the initial millions of parameters to hundreds of billions or even trillions of parameters today, the performance of large language models has shown an approximately logarithmic linear growth trend with the model scale, which has further promoted the widespread adoption of large language models in scientific research and industrial applications.
[0107] The large language model is essentially a generative model with the following characteristics: ① It has strong generalization and multi-round result generation capabilities, enabling it to recognize long queries (such as long voice signals or long text strings). Therefore, it can better handle complex user queries and unrecognized queries. ② It has strong comprehension capabilities, enabling it to better understand user intent and, therefore, can be fine-tuned to identify user intent in conjunction with business operations. ③ It has the ability to generate content. By fine-tuning (or fine-tuning) the base large language model, the fine-tuned language model can be combined with business operations to recall higher-quality content.
[0108] In practical applications, interactive dialogues are presented solely through the interactive dialogue flow; that is, the user's input question description is used to generate and display the result information. This interactive dialogue presentation approach suffers from a single dialogue format. Answering questions solely through the result information often provides limited reference information to the user, forcing the user to engage in multiple rounds of interactive dialogue in an attempt to obtain a more accurate answer.
[0109] To enhance the richness of content presented during interactive conversations, more reference information can be provided to users during each round of interactive conversations. The speech processing solution provided in the embodiments of the present application improves the presentation of interactive conversations. Specifically, it not only provides users with interactive conversation flow information, so that the question description information entered by the user can be directly responded to through the question result information in the interactive conversation flow information, but also provides users with content information related to the interactive conversation flow information. This content information serves as a further information supplement to the interactive conversation flow information, providing users with more reference information. In this way, users can combine the interactive conversation flow information and the content information related to the interactive conversation flow information to jointly make decisions, significantly improving the quality of conversations and, to a certain extent, reducing the number of conversations (or conversation rounds), thereby effectively improving conversation efficiency.
[0110] Specifically, the interactive dialogue presentation process of the voice processing solution provided in the embodiment of the present application can be roughly described as follows: receiving a voice signal input by a user, which is a sound wave signal that carries language information and is generated by the user through a vocal organ (such as vocal cords or oral cavity); the user uses the voice signal to express or convey his or her dialogue intention. Displaying an interactive dialogue interface, which includes an interactive dialogue area and a content display area; wherein, the interactive dialogue flow information triggered by the voice signal input by the user is displayed in the interactive dialogue area, and content information related to the interactive dialogue flow information is displayed in the content display area.
[0111] In practice, it has been found that the voice processing solution provided by the embodiment of the present application has obvious advantages in presenting content during interactive dialogue. The following is an example of comparing the content presentation of the present application solution with the existing voice processing solution to illustrate the advantages of the embodiment of the present application, including:
[0112] A schematic diagram of content presentation in existing speech processing solutions can be found in Figure 1a ; Only the interactive dialogue flow information is displayed in the interactive dialogue interface, and the user finds the information he wants and makes a decision by looking for the problem result information in the interactive dialogue flow information. In the case that the user is not satisfied with the current problem result information, the user needs to input different problem description information multiple times, hoping that the machine can provide more in line with the intention of the problem result information. It is not difficult to find that this single content presentation method can provide limited reference information to the user, which causes the user to have multiple human-computer interactions, reducing the quality and efficiency of the dialogue. However, the schematic diagram of the content presentation of the voice processing solution provided in the embodiment of the present application can be seen in Figure 1b; Content is displayed in a partitioned manner in the interactive dialogue interface, specifically, the interactive dialogue area is divided into an interactive dialogue area and a content display area. Among them, the interactive dialogue area is used to present interactive dialogue flow information, helping users experience streaming questions and answers, which is more in line with users' inquiry habits and improves human-computer interactivity. In addition to the interactive dialogue area, the interactive dialogue interface is also provided with a content display area; the content display area is used to present content information related to the interactive dialogue flow information in the interactive dialogue area as an information supplement to the interactive dialogue flow information, so as to provide users with more reference information. In this way, users can combine richer interactive dialogue flow information and content information related to the interactive dialogue flow information to make decisions more quickly and accurately, significantly improving the dialogue quality and dialogue efficiency of the interactive dialogue.
[0113] The voice processing solution provided in the embodiments of this application serves as an automated question-and-answer solution for human-computer interaction. It can understand, parse, and answer user queries, making the voice processing solution provided in the embodiments of this application applicable to a variety of interactive dialogue scenarios. These interactive dialogue scenarios may include, but are not limited to: ① Customer Support: The intelligent question-and-answer system can serve as a customer support tool to answer common user questions, reduce the workload of customer service staff, and improve customer satisfaction. ② Internal Enterprise Knowledge Base: Enterprises can use the voice processing solution provided in the embodiments of this application to search for a richer knowledge base, build an internal knowledge base, and help employees quickly find the information they need, thereby improving work efficiency. ③ Virtual Assistant: Machines equipped with the voice processing solution provided in the embodiments of this application can serve as virtual assistants for individuals or enterprises, providing functions such as daily task management, scheduling, and reminder services. ④ Online Education: Machines equipped with the voice processing solution provided in the embodiments of this application can be used in the field of online education to provide students with personalized learning resources and real-time question-and-answer services. ⑤ E-commerce: Machines equipped with the voice processing solution provided in the embodiments of this application can help users answer questions and provide shopping recommendations during the shopping process, thereby improving the shopping experience. ⑥ Financial services: A machine equipped with the voice processing solution provided in the embodiments of the present application can provide real-time consulting services to customers of financial institutions such as banks and insurance companies, and answer questions about accounts, transactions, products, etc. ⑦ Medical consultation: A machine equipped with the voice processing solution provided in the embodiments of the present application can provide basic medical consulting services to patients, and answer questions about diseases, treatments, drugs, etc. ⑧ Tourism consultation: A machine equipped with the voice processing solution provided in the embodiments of the present application can provide tourists with real-time tourism information, and answer questions about attractions, hotels, transportation, etc. ⑨ News and information retrieval: A machine equipped with the voice processing solution provided in the embodiments of the present application can help users quickly find the news and information they need, thereby improving the efficiency of information retrieval. ⑩ Media resource scenario: In the media resource scenario, users often have the need to quickly search for multimedia resources, which may include but are not limited to: audio (such as music), video (such as movies or short videos, etc.) and images; in this case, the introduction of the voice processing solution provided in the embodiments of the present application can provide users with richer resource search results during the interactive dialogue process, significantly improving the search efficiency of media resources.
[0114] It should be understood that the above description is merely an exemplary product performance and interactive dialogue scenario provided by the embodiments of this application, and does not limit the product performance and interactive dialogue scenario of the voice processing solution provided by the embodiments of this application. The voice processing solution provided by the embodiments of this application can provide efficient, accurate, and convenient question-and-answer services in various interactive dialogue scenarios, demonstrating high value and practicality in various interactive dialogue scenarios, and helping to improve user experience and satisfaction.
[0115] For ease of explanation, the interactive dialogue scenario involved in the voice processing solution provided in the embodiment of the present application is used as an example of a media resource scenario for introduction. In the media resource scenario, the voice processing solution provided in the embodiment of the present application is applied to a media resource application, which can be understood as a client that integrates the voice processing solution provided in the embodiment of the present application and has the ability to search and play multimedia resources; for example, the voice processing solution provided in the embodiment of the present application is specifically integrated into the media resource application in the form of an intelligent assistant, so that after the intelligent assistant in the media resource application is awakened, human-computer interaction is realized to achieve rapid search of media resources.
[0116] The following combination Figure 2a A schematic diagram of the system architecture of an interactive dialogue system is given; Figure 2a As shown, the interactive dialogue system includes an object 201, a terminal 202, and a server 203. The embodiment of the present application does not limit the number and naming of the object 201, the terminal 202, and the server 203.
[0117] ① Object 201 is a user capable of sending voice signals. Object 201 and terminal 202 can be in the same physical environment or in different physical environments. When object 201 and terminal 202 are in the same physical environment, object 201 can achieve near-field interaction with terminal 202. When object 201 and terminal 202 are in different physical environments, object 201 can achieve remote interaction with terminal 202.
[0118] ② Terminal 202 refers to a terminal device with interactive dialogue capabilities. Terminal devices may include, but are not limited to, smartphones (such as those running Android or running the Internetworking Operating System (IOS)), portable personal computers (such as tablets and smart computers), mobile internet devices (MIDs), in-vehicle devices, head-mounted devices, intelligent chatbots, smart TVs, set-top boxes, projectors, and smart billboards. It should be understood that the terminals 202 involved in or used in different interactive dialogue scenarios may be different or the same.
[0119] For example, in the case where the aforementioned interactive dialogue scenario is a media resource scenario and a large screen (i.e., a display screen with a larger display area), the terminal 202 may refer to a resource playback device that runs a media resource application, is deployed with the voice processing solution provided by the embodiment of the present application, and has a larger display screen. The resource playback device may be referred to as an OTT (Over-The-Top) device or an OTT large screen, including but not limited to the above-mentioned smart TVs and set-top boxes that can be connected to the Internet, smart computers with larger display screens, and electronic billboards connected to the Internet. The resource playback device is a device that transmits content based on the Internet's content transmission method; the resource playback device is different from traditional telecommunications operators, and application service providers provide application services directly through the Internet. A player is deployed in the resource playback device, and the device can be used to play audio and video through the player; the resource playback device also has voice collection capabilities to realize voice signal collection, such as the resource playback device itself is integrated with a signal collector, or an external signal collector is used to realize the collection of the user's voice signal. Considering that the voice processing solution provided in the embodiment of the present application can realize content display in different areas, this content display method in different areas can better utilize the display screen with a larger display area of the OTT device, thereby improving the utilization rate of the display screen.
[0120] ③ Server 203 is the server corresponding to terminal 202, and is used to interact with terminal 202 for data to provide computing and application service support for terminal 202, specifically to provide application services and technical support for machines with interactive dialogue capabilities (such as media resource applications, code levels or intelligent assistants, etc.) running in terminal 202. Server 203 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. For example, server 203 is an independent physical server. In this case, Figure 2a The interactive dialogue shown is a centralized system. For another example, the server 203 is a server cluster composed of multiple physical servers. In this case, Figure 2a The interactive dialogue shown is a distributed system, wherein the terminal 202 and the server 203 can be connected directly or indirectly via wired or wireless communication, which is not limited in this application.
[0121] The speech processing solution provided in the embodiment of the present application can be Figure 2aThe terminal 202 or the server 203 in the system shown in the figure can execute the voice processing scheme, or the terminal 202 and the server 203 can execute the voice processing scheme together; that is, the computer device for executing the voice processing scheme can be the terminal 202 or the server 203, or can also include the terminal 202 and the server 203. Figure 2a The interactive dialogue system shown introduces the flow of the speech processing solution, where:
[0122] Terminal 202 receives a voice signal input by a user and transmits the voice signal to server 203. Server 203 performs search results processing on the voice signal and outputs the generated question result message to terminal 202 via a streaming response. After receiving the question result message, terminal 202 displays an interactive dialogue interface through a resource media application and outputs interactive dialogue flow information and content information related to the interactive dialogue flow information via a streaming response in the interactive dialogue interface. Specifically, the interactive dialogue flow information triggered by the voice signal is displayed in the interactive dialogue area of the interactive dialogue interface, and the content information related to the interactive dialogue flow information is displayed in the content display area of the interactive dialogue interface. In this way, object 201 can make decisions based on the interactive dialogue flow information in the interactive dialogue area and the content information related to the interactive dialogue flow information in the content display area, significantly improving the accuracy of the decision, thereby improving the dialogue quality and object efficiency.
[0123] Thus, the embodiments of the present application display interactive dialogue flow information in the interactive dialogue area of the interactive dialogue interface, helping users experience streaming questions and answers and improving human-computer interactivity. Furthermore, content information related to the interactive dialogue flow information is displayed in the content display area as a supplement to the interactive dialogue flow information, providing users with more reference information. Compared to the traditional display of only interactive dialogue flow information, users can combine the richer interactive dialogue flow information and the content information related to the interactive dialogue flow information to make decisions more quickly and accurately, significantly improving the quality and efficiency of interactive dialogues.
[0124] Based on the above brief introduction to the speech processing solution, interactive dialogue scenario, and interactive dialogue system provided in the embodiments of the present application, the following points should be explained:
[0125] ① The above-mentioned embodiments of this application Figure 2a The system architecture shown is for the purpose of more clearly illustrating the technical solutions of the embodiments of the present application and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. It is known to those skilled in the art that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems. In other words, Figure 2a The architectural diagram of the interactive dialogue system shown is only an exemplary architectural diagram; in actual applications, the number and distribution of computer devices included in the interactive dialogue system may change, and the embodiments of the present application do not limit the architectural diagram of the interactive dialogue system.
[0126] For example, Figure 2a The example of object 201 and terminal 202 being in the same physical environment is used to illustrate the process of terminal 202 collecting the voice signal of object 201 in the near field. However, in actual applications, object 201 and terminal 202 may also be in different physical environments; in this case, object 201 also holds a terminal device (such as a smart phone, etc.) so that terminal 202 can input a voice signal through the terminal device, and the voice signal will be transmitted by the terminal device to terminal 202 or server 203 for result search processing. Figure 2b As shown, the interactive dialogue system further includes a terminal 204 , through which the object 201 implements remote voice control of the terminal 202 .
[0127] ② The collection and processing of relevant data in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations. The acquisition of personal information requires the knowledge or consent of the individual subject (or the legal basis for obtaining the information), and subsequent data use and processing shall be carried out within the scope of authorization of laws and regulations and the subject of personal information. For example, when the embodiments of this application are applied to specific products or technologies, such as when collecting the user's voice signal, the user's permission or consent must be obtained, and the collection, use and processing of relevant data (such as search results for voice signals, etc.) must comply with relevant laws, regulations and standards of the relevant region.
[0128] Based on the voice processing solution described above, the embodiment of the present application proposes a more detailed voice processing method. The voice processing method proposed in the embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0129] See Figure 3 , Figure 3 A flow chart of a speech processing method provided by an exemplary embodiment of the present application is shown; Figure 3 The voice processing method shown can be executed by a computer device, such as a terminal, or a computer device including a terminal and a server. The voice processing method may include but is not limited to steps S301-S303:
[0130] S301: Receive a voice signal and display an interactive dialogue interface.
[0131] Among them, the voice signal refers to the user instruction issued by the user who wants to perform voice control on the computer device. The voice signal is used to express the user's intention to perform voice control on the computer device; for example, the intention expressed by the voice signal is to control the volume of the computer device to increase, or to control the computer device to search for TV series, etc. The embodiment of the present application supports the computer device to receive or obtain the user's voice signal based on near-field communication or long-range communication (or called far-field communication). The following is an introduction to the process of the computer device receiving the user's voice signal based on near-field communication or long-range communication, wherein:
[0132] (1) The computer device receives the user's voice signal based on near-field communication, which means that when the user and the computer device are in the same physical environment, the computer device directly collects the sound emitted by the user to receive the user's voice signal; wherein, the computer device and the user are in the same physical environment can be understood as the distance between the user and the computer device is less than the distance threshold (such as 5 meters). Conversely, if the distance between the user and the computer device is greater than or equal to the distance threshold, it is determined that the two are in different physical environments. For example, a signal collector deployed in a computer device; when the signal collection function of the computer device is turned on, the computer device can collect the sound in the near field range through the signal collector, and thus collect the user's voice signal.
[0133] (2) The computer device receives the user's voice signal based on remote communication, which means that when the user and the computer device are in different physical environments, the computer device receives the user's voice signal transmitted by the signal transmission device, so as to receive the user's voice signal. The signal transmission device refers to a terminal in the same physical environment as the user, which can collect the user's voice signal and transmit the collected voice signal to the computer device; if the signal transmission device is Figure 2b Terminal 204 in the system shown.
[0134] It can be seen that the embodiments of the present application provide two methods for collecting user voice signals, so that users can quickly realize voice control of computer devices in either situation, meet the user's personalized voice control needs, and to a certain extent broaden the usage scenarios of the embodiments of the present application.
[0135] Furthermore, after receiving the user's voice signal, if the computer device determines, based on the user's intent expressed in the voice signal, that the question result information needs to be presented in an interactive dialogue format, the computer device displays an interactive dialogue interface. The interactive dialogue interface is a user interface (UI) provided by the computer device to facilitate human-computer interaction between the user and the computer device. The UI is the medium for interaction and information exchange between the operating system in the computer device and the user, enabling the conversion between the internal form of information and a form acceptable to humans.
[0136] Among them, the interactive dialogue interface includes an interactive dialogue area and a content display area. The interactive dialogue area is used to display the interactive dialogue flow information triggered by the user's voice signal; that is, it is used to present the flow information of the interactive dialogue between the user and the computer device, and the interactive dialogue flow information includes at least one round of interactive dialogue information. The content display area is used to display content information related to the interactive dialogue flow information; that is, the content display area can be understood as a display area opened up in the display screen of the computer device (specifically, the interactive dialogue interface), and the display area is used to present content information that supplements the interactive dialogue flow information, so as to enrich the information available for user reference in the interactive dialogue interface.
[0137] For example, a schematic diagram of displaying an interactive dialogue interface on a display screen corresponding to a computer device (the computer device itself is integrated with a display screen, or the computer device projects a virtual screen (such as a projector projects a virtual screen on a wall)) can be found in Figure 4a ;like Figure 4a As shown, the interactive dialogue interface 401 is divided into a left and right structure, with the left side being the interactive dialogue area 402 and the right side being the content display area 403 (or the left side being the content display area and the right side being the interactive dialogue area).
[0138] based on Figure 4a The following two points need to be explained about the exemplary interactive dialogue interface shown:
[0139] ① In addition to the interactive dialogue interface Figure 4a As shown, in addition to being displayed in the form of an independent user interface (which can make full use of the display screen); in some scenarios (such as using voice control to search for movies in a media resource application), in order to help users better understand that the interactive dialogue interface is a temporary dialogue interface evoked from a computer device, the embodiment of the present application also supports the interactive dialogue interface to be displayed in the form of a floating window. Figure 4bAs shown, the interactive dialogue interface 401 is displayed in the form of a floating window on the service interface of the computer device (such as the system interface of the computer device or the application interface in the media resource application), and the display area of the floating window is smaller than the display area of the service interface provided by the computer device. Figure 4b The floating window form shown is used as an example to introduce it, and is specially explained here.
[0140] ② The display mode of interactive dialogue area and content display area in the interactive dialogue interface is not limited to Figure 4a and Figure 4b The horizontal arrangement display mode shown in the figure may also include but is not limited to: vertical arrangement display and mosaic display. Vertical arrangement display refers to displaying two areas along the vertical direction of the display screen of the computer device; Figure 4c As shown in the figure, the vertical direction of the display screen includes an interactive dialogue area and a content display area. Mosaic display means that one area is displayed on the outer edge of another area to form a mosaic style; Figure 4d As shown in FIG (1), the interactive dialogue area 402 is the rectangular frame area in the middle of the display screen, and the content display area 403 is the display area between the four sides of the rectangular frame area and the border of the display screen; Figure 4d In the attached figure (2), the interactive dialogue area 402 is the rectangular frame area on the upper left side of the display screen, and the content display area 403 is the display area between the right and lower sides of the rectangular frame area and the border of the display screen. Figure 4a The horizontal arrangement display shown is used as an example for introduction and is specifically explained here.
[0141] S302: Displaying interactive dialogue flow information triggered by the voice signal in the interactive dialogue area.
[0142] The interactive conversation flow information triggered by a voice signal refers to the flow information formed by the interactive conversation information corresponding to at least one interactive conversation round initiated by the voice signal; this flow information includes the interactive conversation information corresponding to each interactive conversation round. For example, after receiving a voice signal, a computer device converts the voice signal into question description information in the first interactive conversation round. The computer device generates question result information based on the question description information, forming the first interactive conversation round. If the user wishes to continue the conversation, the computer device collects voice signals for a second interactive conversation round and forms the second interactive conversation round using the same method as the first interactive conversation round. Repeating these steps can obtain the interactive conversation flow information triggered by the voice signal.
[0143] In the embodiment of the present application, two result styles are provided for the interactive dialogue flow information, specifically two information types of question result information are provided, namely, graphic type and plain text type. Among them: ① When the information type of the interactive dialogue information is graphic type, the interactive dialogue information includes modal information such as text and image at the same time; so as to provide users with rich search results in a multimodal manner. For a schematic diagram of a graphic type interactive dialogue information, see Figure 5a ;like Figure 5a As shown, the question result information included in a round of interactive dialogue information 501 includes both images and text. ② When the information type of the interactive dialogue information is pure text type, the interactive dialogue information does not include models such as images, such as only text, links or icons, which do not occupy a large display area. A schematic diagram of a pure text type interactive dialogue information is shown in FIG. Figure 5b ;like Figure 5b As shown, the question result information included in a round of interactive dialogue information 502 only contains text. It should be noted that the result formats of the interactive dialogue flow information are not limited to the above two; the above two result formats do not limit the embodiments of the present application.
[0144] S303: Displaying content information related to the interactive dialogue information in the content display area.
[0145] The content information related to the interactive dialogue information displayed in the content display area is specifically content information related to any round of interactive dialogue information in the interactive dialogue flow information; the content information is different from the any round of interactive dialogue information, but is relevant in content.
[0146] Specifically, assuming that the interactive dialogue area displays the i-th round of interactive dialogue information among N rounds of interactive dialogue information, that is, all or part of the information in the i-th round of interactive dialogue information is displayed in the interactive dialogue area, and all or part of the display area in the interactive dialogue area is occupied by the i-th round of interactive dialogue information, i is an integer, and 1≤i≤N, then the content information related to the interactive dialogue flow information displayed in the content display area includes at least one of the following:
[0147] ① One or more candidate question descriptions that match the conversational intent of the i-th round of interactive dialogue information. The conversational intent of the i-th round of interactive dialogue information refers to the user's intention for engaging in the i-th round of interactive dialogue. Candidate question descriptions that match the conversational intent of the i-th round of interactive dialogue information are question descriptions that are relevant to the conversational intent and available for user selection. Providing this candidate question description information to the user can help improve the user's description of their intent if the currently output question result information does not meet the user's intent, thereby improving the accuracy of the searched and generated question result information and, in turn, enhancing the quality of the conversation.
[0148] For example, a schematic diagram showing candidate question description information in the content display area can be seen in Figure 6 ;like Figure 6 As shown, assuming that the question description information of the i-th round of interactive dialogue information displayed in the interactive dialogue area is how the weather is, and the question result information is the temperature is 20 degrees; then the intentions expressed by the candidate question description information 601 and the candidate question description information 602 displayed in the content display area match the dialogue intention of the i-th round of interactive dialogue information. For example, the candidate question description information 601 and the candidate question description information 602 are both related to the intention of asking about the weather. For example, the candidate question description information 601 may be whether the temperature will drop, and the candidate question description information 602 may be how to dress in 20-degree weather, etc.
[0149] ② Target content information related to the target dialogue information in the i-th round of interactive dialogue information; wherein the target content information includes at least one of the following: detailed information of the target dialogue information, source information of the target dialogue information, and commentary information of the target dialogue information. The target dialogue information is the information selected (by default or user-defined) in the i-th round of interactive dialogue information that needs to be supplemented (i.e., related content information is displayed in the content display area). The target dialogue information is information that describes the problem or information that contains the result of the problem; this embodiment of the application is not limited to this. The detailed information of the target dialogue information can be understood as information that further explains the target dialogue information; the source information of the target dialogue information includes source information such as web pages or journals referenced when generating the target dialogue information (such as journal names, web page addresses, etc.); the commentary information of the target dialogue information can refer to evaluation information generated by a computer device (specifically, a server) for the target dialogue information. For example, if the target dialogue information is a video media resource, the commentary information of the target dialogue information can be the reasons for recommending the video media resource generated by the computer device, etc.
[0150] It is worth noting that the target dialogue information can be a default one or customized by the user in the interactive dialogue area. The following describes two methods for determining the target dialogue information in the interactive dialogue information of round i as an example, where:
[0151] (1) The target dialogue information is the default.
[0152] This embodiment of the application supports pre-configuration of one or more types of target dialogue information. When a pre-defined type of target dialogue information appears in the question description information or question result information in the interactive dialogue area, the content information related to the target dialogue information is displayed in the content display area by default. Through this pre-defined method, important dialogue information during the interactive dialogue can be automatically further explained in the content display area, improving the intelligent and automated content display of the interactive dialogue and enhancing the user's interactive experience.
[0153] Among them, depending on the interactive dialogue scenario, the predefined types of target dialogue information may be different. For example, in a media resource scenario (such as a video scenario), it is predefined that the target dialogue information of the video type needs to display content information in the content display area. Then, when the target dialogue information of the video type (such as the resource image corresponding to the video media resource - a movie or TV drama poster) appears in the interactive dialogue area, the default is to directly display the content information related to the target dialogue information of the video type in the content display area. For another example, in a music scenario, it is predefined that the target dialogue information of the music type needs to display content information in the content display area. Then, when the target dialogue information of the music type appears in the interactive dialogue area, the default is to directly display the content information related to the target dialogue information of the music type (such as the music cover corresponding to the music resource) in the content display area.
[0154] The following is combined with Figure 7 Taking the target conversation information of a predetermined video type in the media resource scenario as an example, and the need to display content information related to the target conversation information of the video type in the content display area, the process of displaying content information related to the target conversation information by default in the content display area is introduced. Figure 7As shown, assume that the interactive dialogue information for round i includes first question description information and first question result information, and that the first question result information includes a resource image 701 corresponding to a video media resource (i.e., the target dialogue information displayed in the interactive dialogue area)—a video poster. Then, the computer device determines that the resource image displayed in the interactive dialogue area is defaulted to the target dialogue information, and then displays the target content information associated with resource image 701 in the content display area by default. This target content information is at least one of one or more content information associated with the target dialogue information, such as, for example, the target content information includes commentary information about the target dialogue information—recommendation reason information 702 for resource image 701.
[0155] Furthermore, the embodiment of the present application also supports preview display or playback of video media resources corresponding to resource images in the interactive dialogue area in the content display area in the media resource scenario, so as to help users make better decisions and enhance the user's video search experience by previewing video media resources in the content display area. As mentioned above, the content information related to the target dialogue information displayed in the content display area may include detailed information; in the case where the target dialogue information in the media resource scenario is a resource image corresponding to a video media resource, the detailed information related to the resource image may include the video media resource corresponding to the resource image, that is, the target content information related to the resource image displayed in the content display area is the video media resource itself. In this case, the computer device plays the video media resource in the content display area through a player (such as automatically playing or playing when clicked by the user).
[0156] During the playback of video media resources in the content display area, the embodiment of the present application also supports users to perform operations such as pausing or continuing the playback of the video media resources played in the content display area according to their own playback needs, so as to meet the user's personalized video browsing needs. Specifically, during the playback of video media resources in the content display area, the computer device receives a pause playback operation for the video content, such as the pause playback operation can be but is not limited to a trigger operation on the pause component (or option, button, control) in the player, or a click operation on any display position in the playback interface of the video media resource. The computer device will respond to the pause playback operation on the video media resource and pause the playback of the video media resource in the content display area. Next, if the computer device detects the user's continue playback operation on the video media resource, the video media resource will continue to be played in the content display area.
[0157] Furthermore, in the media resource scenario, in order to help users quickly play the video media resources when searching for the video media resources they want to watch, the embodiment of the present application also supports displaying a video playback entrance for triggering large-screen playback of video media resources in the content display area (the entrance is displayed as a control, component, button or option on the interface, etc.), thereby realizing a quick jump from the interactive dialogue interface to the media playback interface of the video media resource; compared to requiring the user to exit the interactive dialogue interface and then search for video media resources in the media resource application, the playback path of the video media resource is significantly shortened, thereby improving the user's video media resource search experience.
[0158] An exemplary diagram of previewing a video media resource in a content display area and triggering large-screen playback of the video media resource from the content display area can be found in Figure 8 .like Figure 8 As shown, when the interactive dialogue information of the i-th round includes a resource image corresponding to a video media resource (e.g., resource image 801), and resource image 801 is in focus (i.e., the cursor on the display screen selects resource image 801), the target content information associated with resource image 801, namely, video media resource 802, is displayed in the content display area, and video media resource 802 is played. The user can perform a pause or resume playback operation to independently control video media resource 802. In addition, a video playback entry 803 for video media resource 802 is also displayed in the content display area. When video playback entry 803 is triggered, indicating that the user wants to play video media resource 803 on a large screen (i.e., the display screen plays the video media resource in full screen), the computer device closes the interactive dialogue interface and displays a media playback interface 804 for video media resource 802, so that video media resource 802 can be played in full screen on the media playback interface 804 (e.g., play from the beginning, or play from the playback progress in the content display area).
[0159] (2) The target dialogue information is customized by the user in the interactive dialogue area.
[0160] The embodiment of the present application supports users to customize the target conversation information that needs to display content information in the interactive conversation area according to their own content viewing needs during the interactive conversation process; this gives users more information selection rights, can meet the user's needs for personalized content information viewing, and significantly improve the user's interactive conversation experience.
[0161] In a specific implementation, if a user needs to supplement specific conversation information in the i-th round of interactive conversation information in the interactive conversation area, the user performs an information viewing operation on the specific conversation information, and the specific conversation information is determined as the target conversation information. In this case, in response to the information viewing operation performed on the target conversation information in the i-th round of interactive conversation information in the interactive conversation area, the computer device displays the target conversation information in a selected state in the interactive conversation area, prompting the user to confirm the selected target conversation information. The computer device then displays target content information related to the target conversation information in the content display area. Herein, the following: ① The information viewing operation performed by the user on the target conversation information in the interactive conversation area can be understood as a selection operation on the target conversation information; selection operations include, but are not limited to: selecting the target conversation information → pressing a key → confirming an information confirmation option in an options window; or selecting the target conversation information using an input device such as an electronic pen or mouse. ② The styles of displaying the target conversation information in the interactive conversation area as selected may include, but are not limited to: darkening the background color of the target conversation information; displaying a prompt label adjacent to the target conversation information, the prompt label used to prompt the user that the target conversation information needs to be supplemented; and so on.
[0162] The following describes the operation process of the information viewing operation described above and the exemplary style of the target dialogue information being in the selected state in conjunction with the accompanying drawings. Figure 9a As shown, assuming that the selection operation for the target conversation information is as follows: selecting the target conversation information → pressing a key → confirming the information confirmation option in the options window; if the computer device is a smart computer or smart TV, the target conversation information 901 can be selected by controlling the cursor on the computer device's display screen using a mouse or remote control. Then, a key operation is performed on the mouse or remote control (e.g., right-clicking the mouse, pressing a button on the remote control), and an options window 902 is displayed. This options window 902 includes an information confirmation option 903 for triggering the display of target content information related to the target conversation information 901 in the content display area. The user can continue to control the cursor using the mouse or remote control to confirm information confirmation option 903. At this point, it is determined that the user needs to further supplement the target conversation information 901 in the content display area, and the target conversation information 901 is displayed in a selected state in the interactive conversation area, such as with a darkened background color 904 or a prompt label 905 displayed adjacent to the target conversation information.
[0163] like Figure 9bAs shown, it is assumed that the selection operation of the target dialogue information is: the selection operation of the target dialogue information through an input device such as an electronic pen or a mouse; in the case where the computer device is a touch-screen device such as a smart computer, the target dialogue information 901 can be selected by using an electronic pen or a finger on the display screen of the computer device. For example, the selection operation is a circle operation. At this time, the computer device determines that the user needs to further supplement the target dialogue information 901 in the content display area, and displays the target dialogue information 901 as a selected state in the interactive dialogue area. The selected state is shown in FIG. Figure 9a The above has been explained and will not be repeated here.
[0164] Based on the description of the aforementioned implementation methods (1) and (2), it can be seen that the embodiment of the present application supports predefined target dialogue information or user-defined target dialogue information. When the number of content information related to the target dialogue information in the i-th round of interactive dialogue information (including target content information and candidate question description information) is at least two, the embodiment of the present application supports the one-time display of at least two content information related to the target dialogue information in the content display area; that is, at least two content information are displayed in the content display area at the same time, such as Figure 9b The content display area shown displays content information 906 and content information 907 related to the target conversation information at the same time, so as to facilitate the user to view the content information.
[0165] Considering that there is a large amount of content information related to the target conversation information, the limited display area of the content display area cannot simultaneously display multiple content information. Based on this, embodiments of the present application also support switching the display of multiple content information related to the interactive conversation flow information in the content display area. Specifically, switching the display of multiple content information related to the i-th interactive conversation information displayed in the interactive conversation area. This not only ensures that each piece of content information is fully displayed in the content display area, but also ensures flexible control over the display and hiding of multiple content information. In a specific implementation, assuming that there is at least one piece of content information related to the interactive conversation flow information, the content information may specifically include: content information related to the i-th interactive conversation information (including candidate question description information matching the conversation intent of the i-th interactive conversation information, and target content information related to the target conversation information in the i-th interactive conversation information). In this case, the content display area includes at least one content type option, each content type option corresponding to a type of content information. In this way, in response to a user selecting a target content type option, the computer device displays the content information corresponding to the selected target content type option in the content display area. The target content type option is any one of the at least one content type option.
[0166] For example, a schematic diagram of the content type options in the content display area can be found in Figure 10 ;like Figure 10 As shown, the content display area includes content type options 1001, 1002, and 1003. Content type option 1001 corresponds to the detailed information of the target conversation information, content type option 1002 corresponds to the candidate question description information that matches the conversational intent of the i-th interactive conversation round, and content type option 1003 corresponds to the commentary information of the target conversation information. When a user wishes to view any content information of the target conversation information, they can select the content type option corresponding to that content information, such as content type option 1003. The detailed information of the target conversation information corresponding to content type option 1003 will be displayed in the content display area. Furthermore, if there are multiple target conversation information in the i-th interactive conversation round, when content type option 1003 is selected, detailed information related to each target conversation information will be displayed in the content display area. This facilitates the user to view the detailed information of multiple target conversation information in batches, improving content information viewing efficiency.
[0167] The interactive dialogue area and the content display area in the interactive dialogue interface provided by the embodiment of the present application are linked to each other. The linked display between the interactive dialogue area and the content display area means that: when most of the display area of the interactive dialogue area (such as more than 1 / 2 of the display area) displays the interactive dialogue information of the i-th round, even if the user dynamically views the interactive dialogue information of the i-th round in the interactive dialogue area, the content information related to the interactive dialogue information of the i-th round is always displayed in the content display area; when most of the display area of the interactive dialogue information of the interactive dialogue switches from displaying the interactive dialogue information of the i-th round to displaying the interactive dialogue information of the j-th round, the content information related to the interactive dialogue information of the j-th round is switched to be displayed in the content display area. This method of linked display between the interactive dialogue area and the content display area can keep the content information related to the interactive dialogue information of the i-th round displayed in the content display area when the user is viewing the interactive dialogue information of the i-th round, avoiding problems such as information misalignment caused by the irrelevant content information displayed in the content display area and the interactive dialogue information displayed in the interactive dialogue area.
[0168] The following is combined with Figure 11 , taking the interactive dialogue information of round i and round j in N interactive dialogue information as an example, the process of linkage display between the interactive dialogue area and the content display area mentioned above is introduced; wherein j is a positive integer, j≠i, and 1≤j≤N. Figure 11As shown, during a dialog viewing operation for interactive dialog information 1101 in the interactive dialog area, content information 1102 related to the interactive dialog information in the round i is maintained displayed in the content display area. Specifically, during a dialog viewing operation in the interactive dialog area, if interactive dialog information 1101 in the round i consistently occupies a majority of the display area in the interactive dialog area, it is determined that the dialog viewing operation in the interactive dialog area is viewing information in the round i, and content information 1102 related to the interactive dialog information in the round i is maintained displayed in the content display area. When a dialog viewing operation is detected in the interactive dialog area switching from interactive dialog information 1101 in the round i to interactive dialog information 1103 in the interactive dialog area, specifically, as the dialog query operation continues, the display area of interactive dialog information 1103 in the interactive dialog area becomes larger than the display area of interactive dialog information 1101 in the interactive dialog area, the computer device updates the content information related to interactive dialog information in the round i to content information 1104 related to interactive dialog information in the round j.
[0169] Among them, the conversation viewing operations performed in the interactive conversation area may include at least one of the following: page turning operations (i.e., the interactive conversation information in the interactive conversation area simulates the style of turning pages left and right or up and down in the real world to turn pages and display cross-type conversation information), scrolling operations (such as controlling the mouse or buttons in the remote control to scroll and display cross-type conversation information) and dragging operations (such as the sliding axis in the interactive conversation area to slide and display the interactive conversation information), etc.
[0170] For example, if the conversation viewing operation is a scrolling operation, j = i + 1, and the information type of the interactive conversation information in round i is plain text, and the information type of the interactive conversation information in round j is graphic and text, then according to the page-turning operation, half of the screen is flipped upward in the interactive conversation area to implement page-turning and update display of the interactive conversation flow information, specifically, page-turning and update display of the interactive conversation information in round i. During the continuous scrolling operation, if the interactive conversation information in round i always occupies the majority of the display area of the interactive conversation area, then the content display area always displays content information related to the interactive conversation information in round i. Furthermore, as the scrolling operation continues, the interactive conversation information in round j appears in the interactive conversation area, and if the interactive conversation information in round j occupies a larger display area of the interactive conversation area than the interactive conversation information in round i, then the conversation information related to the interactive conversation information in round j is displayed in the content display area. It is worth noting that the information type of the j-round interactive dialogue information is a graphic and text type. For example, in the media resource scenario, resource images and text coexist in the j-round interactive dialogue information. Considering that resource images are more likely to be information of interest to users than plain text, in the process of updating and displaying the j-round interactive dialogue information, the resource image in the j-round interactive dialogue information in the interactive dialogue area is given priority focus, that is, the resource image is automatically determined as the third dialogue information.
[0171] It should be noted that Figure 11 This is merely an exemplary operation process for performing a conversation viewing operation in an interactive conversation area provided for an embodiment of the present application, and does not limit the embodiment of the present application.
[0172] As previously mentioned, the interactive dialogue flow information includes N rounds of interactive dialogue information, and multiple rounds of interactive dialogue information may be associated with content information. To facilitate users to quickly locate content information related to any round of interactive dialogue information in the content display area, embodiments of the present application support rapid location of any round of interactive dialogue information in the interactive dialogue area, as well as rapid location of content information related to any round of interactive dialogue information in the content display area. This significantly improves the efficiency of locating content information or interactive dialogue information, particularly when N is a large value, allowing users to flexibly locate interactive dialogue information and content information according to their information viewing needs, thereby enhancing their interactive dialogue experience.
[0173] In a specific implementation, the content display area includes dialogue options corresponding to each round of interactive dialogue information in the interactive dialogue flow information; any dialogue option is used to quickly index the corresponding interactive dialogue information in the interactive dialogue area, and quickly index the content information related to the corresponding interactive dialogue information in the content display area. In this implementation, the computer device responds to the user's triggering operation on the target dialogue option in the content display area, displays the content information related to the target interactive dialogue information corresponding to the target dialogue option in the content display area; and displays the target interactive dialogue information in the interactive dialogue area. Figure 12a As shown, assuming N=3, the content display area includes dialogue option 1201 corresponding to the first round of interactive dialogue information, dialogue option 1202 corresponding to the second round of interactive dialogue information, and dialogue option 1203 corresponding to the third round of interactive dialogue information. Furthermore, the second round of interactive dialogue information is currently displayed in the interactive dialogue area, and content information related to the second round of interactive dialogue information is displayed in the content display area. When the user triggers dialogue option 1203 in the content display area, content information related to the third round of interactive dialogue information is displayed in the content display area. Furthermore, the third round of interactive dialogue information is displayed in the interactive dialogue area.
[0174] based on Figure 12a Two points need to be made about the schematic shown:
[0175] ① The dialogue option and the aforementioned content type option can exist simultaneously in the content display area. When the dialogue option and the content type option exist simultaneously in the content display area, when any dialogue option is selected, the user is also allowed to switch the display of different content information related to the interactive dialogue information corresponding to the dialogue option by performing a selection operation on the content type option. Figure 12b As shown, when dialogue option 1202 is selected, the user can continue to select content type options in the content display area to switch between displaying multiple content information related to the second round of interactive dialogue information. This coexistence of dialogue options and content type options greatly increases the user's flexibility in viewing interactive dialogue information and content information from both the dialogue turn dimension and the content information dimension.
[0176] ② As previously mentioned, the number of target conversation information in the i-th round of interactive conversation information may be at least two. This embodiment of the present application allows users to quickly display target content information related to the specified target conversation information in the content display area by performing a content location operation on the target conversation information (specified by the user) in the i-th round of interactive conversation information. This facilitates users viewing the interactive conversation flow information in the interactive conversation area without having to repeatedly view the target conversation information. Instead, the target conversation information is cached, allowing users to quickly retrieve content information related to a specific target conversation information from the interactive conversation area. In a specific implementation, when a content location operation is performed on the target conversation information in the i-th round of interactive conversation information in the interactive conversation area, the target content information related to the specified target conversation information is displayed in the content display area. The content location operation performed on the target conversation information may include, but is not limited to, double-clicking, long-pressing, or single-clicking the target conversation information, which is not limited in this embodiment of the present application.
[0177] In summary, the embodiments of the present application implement partitioned content display in the interactive dialogue interface. Specifically, the interactive dialogue flow information triggered by voice signals is displayed in the interactive dialogue area. This helps users more easily determine the results of questions through question description information and question result information, thereby enhancing the user's question-and-answer experience. Content information related to the interactive dialogue flow is displayed in the content display area, further supplementing the interactive dialogue flow information. This helps users make decisions based on both the interactive dialogue flow information in the interactive dialogue area and the content information displayed in the content display area, effectively improving the quality and efficiency of dialogue.
[0178] See Figure 13 , Figure 13 A flow chart of another speech processing method provided by an exemplary embodiment of the present application is shown; Figure 13 The voice processing method shown can be executed by a computer device, such as a terminal, or a computer device including a terminal and a server. The voice processing method may include but is not limited to steps S1301-S1307:
[0179] S1301: Display the service interface of the resource playback device.
[0180] Among them, the resource playback device provides a globally triggerable voice interaction method; that is, multiple service interfaces of the resource playback device are deployed with voice interaction entrances, which improves the multi-path triggering of voice interaction and enhances the flexibility of triggering voice interaction compared to fixed voice interaction entrance positions. Specifically, a media resource application is deployed in the resource playback device, and an intelligent assistant that integrates the voice processing method provided by the embodiment of the present application is deployed in the media resource application; the embodiment of the present application supports providing voice interaction entrances in each service interface of the resource playback device to achieve globally triggerable voice interaction. Among them, the service interface of the resource playback device includes: the system interface of the resource playback device and the application interface of the media resource application; the system interface of the resource playback device refers to the interface provided by the operating system deployed in the resource playback device, and the application interface of the media resource application refers to the interface provided by the service provider of the media resource application.
[0181] In order to facilitate users to perceive the various stages or progress of voice interaction with the resource playback device, the embodiment of the present application supports displaying voice prompt information in the service interface of the resource playback device, and allows users to perceive the progress or stage of voice interaction through the various states of the voice prompt information. Among them, each state includes: guidance state → awakening state → recognition state → understanding state → execution state (that is, the interactive dialogue interface is displayed for dialogue tasks, and the task execution results are displayed for direct tasks). Among them, the voice prompt information may refer to a voice bar based on automatic speech recognition technology (Automatic Speech Recognition, ASR); ASR technology is a technology that automatically converts human voice signals into corresponding text through AI algorithms, aiming to enable resource playback devices to understand the semantic content expressed by voice signals, thereby realizing speech-to-text (STT) conversion. In this way, the voice prompt information can convert the collected user voice signal into corresponding text (such as question description information) based on ASR technology.
[0182] Specifically, when a user initially enters the service interface of a resource playback device, the service interface includes a voice prompt message in a prompt state. This prompt state can be referred to as a guidance state, and the voice prompt message in the prompt state serves to indicate that the resource playback device has a voice interaction function. If the user desires to interact with the resource playback device through voice, such as invoking a media resource application running on the resource playback device to perform a video search, the user can perform a confirmation operation in response to the voice prompt message in the prompt state, triggering execution of step S1302.
[0183] S1302: In response to a confirmation operation on the voice prompt information in the prompt state, the voice prompt information is converted from the prompt state to the wake-up state.
[0184] When the resource playback device detects a user confirmation action in response to a voice prompt in the prompt state and determines that the user intends to engage in voice interaction, the resource playback device transitions the voice prompt from the prompt state to the wake-up state. The wake-up state can be referred to as the wake-up state. The voice prompt in the wake-up state serves to indicate that the resource playback device is in the voice collection state, so that the user can perceive that the voice prompt is in the wake-up state and begin outputting voice signals.
[0185] It should be noted that the confirmation operation performed by the user on the voice prompt information in the prompt state can be achieved through near-field communication or remote communication; wherein, the relevant content of near-field communication or remote communication can be referred to the relevant description of the relevant content shown in the aforementioned step S301, which will not be repeated here. For example, in the case where the resource playback device is a smart TV, if it is near-field communication, the user can press a button on the remote control of the smart TV to realize the confirmation operation of the voice prompt information in the prompt state; if it is remote communication, when the resource playback device receives the confirmation request transmitted by the signal transmission device on the user side, it determines that the user has performed the confirmation operation on the voice prompt information in the prompt state.
[0186] S1303: When the resource playback device starts to collect voice signals, the voice prompt information is converted from the awakening state to the collecting state.
[0187] When the voice prompt information is in the awake state, the resource playback device detects whether there is a voice signal; if the resource playback device begins to collect the voice signal, the resource playback device converts the voice prompt information from the awake state to the collection state. The collection state can be called the recognition state, and the voice prompt information in the collection state is used to prompt the resource playback device that the voice signal is being collected. The resource playback device can collect voice signals by collecting voice signals emitted by users in the physical environment based on near-field communication, or by receiving voice signals transmitted by a signal transmission device on the user side based on long-distance communication.
[0188] S1304: When the resource playback device finishes collecting the voice signal, the voice prompt information is converted from the collection state to the understanding state.
[0189] When the resource playback device has successfully collected the voice signal, the resource playback device will convert the voice prompt from the collection state to the understanding state; wherein the understanding state can be called the understanding state, and the voice prompt information in the understanding state is used to prompt the resource playback device to perform result search processing based on the voice signal.
[0190] To facilitate understanding of the flow between the various states of the voice prompt information described in steps S1301-S1304, the following Figure 14 A diagram showing the voice prompts in different states. Figure 14 As shown, a voice prompt message 1401 in a guiding state is displayed in the service interface of the resource playback device; the text content included in the voice prompt message 1401 in the guiding state is used to guide the user to perform operations or search for movies through voice control of the resource playback device. The text content can be generated in combination with platform operations, object data (such as the user's online appointment behavior, continued viewing of videos, etc.) and video content (such as actors, commentators, etc.), and the embodiment of the present application does not limit this. When it is detected that the user performs a confirmation operation on the voice prompt message 1401 in the guiding state, the voice prompt message 1401 is switched from the guiding state to the awakening state to prompt the user to start outputting the voice signal. When the resource playback device starts to collect the voice signal, the voice prompt message 1401 in the awakening state will be converted to the recognition state to prompt the user that the voice signal is currently being collected. When the voice signal collection is completed, the voice prompt message 1401 in the recognition state will be converted to the understanding state to inform the user that the voice signal is being searched for results. It should be understood that, Figure 14 The information style, display location and specific content of the voice prompt information shown are all examples and do not limit the embodiments of the present application.
[0191] Among them, the result search processing for the voice signal mainly includes operations such as intent recognition processing and result generation processing for the voice signal, which aims to generate problem result information of the problem description information corresponding to the voice signal to achieve a reply to the user query. In order to achieve more accurate interactive dialogue, the embodiment of the present application supports the result search processing of the voice signal based on the large language model, which aims to generate more accurate problem result information to improve the dialogue quality and dialogue efficiency. In other words, this solution provides an AI voice landing solution through a generative large language model (such as an LLM model); with the help of the better intent understanding ability of the large language model, the generative model is fine-tuned in combination with specific business data (business data in the media resource scenario), so that the fine-tuned generative model can be better combined with the business (such as the business of searching for media resources in the media resource scenario), and can effectively realize content search on resource playback devices, especially OTT devices. In this way, on the one hand, in the interactive dialogue scenario, the fine-tuned generative model can recognize the intent of complex user queries and unseen user queries, thereby improving generalization capabilities. On the other hand, relying on the long text processing capabilities of the generative model, it can understand the text content corresponding to longer voice signals, improving the accuracy of intent recognition and content recall.
[0192] In a specific implementation, the general process of the embodiment of the present application for performing result search processing on the voice signal based on the fine-tuned generative model may include: first, obtaining basic analysis data, which is basic data related to the business to which the voice processing method is applied, including but not limited to: scene data (such as application data of the media resource application, etc.), object data (such as the user's historical identity data, behavior data, etc.), permission data (such as when the user is a member of the media resource application, his permission is greater than that of a non-member) and intent list data (such as the various intentions expressed by the user when searching for movies in the media resource application during the historical time). Then, the fine-tuned generative model is called to perform intent recognition processing on the voice signal based on the basic analysis data to obtain the intent recognition result; specifically, the basic analysis data is given to the fine-tuned generative model, so that the generative model can quickly analyze the basic analysis data to obtain the intent recognition result, and the intent recognition result indicates a voice task, which is the user intention corresponding to the voice signal input by the user.
[0193] Finally, the voice task is executed according to the intention recognition result. If the intention recognition result indicates that the voice task indicated by the voice signal is a dialogue task, that is, voice interaction with the user is required to realize or achieve the intention, then the result generation processing is performed according to the intention recognition result, and question result information corresponding to the voice signal is generated, and the execution of steps S1305-S1307 is triggered; wherein, the generated question result information and the question description information corresponding to the voice signal constitute a round of interactive dialogue information included in the interactive dialogue flow information in the interactive dialogue area. In the media resource scenario, during the result generation processing according to the intention recognition result, the intention recognition result will continue to be classified in more detail; and the result generation processing is performed according to the more detailed intention classification result to generate question result information corresponding to the voice signal.
[0194] On the contrary, if the intention recognition result indicates that the voice task indicated by the voice signal is a direct task, that is, there is no need to interact with the user by voice, and only the corresponding task processing needs to be performed directly, then the direct task is processed and the task processing result (or called direct result) is presented in the resource playback interface; it is worth noting that the embodiment of the present application also supports the use of a fine-tuned generative model to perform task processing on direct tasks, aiming to improve the efficiency of task execution through the generative model and directly process the task results. Among them, direct tasks may include at least one of the following: control tasks, interface switching tasks, and player control tasks, etc. Among them: ① Direct tasks are control tasks, and control tasks indicate the control of the device capabilities of the resource playback device, such as volume control, brightness control, and barrage control, etc. The corresponding task processing results can be experienced through voice prompt information. Such as Figure 15aAs shown, assuming that the intention recognition result indicates that the user intends to turn off the barrage on the display screen, the resource playback device can automatically turn off the barrage on the display screen and display the task execution result in the voice prompt information 1501 to prompt that the barrage has been turned off. ② Direct tasks are interface switching tasks that can be implemented across applications on the resource playback device. The application can refer to different applications running in the resource playback device or different application interfaces in the same application. Figure 15b As shown, assuming that the user's intention indicated by the intention recognition is to jump to the news channel, the channel interface corresponding to the news channel is directly opened in the resource playback device, and the task execution result is displayed in the voice prompt information 1502 to prompt to open the news channel. ③ Direct-access tasks are player control tasks, which can directly open the playback interface of the video media resource when the video media resource is uniquely determined, thereby achieving fast resource playback. Figure 15c As shown, assuming that the intention recognition result indicates that the user intends to open the video media resource, the resource playback device opens the player and plays the video media resource, and displays the task execution result in the voice prompt information 1503 to prompt that the video media resource has been successfully opened.
[0195] The following is combined with Figure 16 , the detailed process of the result search processing for the voice signal described above is introduced: After the user inputs the voice signal through the media resource application, the user request (i.e., user query or voice signal) is determined to enter the access layer (corresponding to the recognition state of the voice prompt information). The access layer will call the AI voice logic layer to perform user request analysis, scenario analysis, rights analysis, and intent analysis (such as also including general logic such as text correction and word segmentation processing) to obtain basic analysis data; among them, the AI voice logic layer is mainly used to implement some UI component construction logic in the service interface to help the access layer complete the analysis of general logic and obtain basic analysis data.
[0196] The voice signal and basic analysis data are then input into the intent decision layer, which is connected to the large language model service provided by the embodiments of this application and the traditional intent capability service of the resource playback device. In this way, the intent decision layer will analyze the voice signal and basic analysis data and use the traditional intent capability service or the large language model service to process the user query. Among them, traditional intent service capabilities may include task execution services, template capabilities, model service capabilities, object data services, search services, and scene matching services. The task execution service can implement task execution for direct tasks. The template capability supports rapid intent recognition through string matching (such as exact string matching, fuzzy string matching, exact slot matching, and fuzzy slot matching). The model service capability can use some traditional models (such as the BERT model) to identify the intent of relatively simple voice signals (such as those with relatively short question description information). The object data service supports returning the user's intent based on the user's historical voice usage. The search service can support querying video media resources through search capabilities for relatively simple user queries. The scene matching service supports scene matching based on string matching. For example, the homepage of a media resource application contains three types of buttons: application, media resource, and function (button). Each type of button is set with corresponding text content to represent its function. Therefore, a simple string pattern matching can be performed between the text content set on the button and the user query to determine whether there is scene information on the homepage that the user needs to click.
[0197] If the intention decision layer recognizes the use of a large language model service to perform intent recognition processing and result generation processing on the user query, the voice signal (i.e., the user query) and the basic analysis data will be input into the intent model service together. The intention model service is connected to at least one generative model for identifying intent, that is, the embodiment of the present application supports the use of multiple generative models to achieve voice interaction. Generative models may include but are not limited to: multi-intention classification models, intent classification models, and scene matching models, etc. Among them, the scene matching model refers to the scene matching of the user query to achieve intent recognition. The core function of the intent classification model and the multi-intention classification model is to perform intent recognition and slot extraction on the query (Query) input by the user; specifically, by using the intent classification model or the multi-intention classification model to perform semantic analysis on the user query, determine the user's main goal or operation intention, and extract relevant key information (slots) from it to identify the user's intention. Among them, the multi-intention classification model can realize the recognition of multiple intentions compared to the intent classification model. For example, if the intent classification model performs single-intent recognition and the input is "I want to listen to songs by singer", and the recognized slot includes singer (singer_name), then the output intent recognition result is music-singer (music-singer). For another example, if the multi-intent classification model performs multi-intent recognition and the input is "Open the music player to play song A", and the recognized slot includes music player, the corresponding intent is open-app (open-app), and the recognized slot also includes song A, then the corresponding intent is music-song (music-song).
[0198] It is worth noting that in order to improve the accuracy of intent recognition, the resource playback device (specifically, the media resource application in the resource playback device) will determine whether it is a multi-round interactive dialogue based on the user's actual operation after processing the first user query; if it is a multi-round, the resource playback device will carry the dialogue information of the historical user request when the user requests it next time. In this way, the generative model jointly determines the intent type of the next user request based on the dialogue information of multiple rounds of interaction. For example, the exemplary code for the next user request to carry at least one round of historical interactive dialogue information is as follows:
[0199] Round1: "role": user, content: "xxxxQuery: Set an alarm answer:" / / User query request to set an alarm
[0200] Round1”role”:system,content:”{\"intent\":\"CLOCK_REMINDAER / CLOCK\"}
[0201] / / Response to user query to set alarm
[0202] Round1: "role": user, content: "Query:5 o'clock alarm answer:" / / User query request to set the alarm at 5 o'clock
[0203] Round 1: "role": system, content: "CLOCK_REMINDAER / CLOCK" / / respond to user query and set the alarm at 5 o'clock
[0204] The above three generative models corresponding to the intent model service can jointly quickly domain the user intention of the voice signal (that is, quickly identify the scope of the intention, such as film and television or non-film and television), significantly improving the recognition efficiency of intent recognition. Among them, the embodiment of the present application supports fine-tuning or fine-tuning of any generative model connected to the intent model service in combination with business data to improve the combination of the generative model and the specific business, so as to better complete the intent recognition under the specific business and improve the accuracy and recognition efficiency of intent recognition. The following is a detailed introduction to the construction of training data for the generative model and the model training process, including:
[0205] (1) Construct training data.
[0206] The general process of constructing training data can be found in the attached Figure 17 ;like Figure 17 As shown, training data can be constructed by uploading training data yourself, or by building training data based on reference data—prompt templates and parameter datasets (including multiple parameters in media resource scenarios). When training data is constructed based on prompt templates and parameter datasets, a large base model is used to generate preliminary training data that meets the business scenario based on the reference data. This preliminary training data is then cleansed (or corrected) to obtain training data. This training data is then published to a data warehouse and combined with training data obtained through other means (i.e., self-uploaded training data) to obtain the final training data used for model training. The training data is then used to train the generative model (i.e., fine-tune or fine-tune it), and the resulting fine-tuned generative model is published for online use.
[0207] Taking the hybrid model based on the generative model as an example, an exemplary prompt word template includes:
[0208] You are an assistant that performs intent recognition and slot extraction on natural speech text. The following is an explanation of intent and slots.
[0209] ###Optional intents (intent values are only these. You must select one or more from the following options. Do not generate other options yourself)###
[0210] ['MUSIC / SONG', 'MUSIC / ALBUM', 'MUSIC / SINGER'...] Intent recognition has only the above options and cannot generate new options by itself
[0211] ###Intent Explanation###
[0212] MUSIC / SONG: Play a song. The entity only contains songs. ##Example##"[Play](ACT_PLA Y)[Blue and White Porcelain](MUSIC_TAG)""[Open](ACT_OPEN)[Blue and White Porcelain](MUSIC_TAG)""[I want to listen](ACT_WANT)[Blue and White Porcelain](MUSIC_TAG)"
[0213] ###Multiple intent combinations (when there is more than one intent value, there is only one combination)###
[0214] [['MUSIC / SONG','MUSIC / ALBUM'],[...]] Multiple intents can only be combined in the above ways. It is forbidden to generate new combinations yourself. ### Special format query intent intervention ####
[0215] The query text in the following format belongs to the specified intent type:
[0216] (ACT_PLAY) TV series: xxxx
[0217]
[0218]
[0219] Based on the above description, perform intent recognition and slot extraction. Do not introduce additional information and strictly follow the information in the question. For intents and slots with limited scope, strictly adhere to it. Output the results in JSON format.
[0220] question:
[0221] [xxxx]
[0222] answer:
[0223] As can be seen, the prompt word template not only provides the goal of this intent recognition (such as "Perform intent recognition and slot extraction. Do not introduce additional information and strictly follow the information in the question. For intents and slots with limited ranges, you must strictly comply. Output the results in JSON format"), but also provides examples of recognition targets (such as "You are an assistant who performs intent recognition and slot extraction on natural speech text. The following is an explanation of intent and slots... answer: "). This helps the generative model learn how to perform target intent recognition based on examples, thereby improving the model recognition performance of the fine-tuned generative model.
[0224] (2) Model training.
[0225] The general process of model training can be found in Figure 18 ;like Figure 18 As shown, first, unlabeled data can be directly obtained online or manually injected into the unlabeled training data set. Then, specific unlabeled data (such as recent data or data that is more relevant to business data) is queried from the unlabeled training data. Data cleaning is performed on the unlabeled data to label the unlabeled training data and to land the labeled training data in the labeled training data set. Then, data is extracted from the labeled training data set, and the extracted training data is converted into a data format according to the requirements of the training platform (such as packaging the training data into a training data set file in json format), and the converted training data is packaged and uploaded to the training data warehouse. Finally, the training data is called from the training data warehouse for model training. The trained generative model can be evaluated to evaluate the model performance. If the performance is good, the generative model can be published online for use. Among them, the above-mentioned data cleaning, data extraction and model evaluation can be performed in the voice data operation platform, which is a platform for R&D personnel to perform visual operations and data queries; the entire model training process is implemented on a one-stop platform for generative models, and the efficiency of model training is improved through this one-stop platform.
[0226] Furthermore, the generative model trained based on the above model can perform intent recognition processing on the speech signal based on the speech signal and basic analysis data to obtain the intent recognition result.
[0227] (1) In the media resource scenario, if the intent recognition result indicates that the speech task corresponding to the speech signal is a conversational task, specifically a media resource search task under the conversational task, the user query is input into the film and television model service. The film and television model service is connected to a sub-intent classification model, which can classify the intent of the user query in more detail, so as to realize high-quality media resource search and generation for the user query based on the more detailed intent classification through the external network (such as other resource databases or platforms independent of the media resource application) and the intranet (i.e., the resource database corresponding to the media resource application).
[0228] The specific implementation process of the film and television model service calling the sub-intent classification model, external network and internal network to search and generate media resources can be found in the attached Figure 19 ;like Figure 19 As shown:
[0229] ① Obtaining multi-round interactive dialogue information: First, perform multi-round extraction on the user query; specifically, obtain multi-round interactive dialogue information from the interactive dialogue flow information reported by the media resource application.
[0230] ② Intranet search: perform intent classification and judgment on multiple rounds of user queries (i.e., at least one round of interactive dialogue information) to determine the search path intent of media resources (such as search path intent including "find recommendation", "find play history", "find media resource name", "find actor name", "find scene", "find list" and "find film and television information", etc.), that is, the dialogue intent of each round of interactive dialogue information in at least one round of interactive dialogue flow information. Then, according to the dialogue intent of each round of interactive dialogue information in at least one round of interactive dialogue flow information, perform intranet resource search processing to obtain intranet resources; specifically, call different interfaces (Application Programming Interface, API) according to different search path intentions, such as calling the corresponding interfaces of media resource information, list, recommendation, user history and search, and search for resources from the database corresponding to the media resource application. It is worth noting that in order to improve the richness of intranet resources, the embodiment of the present application supports the recall of richer intranet resources from the database corresponding to the media resource application from dimensions such as expanding the number of search results returned for intranet resources and generating keywords for content recall. This is more helpful in providing users with accurate media resources based on richer intranet resources.
[0231] The intent classification judgment of multiple rounds of user queries can be implemented by using a sub-intent classification model. The embodiment of the present application also supports fine-tuning the sub-intent classification model using prompt words. The fine-tuning prompt words can be as follows:
[0232] You are a video search AI assistant, and you are provided with user historical queries and latest queries. You need to refer to the user historical queries to classify the intent of the latest query.
[0233] #Intent Definition
[0234] The first-level intentions are: find a movie, find a function, and others
[0235] ##Film and TV
[0236] Definition: Contains film and television search intentions
[0237] ###Second-level intent list under "Movies and TV"
[0238] 1. Find the list: Find new hot related needs
[0239] 2. Find playback history: Find the user's historical viewing behavior
[0240] 3. Find recommendations: recommend personalized movies and TV shows
[0241] 4. Find scenes: movies and TV shows to watch in certain given scenes
[0242] 5. Find similarities: limit the content of similar movies and TV shows
[0243] 6. Find film and television information: limit detailed information on certain aspects of film and television
[0244] 7. Find actor name: Film and television works with the specified name
[0245] 8.Finding movies: Other film and television needs besides the above intentions
[0246] ##other
[0247] Definition: Intent other than "film and television"
[0248] ###Second-level intent list under "Other"
[0249] 1. Chat: Chat related content
[0250] #Query information
[0251] ##User History Query
[0252] {}
[0253] ##Latest Query
[0254] {}
[0255] Note: The output results must be in the "first-level intent, second-level intent" format.
[0256] Multi-level intent
[0257] It can be seen that the prompt word not only gives the goal of this intent recognition (such as "the result must be output in the format of "first-level intent, second-level intent"), but also gives examples of recognition targets (such as "You are a video search AI assistant, providing you with user historical queries and the latest queries. You need to refer to user historical queries to classify the latest queries for intent. ...##Latest Query{}"). This helps the sub-intent classification model learn how to implement intent classification based on examples, thereby improving the model recognition performance of the fine-tuned sub-intent classification model. It should be noted that the sub-intent classification model mentioned in the embodiment of the present application can actually be a generative model, which is not limited to this.
[0258] ③ External Network Search: Based on the conversational intent of each interactive conversation in at least one interactive conversation flow, an external network resource search is performed to obtain external network resources. Specifically, the media resource application uploads the interactive conversation flow to the corresponding backend server. The backend server extracts multiple rounds of interactive conversation information and performs query expansion on the question description in each interactive conversation to obtain more easily understood expanded information. High-quality web content is then searched for through other search paths independent of the media resource application's database (e.g., web search). The retrieved web content is then re-ranked (e.g., by slicing, filtering, and BM25 ranking (a ranking function based on a probabilistic model)) to extract external web content. Finally, the retrieved web content is summarized to extract film and television information summaries, ultimately obtaining external network resources. Optionally, identifier extraction (e.g., extracting the title of a drama (e.g., a video)) is also supported. Based on the extracted identifiers, resources are searched from the database corresponding to the media resource application (or media resource library). It is worth noting that in order to improve the search accuracy of external network resources, the embodiment of the present application supports improving the relevance of the searched external network resources and intentions from dimensions such as relevance and search path selection, thereby improving the accuracy of generating the final question result information.
[0259] ④ Reorganize the intranet resources and the external network resources to generate question result information corresponding to the voice signal. Among them, resource reorganization may include but is not limited to operations such as resource reordering and question result information generation; specifically, it includes reordering the intranet resources searched in the intranet according to the reordering strategy, and then calling a large language model with content generation capabilities to generate question result information based on the intranet resources and external network resources after reordering. For example, in the process of generating picture and text type question result information, CID (Click ID) will be generated to enable the resource image in the question result information to be identified, so as to generate content information related to the resource image, such as recommendation reason information. It is worth noting that the reordering strategy includes a variety of sub-strategies for generating question result information, such as the operation period weighting sub-strategy, the site weighting sub-strategy, the classic content weighting sub-strategy and the popularity weighting sub-strategy; in this way, more accurate question result information can be generated through a rich reordering strategy.
[0260] (2) In the media resource scenario, if the intent recognition result indicates that the speech task corresponding to the speech signal is a dialogue task, specifically a non-media resource search task under the dialogue task, the user query is input into the non-film and television model service. The non-film and television model service is mainly used to process the content of the general vertical intent field and the playback, operation and other related content of video media resources in the media resource application. In particular, for the intent in the operation field, there will be more configurations of application control and TV station operation; specifically, the operation data such as application control and TV station operation will be configured and stored first, and then the slot results (slot results) of the user query analysis mentioned above will be combined with the three generative models mentioned above to perform intent matching from the stored operation data, thereby determining the specific task content of the direct task.
[0261] S1305: Displaying an interactive dialogue interface.
[0262] It should be noted that the specific implementation process shown in step S1305 and Figure 3 The specific implementation process of "displaying the interactive dialogue interface" described in step S301 in the illustrated embodiment is the same and will not be repeated here.
[0263] S1306: Displaying interactive dialogue flow information triggered by the voice signal in the interactive dialogue area.
[0264] S1307: Displaying content information related to the interactive dialogue flow information in the content display area.
[0265] It should be noted that the specific implementation process shown in steps S1306-S1307 and Figure 3The specific implementation process shown in steps S302-S303 in the illustrated embodiment is the same. Please refer to the relevant description of the specific implementation process shown in steps S302-S303, and no further details will be given here.
[0266] It should also be noted that embodiments of the present application also support fully displaying the process of result search processing in the interactive dialogue area, presenting users with a visual effect similar to that of humans thinking about a problem. This can, to a certain extent, alleviate the anxiety of users waiting for the problem result information and also facilitate users' perception that the generation of the problem result information is based on evidence. Furthermore, embodiments of the present application also support displaying the thinking process of the problem result information in the interactive dialogue area and displaying the problem result information at any time, and pausing the information display based on the user's pause request, thereby improving the intelligence and controllability of the information display in the interactive dialogue area and enhancing the user's interactive experience.
[0267] The following is combined with Figure 20 The specific process of displaying the thinking information corresponding to each round of interactive dialogue information and the pause information display in the interactive dialogue area is introduced. Figure 20 As shown in the accompanying figure (a), the question result information is output in the interactive dialogue area using a streaming response method (see the above-mentioned introduction to the streaming response); when the resource playback device (specifically, the server corresponding to the resource playback device) receives a pause request from the user before starting to generate the question result information, the generation of the question result information is paused, and a first notification message 2001 is displayed in the interactive dialogue area. The first notification message 2001 is used to indicate that the generation of the question result information has been paused. Figure 20 As shown in FIG. (b), when the resource playback device is generating question result information, it will first output the generated thinking information 2002; if a user's pause request is received during the generation of the question result information, the output part of the question result information will be retained in the interactive dialogue area, and the generation of the question result information will be paused, and a second notification message 2003 will be displayed in the interactive dialogue area, and the second notification message 2003 is used to indicate that the generation of the question result information has been paused. Figure 20 As shown in Figure (c) in the figure, if the resource playback device receives a pause request from the user after the question result information has been generated, the generated question result information is still displayed in the interactive dialogue area, and a third notification message 2004 is output. The third notification message 2004 is used to prompt that the question result information has been fully displayed.
[0268] It should be noted that the content information in the content display area is also output in a streaming response manner, such as Figure 21The displayed content information is rendered and displayed gradually in the content display area. Optionally, if the content information in the content display area needs to wait until the question result information in the interactive dialogue area is rendered and displayed before it starts to be rendered and displayed, then no matter when a pause request is generated in the interactive dialogue area, the content information in the content display area will be paused; this process of rendering and displaying the interactive dialogue area first and then the content display area can be seen in Figure 20 Optionally, if the content information in the content display area can be rendered and displayed synchronously with the question result information in the interactive dialogue area, then when there is a pause request in the interactive dialogue area, the content information in the content display area is also paused.
[0269] To summarize, on the one hand, traditional OTT devices use remote control searches to search for resources and control, etc.; and the use of remote controls is limited to user groups (such as the elderly and children who cannot use remote controls), resulting in low operating efficiency and long access paths to commonly used paths. However, the embodiments of the present application support media resource searches and device controls (such as volume control) based on voice interaction, thereby achieving a state where all people and multiple tasks (such as resource search tasks and device control tasks) can be completed. On the other hand, before this invention, users, but traditional solutions have poor intent recognition capabilities and poor content recall capabilities for long queries and complex queries. This embodiment of the present application provides a speech processing method based on a generative large language model, which has better intent recognition capabilities and stronger content recall capabilities for long queries and complex queries than can only use traditional word segmentation technology and intent recognition capabilities to perform word segmentation and intent recognition, and then recall content.
[0270] The method of the embodiment of the present application is described in detail above. In order to facilitate the above-mentioned scheme of the embodiment of the present application to be better implemented, accordingly, the device of the embodiment of the present application is provided below. In the embodiment of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuit or memory) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the module or unit function.
[0271] Figure 22 A schematic diagram of the structure of a speech processing device provided by an exemplary embodiment of the present application is shown; the speech processing device can be used to perform Figure 3 and Figure 13 Some or all of the steps in the method embodiment shown. Figure 22 , the device includes the following units:
[0272] The receiving unit 2201 is used to receive a voice signal and display an interactive dialogue interface; the interactive dialogue interface includes an interactive dialogue area and a content display area;
[0273] The processing unit 2202 is configured to display interactive dialogue flow information triggered by the voice signal in the interactive dialogue area; and
[0274] The processing unit 2202 is further configured to display content information related to the interactive dialogue flow information in the content display area.
[0275] In one implementation, the interactive dialogue flow information includes N rounds of interactive dialogue information, each round of interactive dialogue information includes question description information and question result information; N is a positive integer;
[0276] The content information related to the interactive dialog flow information includes at least one of the following:
[0277] One or more candidate question description information that matches the dialogue intent of the i-th round of interactive dialogue information; the i-th round of interactive dialogue information is displayed in the interactive dialogue area, where i is an integer and 1≤i≤N; and
[0278] Target content information related to the target dialogue information in the i-th round of interactive dialogue information; wherein the target content information includes at least one of the following: detailed information of the target dialogue information, source information of the target dialogue information, and commentary information of the target dialogue information; the target dialogue information is default or customized.
[0279] In one implementation, the interactive dialogue flow information includes N rounds of interactive dialogue information; the i-th round of interactive dialogue information includes first question description information and first question result information, and the first question result information includes a resource image corresponding to the video media resource; the resource image is defaulted to the target dialogue information in the i-th round of interactive dialogue information;
[0280] The processing unit 2202 is configured to display content information related to the interactive dialogue flow information in the content display area, specifically to:
[0281] In the content display area, target content information related to the resource image is displayed by default.
[0282] In one implementation, the target content information associated with the resource image is a video media resource; the processing unit 2202 is further configured to:
[0283] Play video media resources in the content display area;
[0284] During playback of the video media resource, receiving a pause operation for the video media resource;
[0285] In response to the pause operation on the video media resource, the video media resource is paused in the content presentation area.
[0286] In one implementation, the interactive dialogue flow information includes N rounds of interactive dialogue information; the processing unit 2202 is configured to display content information related to the interactive dialogue flow information in the content display area, specifically to:
[0287] In response to an information viewing operation performed in the interactive dialogue area on target dialogue information in the i-th round of interactive dialogue information, displaying the target dialogue information in the interactive dialogue area as a selected state;
[0288] Target content information related to the target conversation information is displayed in the content display area.
[0289] In one implementation, the number of content information related to the interactive dialogue flow information is at least one; the content display area includes at least one content type option, each content type option corresponding to a type of content information; the processing unit 2202 is further configured to:
[0290] In response to a selection operation on a target content type option, content information corresponding to the selected target content type option is displayed in the content display area; the target content type option is any one of the at least one content type option.
[0291] In one implementation, the processing unit 2202 is further configured to:
[0292] During the process of performing a dialogue viewing operation on the interactive dialogue information of the i-th round in the interactive dialogue area, maintaining the display of content information related to the interactive dialogue information of the i-th round in the content display area; the dialogue viewing operation includes at least one of the following: a page turning operation, a scrolling operation, and a dragging operation;
[0293] When a dialogue viewing operation is detected in the interactive dialogue area to switch from the i-th round of interactive dialogue information to the j-th round of interactive dialogue information, the content information related to the i-th round of interactive dialogue information will be updated and displayed in the content display area as the content information related to the j-th round of interactive dialogue information; j is a positive integer, j≠i, and 1≤j≤N.
[0294] In one implementation, the content display area includes dialogue options corresponding to each round of interactive dialogue information in the interactive dialogue flow information; the processing unit 2202 is further configured to:
[0295] In response to a triggering operation on a target dialogue option in the content display area, content information related to the target interactive dialogue information corresponding to the target dialogue option is displayed in the content display area; the target dialogue option is a dialogue option corresponding to any round of interactive dialogue information;
[0296] The target interactive dialogue information is displayed in the interactive dialogue area.
[0297] In one implementation, the number of target dialogue information in the i-th round of interactive dialogue information is at least two; the processing unit 2202 is further configured to:
[0298] When the content positioning operation is performed on the target dialogue information designated in the interactive dialogue area, target content information related to the designated target dialogue information is displayed in the content display area.
[0299] In one implementation, the voice processing method is applied to a media resource application, which runs on a resource playback device; the receiving unit 2201, when receiving a voice signal, is specifically configured to:
[0300] Displaying the service interface of the resource playback device, the service interface includes voice prompt information in the prompt state; the voice prompt information in the prompt state is used to prompt that the resource playback device has a voice interaction function;
[0301] In response to a confirmation operation on the voice prompt information in the prompt state, the voice prompt information is converted from the prompt state to the wake-up state; the voice prompt information in the wake-up state is used to prompt the resource playback device to be in the voice collection state;
[0302] When the resource playback device starts to collect voice signals, the voice prompt information is converted from the awakening state to the collection state; the voice prompt information in the collection state is used to prompt the resource playback device that the voice signal is being collected;
[0303] When the resource playback device finishes collecting the voice signal, the voice prompt information is converted from the collection state to the understanding state; the voice prompt information in the understanding state is used to prompt the resource playback device to perform result search processing based on the voice signal.
[0304] In one implementation, the resource playback device is a device that transmits content based on the Internet; wherein,
[0305] The interactive dialogue area and the content display area may be displayed in the interactive dialogue interface in a manner including: horizontal arrangement display, vertical arrangement display, and mosaic display.
[0306] In one implementation, the result search process includes intent recognition processing and result generation processing. The processing unit 2202 is further configured to:
[0307] Obtaining basic analysis data, where the basic analysis data includes one or more of scenario data, object data, permission data, and intent list data;
[0308] Call the fine-tuned generative model to perform intent recognition processing on the voice signal based on the basic analysis data to obtain the intent recognition result;
[0309] If the intention recognition result indicates that the speech task indicated by the speech signal is a dialogue task, result generation processing is performed based on the intention recognition result to generate question result information corresponding to the speech signal; the question result information and the question description information corresponding to the speech signal constitute a round of interactive dialogue information in the interactive dialogue flow information.
[0310] In one implementation, the processing unit 2202 is further configured to:
[0311] If the intention recognition result indicates that the speech task indicated by the speech signal is a direct task, the direct task is processed;
[0312] Among them, direct tasks include at least one of the following: control tasks, interface switching tasks and player control tasks.
[0313] In one implementation, the processing unit 2202 is configured to perform result generation processing based on the intention recognition result, and to generate question result information corresponding to the speech signal, specifically to:
[0314] Obtaining at least one round of interactive dialogue flow information in the interactive dialogue flow information;
[0315] Performing intranet resource search processing based on the conversation intention of at least one round of interactive conversation flow information to obtain intranet resources; and
[0316] Performing external network resource search processing based on the conversation intention of at least one round of interactive conversation flow information to obtain external network resources;
[0317] The intranet resources and the extranet resources are reorganized to generate problem result information corresponding to the voice signal.
[0318] According to one embodiment of the present application, Figure 22The various units in the voice processing device shown can be individually or completely combined into one or several other units to form a whole, or one (or some) of the units can be further divided into multiple functionally smaller units to form a whole, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the voice processing device can also include other units. In actual applications, these functions can also be implemented with the assistance of other units and can be implemented by the collaboration of multiple units. According to another embodiment of the present application, the following can be executed by running on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), and a read-only memory (ROM). Figure 3 and Figure 13 The computer program of each step involved in the corresponding method shown is constructed as follows Figure 22 The processing device based on media application shown in and the speech processing method of the embodiment of the present application are implemented. The computer program can be recorded on a computer readable recording medium, for example, and loaded into the above-mentioned computing device through the computer readable recording medium and run therein.
[0319] Based on the same inventive concept, the principles and beneficial effects of solving the problem by the computer device provided in the embodiment of the present application are similar to the principles and beneficial effects of solving the problem by the speech processing method in the method embodiment of the present application. Please refer to the principles and beneficial effects of the implementation of the method. For the sake of concise description, they will not be repeated here.
[0320] Figure 23 FIG2 shows a schematic diagram of a computer device provided by an exemplary embodiment of the present application. Figure 23, the computer device includes a processor 2301, a communication interface 2302 and a computer-readable storage medium 2303. The processor 2301, the communication interface 2302 and the computer-readable storage medium 2303 can be connected via a bus or other means. The communication interface 2302 is used to receive and send data. The computer-readable storage medium 2303 can be stored in the memory of the computer device, the computer-readable storage medium 2303 is used to store computer programs, and the processor 2301 is used to execute the computer programs stored in the computer-readable storage medium 2303. The processor 2301 (or CPU) is the computing core and control core of the computer device, which is suitable for implementing one or more computer programs, and is specifically suitable for loading and executing one or more computer programs to implement corresponding method processes or corresponding functions.
[0321] The embodiment of the present application also provides a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space that stores the processing system of the computer device. In addition, one or more computer programs suitable for being loaded and executed by the processor 2301 are also stored in the storage space. These computer programs can be one or more computer programs. It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one computer-readable storage medium located away from the aforementioned processor.
[0322] In one embodiment, the computer device may be the desktop device, mobile device or embedded device mentioned in the aforementioned embodiment; one or more computer programs are stored in the computer-readable storage medium; the processor 2301 loads and executes the one or more computer programs stored in the computer-readable storage medium to implement the corresponding steps in the above-mentioned speech processing method embodiment; in a specific implementation, the one or more computer programs in the computer-readable storage medium are loaded and executed by the processor 2301 to execute the steps of each embodiment of the present application; wherein, the steps of each embodiment of the present application can be referred to the relevant description of the aforementioned embodiments, and will not be repeated here.
[0323] Based on the same inventive concept, the principles and beneficial effects of solving the problem by the computer device provided in the embodiment of the present application are similar to the principles and beneficial effects of solving the problem by the speech processing method in the method embodiment of the present application. Please refer to the principles and beneficial effects of the implementation of the method. For the sake of concise description, they will not be repeated here.
[0324] An embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the above-mentioned speech processing method is implemented.
[0325] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0326] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs (one or more). When the computer program is loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The computer program can be stored in a computer-readable storage medium or transmitted by a computer-readable storage medium. The computer program can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. Available media may be magnetic media (eg, floppy disks, hard disks, magnetic tapes), optical media (eg, digital video disks (DVDs)), or semiconductor media (eg, solid state disks (SSDs)).
[0327] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A speech processing method, characterized in that: include: Receive voice signals and display an interactive dialogue interface; The interactive dialogue interface includes an interactive dialogue area and a content display area; Displaying interactive dialogue flow information triggered by the voice signal in the interactive dialogue area; as well as, Content information related to the interactive dialogue flow information is displayed in the content display area.
2. The method according to claim 1, wherein The interactive dialogue flow information includes N rounds of interactive dialogue information, each round of interactive dialogue information includes question description information and question result information; N is a positive integer; The content information related to the interactive dialogue flow information includes at least one of the following: One or more candidate question description information that matches the dialogue intent of the interactive dialogue information in round i; The interactive dialogue information of the i-th round is displayed in the interactive dialogue area, where i is an integer and 1≤i≤N; as well as, Target content information related to the target dialogue information in the i-th round of interactive dialogue information; wherein the target content information includes at least one of the following: detailed information of the target dialogue information, source information of the target dialogue information, and commentary information of the target dialogue information; the target dialogue information is default or customized.
3. The method according to claim 1 or 2, wherein: The interactive dialogue flow information includes N rounds of interactive dialogue information; the i-th round of interactive dialogue information includes first question description information and first question result information, wherein the first question result information includes a resource image corresponding to a video media resource; the resource image is defaulted to the target dialogue information in the i-th round of interactive dialogue information; Displaying content information related to the interactive dialogue flow information in the content display area includes: Target content information related to the resource image is displayed in the content display area by default.
4. The method according to claim 3, wherein The target content information related to the resource image is a video media resource; the method further includes: Play the video media resource in the content display area; During the playback of the video media resource, receiving a pause playback operation for the video media resource; In response to a pause operation on the video media resource, the video media resource is paused in the content display area.
5. The method according to claim 1 or 2, wherein: The interactive dialogue flow information includes N rounds of interactive dialogue information; Displaying content information related to the interactive dialogue flow information in the content display area includes: In response to an information viewing operation performed in the interactive dialogue area on target dialogue information in the i-th round of interactive dialogue information, displaying the target dialogue information in the interactive dialogue area as a selected state; Target content information related to the target conversation information is displayed in the content display area.
6. The method according to claim 1 or 2, wherein: The number of content information related to the interactive dialogue flow information is at least one; the content display area includes at least one content type option, and each content type option corresponds to a type of content information; The method further comprises: In response to a selection operation on a target content type option, displaying the content information corresponding to the selected target content type option in the content display area; The target content type option is any one of the at least one content type option.
7. The method according to claim 1 or 2, wherein: The method further comprises: During a process of performing a dialogue viewing operation on the interactive dialogue information of the i-th round in the interactive dialogue area, maintaining the display of content information related to the interactive dialogue information of the i-th round in the content display area; the dialogue viewing operation includes at least one of the following: a page turning operation, a scrolling operation, and a dragging operation; When a dialogue viewing operation is detected in the interactive dialogue area to switch from the i-th round of interactive dialogue information to the j-th round of interactive dialogue information, the content information related to the i-th round of interactive dialogue information is updated and displayed in the content display area as content information related to the j-th round of interactive dialogue information; j is a positive integer, j≠i, and 1≤j≤N.
8. The method according to claim 1 or 2, wherein: The content display area includes dialogue options corresponding to each round of interactive dialogue information in the interactive dialogue flow information; the method further includes: In response to a triggering operation on a target dialogue option in the content display area, content information related to target interactive dialogue information corresponding to the target dialogue option is displayed in the content display area; the target dialogue option is a dialogue option corresponding to any round of interactive dialogue information; The target interactive dialogue information is displayed in the interactive dialogue area.
9. The method according to claim 8, wherein The number of target dialogue information in the i-th round of interactive dialogue information is at least two; the method further includes: When a content positioning operation is performed on the target dialogue information designated in the interactive dialogue area, target content information related to the designated target dialogue information is displayed in the content display area.
10. The method according to claim 1 or 2, wherein: The method is applied to a media resource application, which runs in a resource playback device; The receiving of the voice signal comprises: Displaying a service interface of the resource playback device, wherein the service interface includes voice prompt information in a prompt state; the voice prompt information in the prompt state is used to prompt that the resource playback device has a voice interaction function; In response to a confirmation operation on the voice prompt information in the prompt state, the voice prompt information is converted from the prompt state to the awake state; the voice prompt information in the awake state is used to prompt that the resource playback device is in the voice collection state; When the resource playback device starts to collect the voice signal, the voice prompt information is converted from the awakening state to the collection state; the voice prompt information in the collection state is used to prompt the resource playback device that the voice signal is being collected; When the resource playback device finishes collecting the voice signal, the voice prompt information is converted from the collection state to the understanding state; the voice prompt information in the understanding state is used to prompt the resource playback device to perform result search processing based on the voice signal.
11. The method according to claim 10, wherein The resource playback device is a device that transmits content based on the Internet content transmission method; wherein, The interactive dialogue area and the content display area are displayed in the interactive dialogue interface in the following ways: horizontal arrangement display, vertical arrangement display and mosaic display.
12. The method according to claim 10, wherein The result search process includes intention recognition processing and result generation processing, and the method further includes: Acquire basic analysis data, where the basic analysis data includes one or more of scenario data, object data, permission data, and intent list data; Calling the fine-tuned generative model to perform intent recognition processing on the voice signal based on the basic analysis data to obtain an intent recognition result; If the intention recognition result indicates that the voice task indicated by the voice signal is a dialogue-type task, result generation processing is performed based on the intention recognition result to generate question result information corresponding to the voice signal; the question result information and the question description information corresponding to the voice signal constitute a round of interactive dialogue information in the interactive dialogue flow information.
13. The method according to claim 12, wherein: The method further comprises: If the intention recognition result indicates that the voice task indicated by the voice signal is a direct task, performing task processing on the direct task; The direct access tasks include at least one of the following: a control task, an interface switching task, and a player control task.
14. The method according to claim 12, wherein The performing result generation processing according to the intention recognition result to generate question result information corresponding to the voice signal includes: Acquire at least one round of interactive dialogue flow information from the interactive dialogue flow information; Performing intranet resource search processing based on the conversation intention of at least one round of interactive conversation flow information to obtain intranet resources; and Performing external network resource search processing based on the conversation intention of at least one round of interactive conversation flow information to obtain external network resources; The intranet resources and the extranet resources are reorganized to generate question result information corresponding to the voice signal.
15. A speech processing device, characterized in that: include: A receiving unit, configured to receive voice signals and display an interactive dialogue interface; The interactive dialogue interface includes an interactive dialogue area and a content display area; A processing unit, configured to display interactive dialogue flow information triggered by the voice signal in the interactive dialogue area; as well as, The processing unit is further configured to display content information related to the interactive dialogue flow information in the content display area.
16. A computer device, characterized in that: a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein when the computer program is executed by the processor, the speech processing method according to any one of claims 1 to 14 is implemented.
17. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the speech processing method according to any one of claims 1 to 14.
18. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the speech processing method according to any one of claims 1 to 14 is implemented.