Intelligent interaction method and device, electronic equipment, computer readable storage medium and computer program product
By performing intention recognition and incremental reasoning on the voice data of the video playback client, intuitive recall results are generated, and the problems of single interaction mode and poor scalability in the existing technology are solved, and the flexibility and diversity of intelligent interaction are achieved.
Patent Information
- Application Number
- CN202510549901.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-25
AI Technical Summary
The existing video platform's voice recognition technology has fixed interaction methods, resulting in a single interactive content and poor scalability, making it impossible to accurately understand the user's complex context and vague expressions, and is prone to misjudgment in intentions when facing new needs.
By identifying the voice data of the video playback client, the target text is generated and intent recognition is performed, the recall results are generated using incremental inference streaming, including the reply text and the playback control image of the target video, and sent to the client through the streaming graphics and text protocol.
It improves the flexibility and diversity of intelligent interaction, and the recall results are intuitive and complete, which is easy for users to understand and operate, simplifies user processes, and improves interaction efficiency and convenience.
Smart Images

Figure CN120378650A_ABST
Abstract
Description
Technical Field
[0001] This application relates to computer technology, and in particular, to an intelligent interaction method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] In today's digital age, the creation and dissemination of video content have grown explosively, and video has become one of the main ways for people to obtain information and entertainment. At the same time, users' demands for the interactivity and convenience of video platforms are increasing day by day, which provides broad space for the development of speech recognition technology in video platforms. In current video platforms, speech recognition technology mainly uses a mode of training based on large-scale corpora to achieve intent recognition, binding each corpus sample to a specific intent category and operation behavior. When a user issues a voice command (query) that hits a certain corpus sample, the corresponding intent recognition result and operation instruction will be sent down. However, this interaction method is fixed, with single interaction content and poor scalability. Moreover, when facing complex contexts, ambiguous expressions, or new types of requirements, problems such as incorrect intent judgment are likely to occur. Summary of the Invention
[0003] Embodiments of this application provide an intelligent interaction method, apparatus, computer-readable storage medium, and computer program product, which can improve the accuracy and diversity of intelligent interaction.
[0004] The technical solution of the embodiments of this application is implemented as follows:
[0005] Embodiments of this application provide an intelligent interaction method, the method comprising:
[0006] Recognize the voice data sent by the video playback client to obtain a target text, and perform intent recognition on the target text to obtain an intent recognition result;
[0007] When it is determined based on the intent recognition result that the intent type is a video recommendation type, perform incremental reasoning based on the intent recognition result and the target text, and streamingly generate a recall result, where the recall result at least includes a response text and a playback control image of at least one target video;
[0008] Send the recall result to the video playback client based on the streaming graphic protocol.
[0009] Embodiments of this application provide an intelligent interaction method, the method comprising:
[0010] In response to the received voice collection instruction, collect voice data and send the voice data to the server;
[0011] Receive the target text recognized by the server for the voice data, and display the target text;
[0012] When the intent type corresponding to the target text is a video recommendation type, receive the recall result sent by the server, where the recall result is incrementally inferred and stream-generated by the server based on the intent recognition result and the target text, and the recall result at least includes a reply text and a playback control image of at least one target video;
[0013] Parse and display the recall result.
[0014] An embodiment of the present application provides an intelligent interaction device, and the device includes:
[0015] An identification module, configured to identify voice data sent by a video playback client to obtain a target text, and perform intent recognition on the target text to obtain an intent recognition result;
[0016] A determination module, configured to, when it is determined based on the intent recognition result that the intent type is a video recommendation type, perform incremental inference based on the intent recognition result and the target text to stream-generate a recall result, where the recall result at least includes a reply text and a playback control image of at least one target video;
[0017] A sending module, configured to send the recall result to the video playback client based on a streaming graphic protocol.
[0018] An embodiment of the present application provides an intelligent interaction device, and the device includes:
[0019] An acquisition module, configured to acquire voice data in response to a received voice acquisition instruction, and send the voice data to the server;
[0020] A receiving module, configured to receive the target text recognized by the server for the voice data, and display the target text;
[0021] The receiving module is further configured to, when the intent type corresponding to the target text is a video recommendation type, receive the recall result sent by the server, where the recall result is incrementally inferred and stream-generated by the server based on the intent recognition result and the target text, and the recall result at least includes a reply text and a playback control image of at least one target video;
[0022] A parsing and display module, configured to parse and display the recall result.
[0023] An embodiment of the present application provides an electronic device, and the electronic device includes:
[0024] A memory, configured to store computer-executable instructions or computer programs;
[0025] A processor, when executing computer-executable instructions or a computer program stored in the memory, implements the intelligent interaction method provided by the embodiments of the present application.
[0026] The embodiments of the present application provide a computer-readable storage medium storing a computer program or computer-executable instructions, which are used to implement the intelligent interaction method provided by the embodiments of the present application when being executed by a processor.
[0027] The embodiments of the present application provide a computer program product including a computer program or computer-executable instructions, which implement the intelligent interaction method provided by the embodiments of the present application when the computer program or computer-executable instructions are executed by a processor.
[0028] The embodiments of the present application have the following beneficial effects:
[0029] In the embodiments of the present application, the voice data sent by the video playback client is recognized to obtain a target text, and the target text is subjected to intent recognition to obtain an intent recognition result; when it is determined based on the intent recognition result that the intent type is video recommendation, incremental reasoning is performed based on the intent recognition result and the target text, and recall results are generated in a streaming manner; based on the streaming graphic protocol, the recall results are sent to the video playback client. In this way, when the voice data sent by the client belongs to video recommendation, incremental reasoning can be performed according to the intent recognition result and the target text, enriching the recall results for the voice data and improving the diversity of intelligent interaction content. Among them, the recall results at least include a reply text and a playback control image of at least one target video. Therefore, the recall results are more intuitive and complete, facilitating the user to understand and select the recall results by combining the text and the image. In addition, through the playback control image of the target video, the user can quickly view the target video of interest, thus simplifying the user's operation process and improving the convenience and flexibility of intelligent interaction. And because the recall results are generated in a streaming manner, the recall results can be presented to the user faster, thereby improving the intelligent interaction efficiency and making the interaction process smoother. Therefore, the flexibility and diversity of intelligent interaction can be improved through the embodiments of the present application. Description of the Drawings
[0030] Figure 1 is a schematic diagram of the architecture of the intelligent interaction system 100 provided by the embodiments of the present application;
[0031] Figure 2A is a schematic diagram of the structure of the server 200 provided by the embodiments of the present application;
[0032] Figure 2B is a schematic diagram of the structure of the terminal 400 provided by the embodiments of the present application;
[0033] Figure 3A It is a schematic flowchart of the intelligent interaction method provided by an embodiment of the present application;
[0034] Figure 3B It is a schematic flowchart of generating a retrieval result provided by an embodiment of the present application;
[0035] Figure 3C It is a schematic flowchart of obtaining multiple result shards provided by an embodiment of the present application;
[0036] Figure 4A It is another schematic flowchart of the intelligent interaction method provided by an embodiment of the present application;
[0037] Figure 4B It is a schematic flowchart of displaying a recall result provided by an embodiment of the present application;
[0038] Figure 4C It is another schematic flowchart of displaying a recall result provided by an embodiment of the present application;
[0039] Figure 5 It is yet another implementation flowchart of the intelligent interaction method provided by an embodiment of the present application;
[0040] Figure 6 It is a schematic diagram of the display state of a guidance state provided by an embodiment of the present application;
[0041] Figure 7 It is a schematic diagram of the display state of a wake-up state provided by an embodiment of the present application;
[0042] Figure 8 It is a schematic diagram of the display state of an identification state provided by an embodiment of the present application;
[0043] Figure 9 It is a schematic diagram of the display state of an understanding state provided by an embodiment of the present application;
[0044] Figure 10A It is a schematic diagram of the display state of an execution state provided by an embodiment of the present application;
[0045] Figure 10B It is a schematic diagram of an interface for presenting an interaction result in a multi-turn dialogue provided by an embodiment of the present application;
[0046] Figure 10C It is another schematic diagram of an interface for presenting an interaction result in a multi-turn dialogue provided by an embodiment of the present application;
[0047] Figure 11 It is yet another schematic flowchart of the intelligent interaction method provided by an embodiment of the present application;
[0048] Figure 12 It is a schematic flowchart of a streaming data transmission provided by an embodiment of the present application;
[0049] Figure 13 It is a schematic flowchart of a card generation method provided by an embodiment of the present application;
[0050] Figure 14 It is a schematic diagram of card display in a dialog box provided by an embodiment of the present application;
[0051] Figure 15 It is a schematic diagram of the display of a question-and-answer interface provided by an embodiment of the present application. Detailed implementation manners
[0052] In order to make the purpose, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0053] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0054] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0055] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0056] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0057] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the personal information subject.
[0058] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0059] 1) Automatic Speech Recognition (ASR): is a technology that converts human speech into text;
[0060] 2) Over The Top (OTT): video, audio and other media services delivered directly to end users via the Internet, bypassing traditional cable, satellite or terrestrial broadcast distribution channels;
[0061] 3) Streaming Graphics and Text Protocol: A protocol used to transmit graphic and text data in network communications. It allows graphic and text data to be transmitted not all at once, but divided into multiple small parts like a stream of water, and transmitted to the receiving end in a certain order. The receiving end can start processing and displaying the received part of the content during the data transmission process.
[0062] 4) Intent recognition: A technology in the field of artificial intelligence that involves understanding the user's intentions and automatically providing corresponding services or responses based on these intentions;
[0063] 5) Incremental reasoning: refers to the ability to gradually and dynamically process new data during model training or reasoning, allowing the model to update its parameters or decision boundaries by learning new data based on existing knowledge, thereby improving the accuracy and adaptability of the model without retraining the entire model.
[0064] In the related art, the background will collect a large amount of corpus from the online for training. These corpus will be fixedly corresponded to specific intents during the training process. When a user initiates a query, the system first performs word segmentation on the query, and then retrieves the intent and its related actions that match the word segmentation result in a large database. If the user's query hits a predetermined intent template, the system will send the intent recognition result and the corresponding recall result to the client. The client displays various recall results to the user based on the intent template defined in the recall result. Among them, the method in the related art has the following problems:
[0065] 1) The understanding of user queries is fixed, and the generated recall results are single and not rich enough;
[0066] 2) During the intelligent interaction process, the complete recall results are sent in one go, so the interaction process is not very intelligent. It does not perform in-depth thinking based on the user's intention and cannot accurately understand the user's intention.
[0067] 3) Since the display of the recall results uses fixed templates, and each intention corresponds to a fixed module, the scalability is very poor. If a new intention is added, it means that both the client and the background need to do development work, and the operation is not flexible enough.
[0068] The embodiments of the present application provide an intelligent interaction method, device, equipment, computer-readable storage medium, and computer program product, which can improve the flexibility and diversity of intelligent interaction. The following describes the exemplary applications of the electronic devices provided by the embodiments of the present application. The electronic devices provided by the embodiments of the present application for implementing the intelligent interaction method can be implemented as various types of terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart TVs, vehicle terminals, OTT devices, etc., or can also be implemented as servers.
[0069] See Figure 1 , Figure 1 is the schematic architecture diagram of the intelligent interaction system 100 provided by the embodiments of the present application. As Figure 1 shown, a video playback client is installed in the terminal 400, which is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two. Among them, in response to the received voice collection instruction, the terminal 400 collects voice data through the voice collection device carried by itself, or collects voice data through the voice collection device in the remote control device of the terminal 400, and sends the voice data to the server 200. The server 200 is used to identify the voice data sent by the terminal 400 to obtain the target text and send it to the terminal 400. The terminal 400 receives the target text recognized by the server for the voice data and displays the target text on the display interface 401. After obtaining the target text, the server 200 will also perform intention recognition on the target text to obtain the intention recognition result. When the intention type is determined to be the video recommendation type based on the intention recognition result, incremental reasoning is performed based on the intention recognition result and the target text, and the recall result is generated in a streaming manner. The recall result at least includes the reply text and the playback control images of at least one target video; based on the streaming graphic protocol, the recall result is sent to the terminal 400. The terminal 400 receives the recall result sent by the server, parses the recall result, and displays the recall result on the display interface 401.
[0070] Taking the server described above as an example of the electronic device for intelligent interaction, see Figure 2A , Figure 2AIt is a schematic structural diagram of the server 200 provided by an embodiment of the present application. Figure 2A The illustrated server 200 includes: at least one processor 210, a memory 230, and at least one network interface 220. Each component in the server 200 is coupled together through a bus system 240. It can be understood that the bus system 240 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2A all kinds of buses are labeled as the bus system 240.
[0071] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0072] The memory 230 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. Optionally, the memory 230 includes one or more storage devices that are physically remote from the processor 210.
[0073] The memory 230 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), and the volatile memory can be a random access memory (RAM, Random Access Memory). The memory 230 described in the embodiments of the present application is intended to include any suitable type of memory.
[0074] In some embodiments, the memory 230 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.
[0075] An operating system 231, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0076] A network communication module 232, for reaching other electronic devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.
[0077] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2A Shown in the memory 230 is the intelligent interaction device 233, which may be software in the form of a program and a plug-in, etc., including the following software modules: an identification module 2331, a determination module 2332, and a sending module 2333. These modules are logical, so they can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.
[0078] Taking the electronic device for intelligent interaction as the above-mentioned terminal as an example, see Figure 2B , Figure 2B which is a schematic structural diagram of the terminal 400 provided in the embodiments of the present application. Figure 2B The shown terminal 400 includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the terminal 400 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. The bus system 440 includes, in addition to a data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2B all kinds of buses are labeled as the bus system 440.
[0079] The processor 410 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0080] The user interface 430 includes one or more output devices 431 capable of presenting media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.
[0081] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disk drives, etc. The memory 450 optionally includes one or more storage devices that are physically located away from the processor 410.
[0082] The memory 450 includes volatile memory, non-volatile memory, or both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM), and the volatile memory can be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0083] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are exemplarily described below.
[0084] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0085] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wi-Fi (Wireless Fidelity), and Universal Serial Bus (USB), etc.;
[0086] The presentation module 453 is used to enable the presentation of information (such as a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (such as a display screen, a speaker, etc.).
[0087] The input processing module 454 is used to detect and translate one or more user inputs or interactions from one of one or more input devices 432.
[0088] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 2B Shown is the intelligent interaction device 455 stored in the memory 450, which can be software in the form of programs and plugins, etc., including the following software modules: the acquisition module 4551, the reception module 4552, and the parsing and display module 4553. These modules are logical, and thus can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.
[0089] In some other embodiments, the device provided by the embodiments of the present application may be implemented in a hardware manner. As an example, the device provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the intelligent interaction method provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.
[0090] The intelligent interaction method provided by the embodiments of the present application may be executed by an electronic device. Taking the electronic device as a server as an example, the intelligent interaction method provided by the embodiments of the present application will be described below.
[0091] See Figure 3A , Figure 3A which is a schematic flowchart of the intelligent interaction method provided by the embodiments of the present application, and will be described in conjunction with Figure 3A the steps shown.
[0092] In step 101, the voice data sent by the video playing client is recognized to obtain a target text, and the target text is subjected to intent recognition to obtain an intent recognition result.
[0093] Here, the video playback client is a software application installed on a terminal for playing various video contents. Taking an Internet TV service (OTT) device as an example, the video playback client is an important tool for users to interact with various video contents and related services, providing users with rich and diverse video contents, high-quality playback experiences, and personalized services, meeting the needs of different users for Internet TV viewing. Voice data is an information carrier recorded and stored in the form of sound, which is data obtained after digitizing human voice signals. Speech recognition is the process of converting voice data into text, also known as automatic speech recognition. When a user says "Find a comedy movie" to the device of the video playback client, the video playback client will collect this voice data and send it to the server side. Through the speech recognition system on the server side to process and analyze the voice data, it is finally converted into a target text like "Find a comedy movie". Then, by performing intent recognition on the target text to understand the true intention or purpose behind the target text, an intent recognition result is obtained. The intent recognition result is used to describe the intention expressed by the target text, such as "Search for a certain movie", "Close a certain application", etc.
[0094] In some embodiments, "Performing intent recognition on the target text to obtain an intent recognition result" in step 101 can be achieved through the following process, including:
[0095] Performing word segmentation on the target text to obtain multiple target word segments; based on the multiple target word segments, determining a target intent service, and based on the target intent service, performing intent recognition on the multiple target word segments to obtain an intent recognition result.
[0096] Here, word segmentation processing is the process of splitting a continuous text sequence into individual words or phrases according to certain rules and algorithms. Among them, the target text can be cleaned first, including operations such as removing punctuation marks, numbers, special characters, and unifying the case of the text, to simplify subsequent processing. Then, a word segmentation algorithm is used to cut the text into words or phrases to obtain multiple target word segments. For example, for the target text "I want to watch movie A", the target word segments obtained through word segmentation processing may be: "I", "want to watch", "movie A". The target intent service is a service aimed at understanding the true intention behind the target word segments, which can be a traditional intent service or a large model intent service. Through the target intent service, intent recognition can be performed on the multiple target word segments to obtain the corresponding intent recognition result.
[0097] In some embodiments, the target intent service can be determined based on the total number of word segments of the target word segmentation. Further, the total number of word segments can be compared with a preset number threshold. When the total number of word segments is greater than the preset number threshold, the first intent service is determined as the target intent service; when the total number of word segments is less than or equal to the number threshold, the second intent service is determined as the target intent service. The intent recognition ability of the first intent service is better than that of the second intent service. However, generally, the intent recognition efficiency of the first intent service is lower than that of the second intent recognition service. The number threshold is a preset positive integer and can be flexibly adjusted according to actual needs. Exemplarily, the number threshold can be 3. When the total number of word segments is greater than the number threshold, it indicates that the target text is a complex text. At this time, the first intent service with stronger intent recognition ability is determined as the target intent service, thereby ensuring the accuracy of intent recognition. When the number of word segments is less than or equal to the number threshold, it indicates that the target text is a simple text. At this time, the second intent service with weaker intent recognition ability but higher intent recognition efficiency is determined as the target intent service, which can improve the intent recognition efficiency on the basis of ensuring accurate intent recognition results.
[0098] In some embodiments, the first intent service can be a large model intent service, which can understand more complex, ambiguous, and diverse semantic expressions. Based on the large language model, it conducts unsupervised or supervised pre-training on massive text data, learns the general patterns, semantic representations, and context relationships of the language, can automatically extract high-level features of the text, and has strong semantic understanding, context awareness, and generalization reasoning capabilities. The second intent service refers to the traditional intent service, which can analyze some simple text contents and mostly relies on rule-based methods, parsing semantics by manually writing a large number of rules. The traditional intent service also uses some traditional machine learning algorithms, such as Naive Bayes, Support Vector Machine, etc. When performing intent recognition, it needs to extract text features first and then classify and recognize intents based on these features.
[0099] In an embodiment of the present application, determine the total number of word segments of the target word segmentation; determine the target intent service based on the total number of word segments. Specifically, when the total number of word segments is greater than a preset number threshold, determine the first intent service as the target intent service; when the total number of word segments is less than or equal to the number threshold, determine the second intent service as the target intent service. In this way, when the total number of word segments of the target word segmentation is greater than the preset number threshold, it indicates that the target text is relatively long. At this time, using the first intent service as the target intent service for intent recognition can improve the accuracy of intent recognition. When the total number of word segments of the target word segmentation is less than or equal to the preset number threshold, it indicates that the target text is relatively simple. At this time, using the second intent service as the target intent service for intent recognition can improve the efficiency of intent recognition. In this way, by judging the complexity of the target text based on the total number of word segments of the target word segmentation and selecting an appropriate intent service, it is possible to optimize resource allocation and improve the overall performance and efficiency of the system while ensuring the quality of intent recognition.
[0100] In some embodiments, the target intent service can also be determined by performing text component analysis on the target text, specifically including:
[0101] Perform text component analysis on the target text to obtain the text components of each target word segment; determine the target intent service based on the text components of each target word segment.
[0102] Here, tools such as NLTK, Stanford CoreNLP, and AllenNLP in the natural language processing toolkit and library can be used to perform lexical analysis, syntactic analysis, or semantic analysis on the target text to obtain the text components of each target word segment. The text components can include the subject, predicate, object, attributive, adverbial, complement, etc. Among them, the subject, predicate, and object are basic components, and the attributive, adverbial, and complement are supplementary components. The preset component can be one or more of the supplementary components. When the text components of each target word segment include the preset component, it indicates that the text structure of the target text is relatively complex. At this time, determine the first intent service as the target intent service; when the text components of each target word segment do not include the preset component, it indicates that the text structure of the target text is relatively simple. At this time, determine the second intent service as the target intent service.
[0103] In an embodiment of the present application, text component analysis is performed on the target text to obtain the text components of each target word segment; based on the text components of each target word segment, a target intent service is determined, where when the text components of each target word segment include a preset component, the first intent service is determined as the target intent service; when the text components of each target word segment do not include the preset component, the second intent service is determined as the target intent service. In this way, when the text components of the target word segment include the preset component, it indicates that the semantics of the target text is complex. At this time, using the first intent service as the target intent service for intent recognition can improve the accuracy of intent recognition. When the text components of the target word segment do not include the preset component, it indicates that the semantics of the target text is simple. At this time, using the second intent service as the target intent service for intent recognition can improve the efficiency of intent recognition. In this way, the complexity of the target text can be judged based on the text components of the target word segment, and a suitable intent service can be selected, optimizing resource allocation and improving the overall performance and benefits of the system while ensuring the quality of intent recognition.
[0104] In an embodiment of the present application, word segmentation processing is performed on the target text to obtain a plurality of target word segments; based on the plurality of target word segments, a target intent service is determined, and based on the target intent service, intent recognition is performed on the plurality of target word segments to obtain an intent recognition result. In this way, the complexity of the target text can be judged based on the target word segments, and a suitable target intent service can be selected, improving the accuracy and efficiency of intent recognition for the target word segments.
[0105] In step 102, when it is determined based on the intent recognition result that the intent type is a video recommendation type, incremental reasoning is performed based on the intent recognition result and the target text, and a recall result is generated in a streaming manner.
[0106] Here, the intent type refers to the category of the purpose or requirement corresponding to the intent recognition result, and the video recommendation type means that the intent recognition result indicates that the user wants the system to make video recommendations according to the requirement, without specifying a designated video to watch. For example, if the user says to the video playback client "Recommend me science fiction movies with wonderful special effects", then the intent type is the video recommendation type. Incremental reasoning means that the reasoning process is not completed in one go, but gradually deepens and improves as more information is obtained or with further interaction from the user. Streaming generation of recall results means that the system does not give a response only after the entire reasoning process is completely finished, but during the reasoning process, once there are some definite recall results, it starts to output information to the user in real time. For example, first, based on the information "science fiction movies", all science fiction movies are screened out from the movie database, and then all science fiction movies will be presented to the user as partial recall results. Then, according to the condition "with wonderful special effects", all the previously screened science fiction movies are further evaluated and screened, and the screened movies are presented to the user as another part of the recall results in a streaming manner. Among them, the recall results at least include the response text and the image of the playback control of at least one target video. The response text is used to reply to the target text corresponding to the voice data. Continuing with the above example, the target text is "Recommend me science fiction movies with wonderful special effects", and the response text can be "The following are some recommendations for science fiction movies with amazing visual effects and grand worldviews, covering different styles and eras, guaranteed to feast your eyes". The target video is the recommended video that meets the user's needs determined according to the intent recognition result, and the image of the playback control of the target video refers to the image related to the target video with the function of controlling video playback. For example, the playback control image is the poster of a certain movie, and the playback control is set in part of the poster. By clicking on the poster, the playback control can be triggered, and then directly jump to the video playback page to play the corresponding movie.
[0107] In some embodiments, referring to Figure 3B , "Performing incremental reasoning based on the intent recognition result and the target text, and streaming generating recall results" in step 102 can be implemented through steps 1021 to 1023, including:
[0108] In step 1021, based on the intent recognition result, the target text is decomposed to obtain multiple sub - problems with logical relationships.
[0109] Here, the intended meaning of the target text may be multi-layered. For example, for the target text "Recommend to me the time-travel TV series starring A released in the last year", the corresponding intention recognition results include multi-layer intentions: "released in the last year", "starring A", and "time-travel TV series". At this time, the target text can be decomposed according to each layer of intention to obtain multiple sub-questions with logical relationships among them. Each sub-question corresponds to one layer of intention. For example, the sub-questions can include: sub-question 1 "time-travel TV series", sub-question 2 "starring A", and sub-question 3 "released in the last year".
[0110] In step 1022, for each sub-question, retrieve the target knowledge information for the sub-question in the knowledge database, and based on the target knowledge information and the reasoning result of the previous sub-question, determine the reasoning result of the sub-question.
[0111] Here, the knowledge database is a database system specially used for storing, managing, and organizing knowledge. For example, for sub-question 1 "time-travel TV series", all time-travel TV series can be found in the knowledge database as the target knowledge information of sub-question 1 and used as the reasoning result of sub-question 1; then, for sub-question 2 "starring A", all the film and television works starring A can be found in the knowledge database as the target knowledge information of sub-question 2, and the intersection of the target knowledge information of sub-question 2 and the reasoning result corresponding to sub-question 1 is determined as the reasoning result of sub-question 2; alternatively, among the reasoning results of sub-question 1, that is, the target knowledge information of sub-question 1, directly screen out the TV series starring A as the reasoning result of sub-question 2; then, for sub-question 3 "released in the last year", all the film and television works released in the last year can be found in the knowledge database as the target knowledge information of sub-question 3, and the intersection of the target knowledge information of sub-question 3 and the reasoning result of sub-question 2 is determined as the reasoning result of sub-question 3; alternatively, among the reasoning results corresponding to sub-question 2, directly screen out the TV series released in the last year as the reasoning result of sub-question 3.
[0112] In some embodiments, the reasoning results of each sub-question can also be sent to the video playback client, and the video playback client displays the reasoning results so that the user can intuitively understand the thinking process of the large language model and can, after the user gets a recall result that does not meet their requirements, make targeted adjustments to the question based on the reasoning results, thereby improving the matching degree between the recall result and the user's needs.
[0113] In step 1023, recall results are generated in a streaming manner based on the reasoning results of each sub-question and the intention recognition results.
[0114] Here, according to the inference results of each sub-question obtained through incremental inference, a search strategy is formulated. Determine from which data sources to conduct the search, such as video platforms, knowledge bases, document libraries, etc., and how to set search keywords and filtering conditions. For example, if it is inferred that the user needs to search for science fiction movies with wonderful special effects, then the search strategy may be to search on the video platform for videos labeled with "wonderful special effects" and classified as "science fiction movies". Then, conduct a real-time search according to the search strategy and return the search results to the video playback client in a streaming manner.
[0115] In some embodiments, each sub-question will have its independent inference and analysis process. As the processing progresses, the inference results and intention recognition results of each sub-question are gradually used to generate the recall results for each part, thereby achieving the streaming generation of recall results.
[0116] In the embodiments of the present application, based on the intention recognition results, the target text is decomposed to obtain multiple sub-questions with logical relationships; for each sub-question, target knowledge information for the sub-question is retrieved from the knowledge database, and based on the target knowledge information and the inference result of the previous sub-question, the inference result of the sub-question is determined; recall results are generated in a streaming manner based on the inference results and intention recognition results of each sub-question. In this way, through the gradual and in-depth inference and analysis of the target text and the streaming generation of recall results, it is possible to better simulate the way of human conversation, make the interaction process more natural and efficient, and improve the intelligence of intelligent interaction.
[0117] In some embodiments, the recall results include multiple result shards. Refer to Figure 3C , step 1023 can be implemented through steps 10231 to 10233, including:
[0118] In step 10231, based on the inference results of each sub-question, a reply result for each sub-question is generated.
[0119] Here, the inference results of each sub-question may include data of multiple data types, such as text data and image data. Abstract the data of each data type into the form of the same type of card to obtain the reply data corresponding to this type, and then combine all the reply data together as the reply result of the sub-question. Therefore, the reply result includes reply data of at least one data type.
[0120] In step 10232, the reply data corresponding to each data type is segmented to obtain multiple result data packets.
[0121] Here, the response data corresponding to each data type will be split to obtain multiple result data packets. For example, after splitting the response data 1 (text type), result data packets 1, 2, and 3 can be obtained; after splitting the response data 2 (image type), result data packets 4 and 5 can be obtained.
[0122] In step 10233, based on a preset protocol, the intent type, each data type, and each result data packet in the intent recognition result are respectively subjected to conversion processing to obtain multiple result shards.
[0123] Here, the preset protocol is used to perform data conversion on the intent type, data type, and result data packet based on a preset data format to obtain data in the preset data format as the result shards. The data type includes at least text type and image type.
[0124] In the embodiment of the present application, the recall result includes multiple result shards. Based on the inference result of each sub-question, the response result of each sub-question is generated, and the response result includes response data of at least one data type; the response data corresponding to each data type is split to obtain multiple result data packets; based on a preset protocol, the intent type, each data type, and each result data packet in the intent recognition result are respectively subjected to conversion processing to obtain multiple result shards. In this way, by splitting the data, the data transmission efficiency can be improved, and by performing conversion processing on the data through the preset protocol, the data transmission efficiency and processing efficiency can be further improved.
[0125] In step 103, based on the streaming graphic protocol, the recall result is sent to the video playback client.
[0126] Here, the streaming graphic protocol is a communication protocol for transmitting graphic information in a streaming manner over the network. It allows the graphic data to be transmitted in the form of a continuous stream rather than sending the entire file at once. Through the streaming graphic protocol, the recall result can be sent to the video playback client in a streaming manner.
[0127] In some embodiments, the recall result can be sent to the video playback client through the following process, including: sequentially sending multiple result shards to the video playback client based on the streaming graphic protocol. In this way, by sequentially sending multiple result shards to the video playback client in a streaming manner, the real-time nature of the recall result is ensured, enabling the video playback client to receive the result shards faster and start displaying the corresponding graphic content, shortening the waiting time of the user, improving the user experience, and thus increasing user stickiness.
[0128] In the embodiments of the present application, the voice data sent by the video playback client is recognized to obtain the target text, and the target text is subjected to intent recognition to obtain the intent recognition result; when it is determined based on the intent recognition result that the intent type is video recommendation, incremental reasoning is performed based on the intent recognition result and the target text, and the recall result is generated in a streaming manner; based on the streaming graphic protocol, the recall result is sent to the video playback client. In this way, when the voice data sent by the client belongs to video recommendation, incremental reasoning can be performed according to the intent recognition result and the target text, enriching the recall result for the voice data and improving the diversity of intelligent interaction content. Among them, the recall result at least includes the reply text and the playback control images of at least one target video. Therefore, the recall result is more intuitive and complete, facilitating the user to understand and select the recall result by combining the text and the image. In addition, through the playback control image of the target video, the user can quickly view the target video of interest, thus simplifying the user's operation process and improving the convenience and flexibility of intelligent interaction. And because the recall result is generated in a streaming manner, the recall result can be presented to the user faster, making the interaction process smoother. Therefore, the flexibility and diversity of intelligent interaction can be improved through the embodiments of the present application.
[0129] In some embodiments, when it is determined based on the intent recognition result that the intent type is the operation control type, the video playback client can also execute the control operation corresponding to the control instruction through the following process, which specifically includes:
[0130] Based on the intent recognition result and the target text, determine the control instruction; send the control instruction to the video playback client so that the video playback client executes the control operation corresponding to the control instruction.
[0131] Here, when the intent type is the operation control type, it indicates that the user wants to control the system to execute the corresponding operation, such as "open a certain channel", "increase the volume", "turn off the barrage", etc. At this time, the server will convert the user's intent into a specific control instruction according to the intent recognition result and the target text, such as the control instruction to turn off the barrage. The control instruction is used to control the video playback client to execute a specific control operation, and the control instruction is sent to the video playback client, so that the video playback client executes the corresponding control operation according to the control instruction, such as turning off the barrage.
[0132] In the embodiments of the present application, when it is determined based on the intent recognition result that the intent type is the operation control type, the control instruction is determined based on the intent recognition result and the target text; the control instruction is sent to the video playback client so that the video playback client executes the control operation corresponding to the control instruction. In this way, when the intent type is the operation control type, the corresponding control instruction can be accurately determined according to the user's intent, and the video playback client can execute the corresponding control operation, improving the control accuracy and efficiency of intelligent interaction.
[0133] In some embodiments, description information of each target video and target question text may also be generated and sent to the video playback client, specifically including:
[0134] Generate description information for each target video; for each target video, obtain target question text that meets relevant conditions with the target video from historical question text; send the description information and target question text of each target video to the video playback client.
[0135] Here, the description information of the target video is a textual description of aspects such as the release time, rating, leading actors, content, features, and attributes of the target video. The video details introduction can be directly used as the description information, or relevant information can be collected from the comment area, relevant websites, etc. as the description information. Historical question text refers to all questions asked by users in the past period, including various types of questions. Target question text refers to questions in the historical question text that meet relevant conditions with the target video and are used to guide users to conduct the next round of Q&A. The relevant conditions may refer to that the historical question text contains keywords of the target video (such as video name, core content, theme, etc.); or, the relevant conditions may be that the relevance between the historical question text and the target video is greater than a preset relevance value. For example, using natural language processing techniques, such as word vector models, semantic similarity calculations (such as cosine similarity, edit distance), etc., calculate the semantic similarity between the historical question sample and the keywords of the target video, and filter out questions with a similarity higher than a certain threshold as the target question text. Then, send the description information and target question text of each target video to the video playback client.
[0136] Exemplarily, when the user asks "Recommend funny videos", the target video "Movie A" is determined, and the description information of Movie A "Movie A tells a story of..." is generated. Then, the target question text "What is the rating of Movie A?" that meets relevant conditions with Movie A is obtained from the historical question text. At this time, "Movie A tells a story of..." and "What is the rating of Movie A?" can be sent to the video playback client, and "Movie A tells a story of..." is displayed on the video playback client, "You can also ask questions like 'What is the rating of Movie A?', 'What other movies has the leading actor B in Movie A starred in?', and let us recommend more content that meets your needs.'"
[0137] In the embodiments of the present application, description information for each target video is generated; for each target video, target problem texts that meet relevant conditions with the target video are obtained from historical problem texts; and the description information and target problem texts of each target video are sent to a video playback client. In this way, the description information and target problem texts of the target video can help users obtain videos of interest more accurately, thereby improving the accuracy and efficiency of video recommendation.
[0138] Next, taking the electronic device that implements the intelligent interaction method of the embodiments of the present application as a terminal (video playback client device), the intelligent interaction method provided by the embodiments of the present application will be described.
[0139] See Figure 4A , Figure 4A which is another flowchart of the intelligent interaction method provided by the embodiments of the present application, and will be described in combination with Figure 4A the steps shown.
[0140] In step 201, in response to the received voice collection instruction, voice data is collected and sent to the server.
[0141] Here, the user can trigger the voice collection function through keywords, buttons, gestures, etc., so that the video playback client receives the voice collection instruction and starts voice collection to obtain the corresponding voice data. For example, when the user says a specified wake-up word or presses a specified button on the remote control device, the voice recognition system of the video playback client device will receive the voice collection instruction. At this time, the voice collection function will be activated and subsequent voice instructions such as "play movie" and "recommend the latest movie" will be collected. After that, the video playback client will send the voice data to the server and wait for the server to process it.
[0142] In step 202, the target text recognized by the server for the voice data is received and the target text is displayed.
[0143] Here, the target text is the text content corresponding to the voice data. Since voice data is an information carrier recorded and stored in the form of sound and is the data obtained after digitizing human voice signals. Therefore, it is necessary for the server to convert the voice data into text through voice recognition to obtain the target text and send it to the video playback client. After receiving the target text sent by the server, the video playback client will display the target text on the display interface.
[0144] In step 203, when the intent type corresponding to the target text is a video recommendation type, the recall result sent by the server is received.
[0145] Here, the intent type refers to the category of the purpose or requirement corresponding to the intent recognition result, and the video recommendation type means that the intent recognition result indicates that the user wants the system to recommend relevant videos according to the requirement, without specifying a designated video to watch. For example, when the user says "Recommend scary TV dramas for me" to the video playback client, then the intent type is the video recommendation type. The recall result is incrementally inferred and streamed by the server based on the intent recognition result and the target text. The recall result includes at least a response text and an image of the playback control of at least one target video. Among them, incremental inference means that the inference process is not achieved overnight, but gradually deepens and improves with the acquisition of more information or further interaction of the user. Streaming the generation of the recall result means that the system does not give a response after the entire inference process is completely over, but during the inference process, once there is a partially determined recall result, it starts to output information to the user in real time.
[0146] In step 204, the recall result is parsed and displayed.
[0147] Here, through parsing, the recall result sent by the server can be converted into a meaningful and understandable form and displayed on the display interface in a streaming manner to help the user better understand and obtain the key content in the recall result.
[0148] In the embodiment of the present application, in response to the received voice collection instruction, voice data is collected and sent to the server; the target text recognized by the server for the voice data is received and displayed; when the intent type corresponding to the target text is the video recommendation type, the recall result sent by the server is received. The recall result is incrementally inferred and streamed by the server based on the intent recognition result and the target text. The recall result includes at least a response text and an image of the playback control of at least one target video; the recall result is parsed and displayed. In this way, when the voice data belongs to video recommendation, incremental inference can be performed according to the intent recognition result and the target text to enrich the recall result for the voice data and improve the diversity of intelligent interaction content. Among them, the recall result includes at least a response text and an image of the playback control of at least one target video. Therefore, the recall result is more intuitive and complete, facilitating the user to understand and select the recall result by combining the text and the image. In addition, through the image of the playback control of the target video, the user can quickly watch the target video of interest, thus simplifying the user's operation process and improving the convenience and flexibility of intelligent interaction. And because the recall result is generated in a streaming manner and is also streamed and displayed on the screen when the recall result is displayed, the recall result can be presented to the user faster, making the interaction process smoother. Therefore, the flexibility and diversity of intelligent interaction can be improved through the embodiment of the present application.
[0149] In some embodiments, the recall result includes multiple result shards. SeeFigure 4B , Step 204 can be implemented through Steps 2041 to 2043, including:
[0150] In Step 2041, based on a preset protocol, parse and process multiple result shards to obtain multiple result protocol objects, and obtain the intent type from the first result protocol object.
[0151] Here, the recall results include multiple result shards. The result shards are obtained by the server based on a preset protocol through conversion processing of the intent type, each data type, and each result data packet in the intent recognition results. Each result protocol object includes the corresponding intent type, data type, and result data packet. By parsing and processing multiple result shards through the preset protocol, multiple result protocol objects can be obtained. From the first result protocol object, obtain the intent type therein.
[0152] In Step 2042, sequentially send multiple result protocol objects to the result processor corresponding to the intent type.
[0153] Here, the result processor is used to process each result protocol object to display the recall results. There are two types of result processors. One is the result processor corresponding to the video recommendation type, which is used to stream-display the recall results. This result processor can also be used to process voice data of the chatting type. The other is the result processor corresponding to the direct access type, which is used to display the recall results at one time. The direct access type includes the video search type and the operation control type. Among them, the video search type includes the type of directly opening a certain player or the type of opening other applications that requires cross-platform.
[0154] In some embodiments, the following process can be used to sequentially send multiple result protocol objects to the protocol processor corresponding to the intent type, specifically including:
[0155] Allocate the first result protocol object to the result processor corresponding to the intent type through a protocol dispatcher; establish a data transmission channel between the protocol dispatcher and the result processor; and sequentially send each result protocol object to the result processor through the data transmission channel.
[0156] Here, the protocol dispatcher is used to allocate the result protocol objects to their corresponding result processors. Since the intent type is the video recommendation type, the corresponding result processor is used to stream-display the recall results. At this time, the protocol dispatcher first allocates the first result protocol object to the result processor corresponding to the intent type, and establishes a data transmission channel between the protocol dispatcher and the result processor. Among them, the data transmission channel refers to the path or medium for transmitting data in a computer system, network, or other data processing environment. The data transmission channel is like a "highway" that allows data to be transmitted and exchanged between different devices, systems, or locations. After that, through the data transmission channel, each result protocol object is sequentially sent to the result processor to implement the streaming sending process of each result protocol object.
[0157] In the embodiment of the present application, the first result protocol object is allocated to the result processor corresponding to the intent type through the protocol dispatcher; a data transmission channel is established between the protocol dispatcher and the result processor; and each result protocol object is sequentially sent to the result processor through the data transmission channel. In this way, the streaming sending of each result protocol object can be achieved, and the sending efficiency of each result protocol object can be improved.
[0158] In step 2043, each result protocol object is processed by the result processor to display the recall results.
[0159] Here, the result processor is used to comprehensively process each result protocol object. For example, each result protocol object is sorted according to its relevance to the target text, and each result protocol object is displayed in a certain format so that users can clearly understand and view the results.
[0160] In the embodiment of the present application, multiple result shards are parsed and processed based on a preset protocol to obtain multiple result protocol objects, and the intent type is obtained from the first result protocol object; the multiple result protocol objects are sequentially sent to the result processor corresponding to the intent type; and each result protocol object is processed by the result processor to display the recall results. In this way, the result processor can effectively process each result protocol object, display the recall results in a manner that better meets user needs and is easy to understand, and improve the accuracy of video recommendation and the user experience.
[0161] In some embodiments, referring to Figure 4C , step 2043 can be implemented through steps 20431 to 20433, including:
[0162] In step 20431, the result processor determines the display layout of the recall results based on the intent type.
[0163] Here, the display layout includes display areas corresponding to response data of different data types in the Q&A interface. That is to say, response data of different data types each have corresponding display areas in the Q&A interface. Among them, the data types of the response data include text type and image type. The response data of the image type can be the poster of the target video recommended for the target user.
[0164] In step 20432, result data packets are obtained from each result protocol object, and the result data packets of the same data type are assembled to obtain response data of each data type.
[0165] Here, the result protocol object includes the corresponding intent type, data type, and result data packet. The corresponding result data packets are obtained from each result protocol object, and then the result data packets of the same data type are assembled to obtain the response data corresponding to each data type, such as response data of the text type and response data of the image type.
[0166] In step 20433, the response data of each data type are stream-displayed in the display areas corresponding to each data type.
[0167] Here, the response data of each data type are stream-displayed respectively in the display areas corresponding to the data types. For example, the response data of the question type are stream-displayed in display area 1 corresponding to the text type, and the response data of the image type are stream-displayed in display area 2 corresponding to the image type, realizing the partition display of different types of response data.
[0168] In the embodiments of the present application, the result processor determines the display layout of the recall results based on the intent type. The display layout includes display areas corresponding to response data of different data types in the Q&A interface; result data packets are obtained from each result protocol object, and the result data packets of the same data type are assembled to obtain response data of each data type; the response data of each data type are stream-displayed in the display areas corresponding to each data type. In this way, different types of response data can be partitioned and stream-displayed, making the recall results clearer and more intuitive, facilitating user understanding, and thus improving the efficiency of video recommendation.
[0169] In some embodiments, the description information of the selected target video and the target question text can also be obtained through the following process, specifically including:
[0170] Receive the description information of each target video and the target question text sent by the server; in response to receiving a selection instruction for the target video, display the description information of the selected target video and the target question text in the Q&A interface.
[0171] Here, the description information of the target video is a textual description of aspects such as the content, features, and attributes of the target video. The video details introduction can be directly used as the description information, or relevant information can be collected from the comment area, related websites, etc. as the description information. The target question text refers to the questions in the historical question text that meet the relevant conditions with the target video and is used to guide the user to conduct the next round of Q&A. In response to receiving a selection instruction for the target video, the description information of the selected target video and the target question text are displayed in the Q&A interface. When the user clicks, touches (for touch devices), or uses other interaction methods (such as selecting after hovering the mouse, etc.) to operate on a certain target video, the interaction monitoring module of the system will sense this operation and identify it as a selection instruction for this target video, and then display the description information of the selected target video and the target question text in the Q&A interface.
[0172] In the embodiment of the present application, receive the description information of each target video and the target question text sent by the server; in response to receiving a selection instruction for the target video, display the description information of the selected target video and the target question text in the Q&A interface. In this way, it can be realized that in response to the user's selection operation on the target video, the relevant description information and the target question text are accurately displayed in the Q&A interface, providing the user with richer video-related information and interaction experience, and improving the diversity of intelligent interaction.
[0173] In some embodiments, when the intent type corresponding to the target text is the operation control type, the control operation corresponding to the control instruction can also be executed through the following process, which specifically includes:
[0174] When the intent type corresponding to the target text is the operation control type, receive the control instruction sent by the server; execute the control operation corresponding to the control instruction.
[0175] Here, when the intent type is the operation control type, it indicates that the user wants to control the system to execute the corresponding operation, such as "open a certain channel", "increase the volume", "turn off the barrage", etc. At this time, the server will convert the user's intent into a specific control instruction according to the intent recognition result and the target text, such as the control instruction to turn off the barrage. After receiving the control instruction sent by the server, the video playback client will execute the specific control operation according to the control instruction to complete the user's instruction.
[0176] In the embodiment of the present application, when the intent type corresponding to the target text is the operation control type, receive the control instruction sent by the server; execute the control operation corresponding to the control instruction. In this way, the user can complete the control instruction only by speaking, greatly improving the convenience of operation and the flexibility and diversity of intelligent interaction.
[0177] The following provides an intelligent interaction method, which is applied toFigure 1 The intelligent interaction system shown, refer to Figure 5 , Figure 5 is another schematic diagram of the implementation process of the intelligent interaction method provided by the embodiments of the present application. The following will be described in conjunction with Figure 5 for illustration.
[0178] In step 301, the terminal 400 responds to the received voice collection instruction, collects voice data, and sends the voice data to the server 200.
[0179] Here, the user can trigger the voice collection function through keywords, buttons, gestures, etc., so that the terminal receives the voice collection instruction and starts to collect voice data to obtain the corresponding voice data. For example, when the user says the specified wake-up word or presses the specified button on the specified device, the voice recognition system of the terminal will receive the voice collection instruction. At this time, the voice collection function will be activated and start to collect subsequent voice instructions, such as "play a movie", "recommend the latest movie", etc. Then, the terminal will send the voice data to the server and wait for the server to process it.
[0180] In step 302, the server 200 recognizes the voice data sent by the terminal 400 to obtain the target text, sends it to the terminal 400, and performs intent recognition on the target text to obtain the intent recognition result.
[0181] Here, the voice data is an information carrier recorded and stored in the form of sound, which is the data obtained after digitizing the human voice signal. For example, when the user says "find a comedy movie", the voice recognition system on the server side processes and analyzes the voice data, and finally converts it into a target text such as "find a comedy movie". Then, by performing intent recognition on the target text to understand the true intention or purpose behind the target text, the intent recognition result is obtained. The intent recognition result is used to describe the intention expressed by the target text, such as "search for a certain movie", "close a certain application", etc.
[0182] In step 303, the terminal 400 receives the target text recognized by the server 200 for the voice data and displays the target text.
[0183] Here, after the video playback client receives the target text sent by the server, it will display the target text on the display interface.
[0184] In step 304, when the server 200 determines that the intent type is the video recommendation type based on the intent recognition result, it performs incremental reasoning based on the intent recognition result and the target text, generates the recall result in a streaming manner, and sends it to the terminal 400 based on the streaming graphic protocol.
[0185] Here, the intent type refers to the category of the purpose or requirement corresponding to the intent recognition result. The video recommendation type means that the intent recognition result indicates that the user wants the system to recommend relevant videos according to the requirement, without specifying a designated video to be watched. For example, when the user says to the video playback client, "Recommend scary TV dramas for me", then the intent type is the video recommendation type. Incremental reasoning means that the reasoning process is not achieved overnight, but gradually deepens and improves as more information is obtained or with further interaction from the user. Streaming generation of recall results means that the system does not give a response only after the entire reasoning process is completely finished, but during the reasoning process, once there are some definite recall results, it starts to output information to the user in real time. For example, first, based on the information "science fiction movies", all science fiction movies are screened out from the movie database, and then all science fiction movies will be presented to the user as partial recall results. Then, based on the condition "with wonderful special effects", all the previously screened science fiction movies are further evaluated and screened, and the screened movies are presented to the user as another part of the recall results in a streaming manner. Among them, the recall results at least include a response text and an image of the playback control of at least one target video. The response text is used to reply to the target text corresponding to the voice data. The target video is a recommended video that meets the user's requirements determined according to the intent recognition result. The image of the playback control of the target video refers to an image related to the target video with the function of controlling video playback. For example, the playback control image is the poster of a certain movie, and there is a playback control set on part of the poster. By clicking on the poster, the playback control can be triggered, and then directly jump to the video playback page to play the corresponding movie.
[0186] In step 305, the terminal 400 receives the recall results sent by the server 200, parses and displays the recall results.
[0187] Here, through parsing, the recall results sent by the server can be converted into a meaningful and understandable form and displayed on the display interface in a suitable manner to help the user better understand and obtain the key content in the recall results.
[0188] In the embodiments of the present application, the terminal collects voice data in response to the received voice collection instruction, and sends the voice data to the server. The server recognizes the voice data sent by the video playback client to obtain the target text, and sends it to the terminal, and the terminal displays the target text. The server also performs intent recognition on the target text to obtain an intent recognition result; when it is determined that the intent type is video recommendation based on the intent recognition result, incremental reasoning is performed based on the intent recognition result and the target text, and the recall result is generated in a streaming manner. Based on the streaming graphic protocol, the recall result is sent to the video playback client. The terminal receives the recall result and parses and displays the recall result. In this way, when the voice data sent by the client belongs to video recommendation, incremental reasoning can be performed according to the intent recognition result and the target text, enriching the recall result for the voice data and improving the diversity of intelligent interaction content. Among them, the recall result at least includes the reply text and the playback control images of at least one target video. Therefore, the recall result is more intuitive and complete, facilitating the user to understand and select the recall result by combining the text and the image. In addition, through the playback control image of the target video, the user can quickly view the target video of interest, thus simplifying the user's operation process and improving the convenience and flexibility of intelligent interaction. And because the recall result is generated in a streaming manner, the recall result can be presented to the user faster, making the interaction process smoother. Therefore, the flexibility and diversity of intelligent interaction can be improved through the embodiments of the present application.
[0189] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.
[0190] The embodiments of the present application provide an intelligent interaction method for streaming and generating an AI voice interaction interface through a streaming graphic protocol. On the one hand, with the help of the streaming graphic protocol, the recommended results to be displayed can be constructed faster and more flexibly. On the other hand, with the help of the thinking and understanding ability of the large language model, it can be sent while thinking. And by fine-tuning the recommended results based on the large language model, the recommended results can be better combined with the business. By understanding the user's intent through the large language model and generating the corresponding recommended results, and combining the recommended results with the video playback client, an intelligent interaction method implemented on OTT can be realized, thereby improving the flexibility of the video playback client operation and the efficiency of the user searching for videos.
[0191] In the embodiments of the present application, globally triggerable AI voice interaction is set in the video playback application. The externally visible entry of the AI voice interaction is set on the home page, personal center, viewing history, filtering page of the video playback client, or interfaces such as the manufacturer's desktop. The voice bar status of the automatic speech recognition technology is as follows, including:
[0192] 1) Guidance state: The guidance state is displayed when the user initially enters the video playback application. See Figure 6, Figure 6 is a schematic diagram of the display state in the guiding state provided by an embodiment of the present application. As Figure 6 shown, the prompt word 61 is used to prompt the user to activate the voice interaction function through the corresponding remote control button. If the user has the habit of continuing to watch, the guiding state can prompt the user through the prompt word 62 that after activating the voice interaction function, the user can say "continue playing" in voice to continue watching a certain video that has been watched before. Alternatively, the prompt word 62 can also be set to "please tell me to open a certain channel", which is used to guide the user to perform operations such as opening an application with a long original path by saying "open a certain channel" in voice after activating the voice interaction function. Among them, the specific settings of the guiding state can be combined with guiding prompts such as platform operation, behavior (appointment online behavior, continuing to watch), and video content (highlights, actors).
[0193] 2) Wake-up state: The voice interaction function has been activated and enters the voice input state. As Figure 7 shown, Figure 7 is a schematic diagram of the display state in the wake-up state provided by an embodiment of the present application. At this time, the icon shown as 71 can be displayed on the display interface to prompt the user that voice input can be performed. The user can perform voice input in this state, and the video playback client can collect voice data.
[0194] 3) Recognition state: Recognize the voice data, and quickly recognize the voice data through general logics such as parsing scene information, manufacturer authentication, intent list processing, text error correction, and word segmentation to obtain the target text. As Figure 8 shown, Figure 8 is a schematic diagram of the display state in the recognition state provided by an embodiment of the present application. The voice data input by the user is recognized as the target text 81 and is displayed in the voice bar.
[0195] 4) Understanding state: Perform intent recognition on the target text to obtain the corresponding intent recognition result. As Figure 9 shown, Figure 9 is a schematic diagram of the display state in the understanding state provided by an embodiment of the present application. The dynamic effect of the voice bar will switch to thinking, indicating that intent recognition is being performed on the target text.
[0196] 5) Execution state: For non-video recommendation types or chat types of intents, use the film and television model service and the intent recognition results, combined with some operation data of the video, to perform content recall. For general vertical intents (video playback type or operation control type), a traditional large language model will be used for content recall. The execution results are divided into four types: Type 1 corresponds to video recommendation types or chat types, and the speech bar will pull up a half-screen panel from the bottom of the interface; Types 2, 3, and 4 are all direct access types, corresponding to video playback types or operation control types. Among them, Type 2 is a cross-application direct access type, Type 3 is a direct access type with a unique result that directly opens the player or the details page, and Type 4 is a direct access type for operation execution. For Types 2, 3, and 4, the speech bar will display a prompt word indicating successful execution, such as Figure 10A as shown, the prompt word 1001 "Playing 《XX》 for you" in the speech bar is used to prompt that the specified operation has been completed.
[0197] Among them, in the execution state, for video recommendation types or chat types, the speech bar will also generate a floating interface, on which the text obtained by recognizing the user's input speech and the recall results returned by the large language model through conversational interaction are displayed. This can solve the problem that users are not clear about the ability boundary of the large model and can also make full use of the OTT large screen. The floating interface is divided into a left-right structure. The left area is used to display the streaming conversational interaction process; the right area is used to display the description information of the selected target video, etc., according to the content of the streaming conversational interaction, or provide targeted prompt words (corresponding to the target question text in other embodiments) returned based on the user's input speech data, such as suggestions for "what else can I ask". Among them, the recall results include two result styles: pure text and graphic types. For the pure text type of conversational interaction: on the left is the content of the streaming conversational interaction, including the question input by the user and the recall results. On the right, relevant questions for further recall are continued based on the recalled content to help the user have a better second-round conversation. For the graphic type of conversational interaction: on the left is the content of the streaming conversational interaction, including the question input by the user and the recall results. The relevant content of the pre-positioned video details page is included in the recall results. Through the playback control image of the target video in the recall results, the user's decision-making path can be shortened. On the right, richer description information is displayed according to the poster focused by the user (corresponding to the playback control image of the target video in other embodiments), and the description information can include the release time, leading actors of the target video, and recommended reasons generated by the large model, etc.
[0198] For the interaction method of result presentation in multi-round conversations, the next round of conversation can only be focused after the content in the left area of the floating interface has scrolled through. As Figure 10BAs shown in the figure, during the process of turning the page down, the pure text 1002 in the left area, the poster 1003 of the target video gains focus and moves down in sequence. The poster 1003 coexists with the pure text 1002, and the page turns down half-screen by half-screen. The right area is used to display the description information 1004 of the selected target video and the prompt words 1005 returned according to the user's question. When the content in the left area exceeds one screen, it turns the page down to get Figure 10C . As Figure 10C shown, other content 1006 of the streaming conversation interaction continues to be displayed downward in the left area, and the prompt words 1005 "You may also want to ask" in the right area remain unchanged. Among them, during the process of recalling the results, it takes time for the large model to return the recall results, and during this period, the "thinking" process of the model needs to be displayed. And, during the process of recalling the results, the user can pause at any time.
[0199] See Figure 11 . Figure 11 FIG. is another flowchart of the intelligent interaction method provided by the embodiment of the present application. After the user 111 initiates the voice, the voice collector 1121 in the video playback client 112 will collect the voice data and send it to the speech recognizer 1131 in the server 113. The speech recognizer 1131 will convert the voice data into the target text, and then initiate an intent recognition request to the intent decision layer 1132. The intent decision layer 1132 performs intent recognition on the target text to obtain the intent recognition result. Among them, after receiving the speech awareness request, the intent decision layer 1132 will interface with the large model intent service 1133 (corresponding to the first intent service in other embodiments), or can also interface with the traditional intent service (corresponding to the second intent service in other embodiments). Among them, the traditional intent service is mainly used to process the intent recognition of some simple texts, and the large model intent service 1133 is mainly used to process the intent recognition of some long texts and complex problems. If the intent decision layer determines that the intent recognition result is a direct access type or an operation control type, it will issue the complete recall result, and the video playback client 112 will process it all at once. If the intent decision layer 1132 determines that the intent recognition result is a video recommendation type or a chat type, it will enter the in-depth thinking process, and during the thinking process, multiple result shards will be generated in a streaming manner as the recall result, and the recall result will be transmitted to the video playback client 112 by streaming the intent result shards. Figure 11 FIG. takes the case where the intent recognition result is a video recommendation type or a chat type as an example. Among them, when performing streaming data transmission, the Server-Sent Events (SSE) technology in the Hyper Text Transfer Protocol (HTTP) is used to implement. The process is as Figure 12 shown. Figure 12It is a schematic flowchart of a streaming data transmission provided by an embodiment of the present application. Among them, the video playback client 112 first sends a streaming graphic protocol request to the server 113, and the server 113 streams and sends result shards to the client 112 according to the streaming graphic protocol request. After receiving the result shards, the video playback client 112 can parse the result shards based on a preset protocol 1122 to obtain a result protocol object. The data format of the result protocol object is specified in the preset protocol. The video playback client 112 obtains the intent type from the first result protocol object and distributes the first result protocol object to the result processor 1124 corresponding to the intent object through the protocol dispatcher 1123. After processing the first result protocol object, the result processor 1124 establishes an independent intent protocol distribution channel with the protocol dispatcher, and subsequent result protocol objects will be directly transmitted to this independent result processor along this channel, improving the distribution efficiency of the result protocol objects.
[0200] Among them, since the reply data obtained by the background search contains data of different data types such as text and images, the background will abstract the reply data of the same data type into a kind of card, and the complete card will be split into multiple result data packets by the background and placed in the result shards, and each result shard contains one result data packet. As Figure 13 shown, Figure 13 It is a schematic flowchart of a card generation method provided by an embodiment of the present application. After receiving the result shards, the result processor 1124 of the video playback client assembles the multiple result shards to obtain cards of different data types, including a text card 131 and an image list card 132, and then streams the text card 131 and the image list card 132 onto the screen respectively and displays them in a dialog box 133 (corresponding to the Q&A interface in other embodiments).
[0201] In the dialog box, each card has a corresponding display area. As Figure 14 shown, Figure 14 It is a schematic diagram of card display in a dialog box provided by an embodiment of the present application, which includes: inquiry cards 141 and 142 for presenting questions issued by the user; reply cards 143 and 144 for presenting replies to the questions issued by the user; the reply card 143 includes a text card 1431 for presenting text content and a poster list card 1432 for presenting multiple poster cards corresponding to the target video; the reply card 144 includes a text card 1441 for presenting text content; the album details card 145 corresponds to the poster cards in the poster list card 1432 and is used to present the description information of the target video; the recommended question language card 146 is used to present the target question text.
[0202] See Figure 15, Figure 15 1 is a display schematic diagram of a question-and-answer interface provided in an embodiment of the present application, which includes a target text 151 , a target video 152 , description information 153 and a target question text 154 .
[0203] In the related art, after the user initiates a voice request, the background will return the recall result at one time after understanding the user's complete intention as a whole. Therefore, the overall time consumed is long, and the returned content is very simple and not user-friendly. In the application embodiment, the recall result is generated in a streaming manner based on a large language model, and the recall result is understood, generated, and displayed in a streaming manner. In this way, the recall result can be delivered to the client for display more quickly. In addition, the content of the recall result is not fixed. It is a graphic result generated by deep thinking based on a large language model. The recall result is more user-friendly and the content is richer. Users can also directly play recommended videos through the recall results, which fully reflects the flexibility and diversity of intelligent interaction and improves the user experience.
[0204] The embodiments of the present application can handle the multi-intent semantics of users, provide users with diversified search results through streaming graphics and text protocols, and perform video search and intelligent interaction based on the user's past behavior preferences and intent recognition results on the intelligent interaction logic of the long video platform. In addition, traditional OTT products are limited by remote control operation, with low operating efficiency and too long to reach commonly used paths. However, large-screen users have a wide coverage, and it is not convenient to search for videos through remote controls for different user groups. The voice technology in the related art remains at the level of fixed instructions and simply replacing remote control operations, and has not achieved a state where all people and multiple tasks can be completed. The operation method that transcends the remote control constructed by the embodiments of the present application can improve the user's operating efficiency and provide a new experience for the user.
[0205] The following is a description of an exemplary structure of the intelligent interaction device 233 provided in the embodiment of the present application implemented as a software module. In some embodiments, Figure 2A As shown, the software modules stored in the intelligent interaction device 233 of the memory 230 may include:
[0206] The recognition module 2331 is used to recognize the voice data sent by the video playback client to obtain the target text, and perform intent recognition on the target text to obtain the intent recognition result;
[0207] A determination module 2332 is used for, when it is determined based on the intent recognition result that the intent type is a video recommendation type, performing incremental reasoning based on the intent recognition result and the target text, and streamingly generating a recall result, wherein the recall result at least includes a reply text and a playback control image of at least one target video;
[0208] A sending module 2333, configured to send the recall result to the video playback client based on a streaming graphic protocol.
[0209] In some embodiments, the recognition module 2331 is further configured to perform word segmentation on the target text to obtain a plurality of target word segments; determine a target intent service based on the plurality of target word segments, and perform intent recognition on the plurality of target word segments based on the target intent service to obtain an intent recognition result.
[0210] In some embodiments, the determination module 2332 is further configured to determine the total number of word segments of the target word segments; determine the target intent service based on the total number of word segments, where when the total number of word segments is greater than a preset number threshold, the first intent service is determined as the target intent service; when the total number of word segments is less than or equal to the number threshold, the second intent service is determined as the target intent service.
[0211] In some embodiments, the determination module 2332 is further configured to perform text component analysis on the target text to obtain the text components of each of the target word segments; determine the target intent service based on the text components of each of the target word segments, where when the text components of each of the target word segments include a preset component, the first intent service is determined as the target intent service; when the text components of each of the target word segments do not include the preset component, the second intent service is determined as the target intent service.
[0212] In some embodiments, the determination module 2332 is further configured to decompose the target text based on the intent recognition result to obtain a plurality of sub-questions with logical relationships; for each sub-question, retrieve target knowledge information for the sub-question in a knowledge database, and determine the reasoning result of the sub-question based on the target knowledge information and the reasoning result of the previous sub-question; streamingly generate a recall result based on the reasoning result of each sub-question and the intent recognition result.
[0213] In some embodiments, the determination module 2332 is further configured to generate a reply result for each sub-question based on the reasoning result of each sub-question, where the reply result includes reply data of at least one data type; perform segmentation processing on the reply data corresponding to each data type to obtain a plurality of result data packets; perform conversion processing on the intent type in the intent recognition result, each data type, and each result data packet respectively based on a preset protocol to obtain a plurality of result shards.
[0214] In some embodiments, the sending module 2333 is further configured to sequentially send the plurality of result shards to the video playback client based on the streaming graphic protocol.
[0215] In some embodiments, the determining module 2332 is further configured to, when determining that the intent type is an operation control type based on the intent recognition result, determine a control instruction based on the intent recognition result and the target text; and send the control instruction to the video playback client so that the video playback client performs a control operation corresponding to the control instruction.
[0216] In some embodiments, the sending module 2333 is further configured to generate description information for each of the target videos; for each target video, obtain a target question text that satisfies relevant conditions with the target video from historical question texts; and send the description information and the target question text of each target video to the video playback client.
[0217] Next, the exemplary structure of the intelligent interaction device 455 provided in the embodiments of the present application as a software module will be continued. In some embodiments, as Figure 2B shown, the software module in the intelligent interaction device 455 stored in the memory 450 may include:
[0218] An acquisition module 4551, configured to acquire voice data in response to a received voice acquisition instruction and send the voice data to the server;
[0219] A receiving module 4552, configured to receive the target text recognized by the server for the voice data and display the target text;
[0220] The receiving module 4552 is further configured to, when the intent type corresponding to the target text is a video recommendation type, receive a recall result sent by the server, where the recall result is incrementally inferred and stream-generated by the server based on the intent recognition result and the target text, and the recall result includes at least a reply text and a playback control image of at least one target video;
[0221] A parsing and display module 4553, configured to parse and display the recall result.
[0222] In some embodiments, the parsing and display module 4553 is further configured to perform parsing processing on multiple result shards based on a preset protocol to obtain multiple result protocol objects, and obtain an intent type from the first result protocol object; sequentially send the multiple result protocol objects to a result processor corresponding to the intent type; and process each result protocol object through the result processor to display the recall result.
[0223] In some embodiments, the parsing and display module 4553 is further configured to allocate the first result protocol object to the result processor corresponding to the intent type through a protocol dispatcher; establish a data transmission channel between the protocol dispatcher and the result processor; and sequentially send each of the result protocol objects to the result processor through the data transmission channel.
[0224] In some embodiments, the parsing and display module 4553 is further configured to determine, by the result processor, a display layout of the recall result based on the intent type, where the display layout includes display areas corresponding to response data of different data types in a question-and-answer interface; obtain result data packets from each of the result protocol objects, and assemble the result data packets of the same data type to obtain response data of each data type; and stream-display the response data of each data type in the display area corresponding to each data type.
[0225] In some embodiments, the receiving module 4552 is further configured to receive description information of each of the target videos and the target question text sent by a server; and in response to receiving a selection instruction for a target video, display the description information of the selected target video and the target question text in the question-and-answer interface.
[0226] In some embodiments, the receiving module 4552 is further configured to, when the intent type corresponding to the target text is an operation control type, receive a control instruction sent by the server; and execute a control operation corresponding to the control instruction.
[0227] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions, and the computer program or computer-executable instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium, and the processor executes the computer program or computer-executable instructions, so that the electronic device executes the intelligent interaction method in the embodiments of the present application as described above.
[0228] An embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions or a computer program are stored. When the computer-executable instructions or the computer program are executed by a processor, the processor will be caused to execute the intelligent interaction method provided by the embodiments of the present application. For example, as Figure 3A or Figure 4A shown in the intelligent interaction method.
[0229] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0230] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0231] As an example, the computer-executable instructions may or may not correspond to a file in a file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program in question, or, stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or portions of code).
[0232] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0233] In summary, in the embodiments of the present application, the voice data sent by the video playback client is recognized to obtain a target text, and the target text is subjected to intent recognition to obtain an intent recognition result; when it is determined based on the intent recognition result that the intent type is video recommendation, incremental reasoning is performed based on the intent recognition result and the target text, and the recall result is generated in a streaming manner; based on the streaming graphic protocol, the recall result is sent to the video playback client. In this way, when the voice data sent by the client belongs to video recommendation, incremental reasoning can be performed according to the intent recognition result and the target text, enriching the recall result for the voice data and improving the diversity of intelligent interaction content. Among them, the recall result at least includes a response text and a playback control image of at least one target video. Therefore, the recall result is more intuitive and complete, facilitating the user to understand and select the recall result by combining the text and the image. In addition, through the playback control image of the target video, the user can quickly view the target video of interest, thus simplifying the user's operation process and improving the convenience and flexibility of intelligent interaction. And, since the recall result is generated in a streaming manner, the recall result can be presented to the user faster, making the interaction process smoother. Therefore, the flexibility and diversity of intelligent interaction can be improved through the embodiments of the present application.
[0234] As described above, the above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are all included within the protection scope of the present application.
Claims
1. An intelligent interaction method, characterized in that, The method includes: Recognize the voice data sent by the video playback client to obtain the target text, and perform intent recognition on the target text to obtain an intent recognition result; When it is determined based on the intent recognition result that the intent type is a video recommendation type, perform incremental reasoning based on the intent recognition result and the target text, and streamingly generate a recall result, where the recall result at least includes a reply text and a playback control image of at least one target video; Send the recall result to the video playback client based on the streaming graphic protocol.
2. The method according to claim 1, wherein The performing intent recognition on the target text to obtain an intent recognition result includes: Perform word segmentation on the target text to obtain multiple target word segments; Based on the multiple target word segments, determine a target intent service, and perform intent recognition on the multiple target word segments based on the target intent service to obtain an intent recognition result.
3. The method according to claim 2, characterized in that, The determining a target intent service based on the multiple target word segments includes: Determine the total number of word segments of the target word segments; Determine the target intent service based on the total number of word segments. Among them, when the total number of word segments is greater than a preset number threshold, determine the first intent service as the target intent service; when the total number of word segments is less than or equal to the number threshold, determine the second intent service as the target intent service.
4. The method according to claim 2, characterized in that, The determining a target intent service based on the multiple target word segments includes: Perform text component analysis on the target text to obtain the text components of each target word segment; Determine the target intent service based on the text components of each target word segment. Among them, when the text components of each target word segment include a preset component, determine the first intent service as the target intent service; when the text components of each target word segment do not include a preset component, determine the second intent service as the target intent service.
5. The method according to claim 1, wherein Performing incremental reasoning based on the intent recognition result and the target text, and streamingly generating a recall result includes: Based on the intent recognition result, decompose the target text to obtain multiple sub-questions with logical relationships; For each sub-question, retrieve target knowledge information for the sub-question in the knowledge database, and determine the reasoning result of the sub-question based on the target knowledge information and the reasoning result of the previous sub-question; Streamingly generate a recall result based on the reasoning result of each sub-question and the intent recognition result.
6. The method according to claim 5, wherein The recall result includes multiple result shards. The streamingly generating a recall result based on the reasoning result of each sub-question and the intent recognition result includes: Generate a reply result for each sub-question based on the reasoning result of each sub-question, where the reply result includes reply data of at least one data type; Perform segmentation processing on the reply data corresponding to each data type to obtain multiple result data packets; Based on a preset protocol, perform conversion processing on the intent type in the intent recognition result, each data type, and each result data packet respectively to obtain multiple result shards.
7. The method according to claim 6, characterized in that, Sending the recall result to the video playback client based on the streaming graphic protocol includes: Based on the streaming graphic protocol, sequentially sending multiple result shards to the video playback client.
8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: When it is determined that the intent type is an operation control type based on the intent recognition result, determining a control instruction based on the intent recognition result and the target text; Sending the control instruction to the video playback client so that the video playback client executes the control operation corresponding to the control instruction.
9. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Generating description information for each of the target videos; For each target video, obtaining a target question text that satisfies relevant conditions with the target video from historical question texts; Sending the description information and the target question text of each target video to the video playback client.
10. An intelligent interaction method, characterized in that, The method includes: In response to a received voice collection instruction, collecting voice data and sending the voice data to the server; Receiving the target text recognized by the server for the voice data and displaying the target text; When the intent type corresponding to the target text is a video recommendation type, receiving a recall result sent by the server, where the recall result is incrementally inferred and stream-generated by the server based on the intent recognition result and the target text, and the recall result at least includes a reply text and a playback control image of at least one target video; Parsing and displaying the recall result.
11. The method according to claim 10, wherein, The recall result includes multiple result shards, and parsing and displaying the recall result includes: Performing parsing processing on multiple result shards based on a preset protocol to obtain multiple result protocol objects, and obtaining the intent type from the first result protocol object; Sequentially sending multiple result protocol objects to a result processor corresponding to the intent type; Processing each result protocol object through the result processor to display the recall result.
12. The method according to claim 11, characterized in that, The sequentially sending multiple result protocol objects to a protocol processor corresponding to the intent type includes: Allocating the first result protocol object to a result processor corresponding to the intent type through a protocol dispatcher; Establishing a data transmission channel between the protocol dispatcher and the result processor; Sequentially sending each result protocol object to the result processor through the data transmission channel.
13. The method according to claim 11, wherein Processing each result protocol object through the result processor to display the recall result: Based on the intent type, determining a display layout of the recall result through the result processor, where the display layout includes display areas corresponding to reply data of different data types in a question-and-answer interface; Obtaining result data packets from each result protocol object and assembling result data packets of the same data type to obtain reply data of each data type; Stream-displaying the reply data of each data type in the display area corresponding to each data type.
14. The method according to any one of claims 10 to 13, characterized in that, The method further includes: Receiving the description information and target question text of each target video sent by the server; In response to receiving a selection instruction for a target video, display the description information of the selected target video and the target question text in the Q&A interface.
15. The method according to any one of claims 10 to 13, characterized in that, The method further includes: When the intent type corresponding to the target text is an operation control type, receive a control instruction sent by the server; Execute the control operation corresponding to the control instruction.
16. An intelligent interaction device, characterized in that, The device includes: An identification module, configured to identify the voice data sent by the video playback client to obtain a target text, and perform intent recognition on the target text to obtain an intent recognition result; A determination module, configured to, when it is determined based on the intent recognition result that the intent type is a video recommendation type, perform incremental reasoning based on the intent recognition result and the target text to streamingly generate a recall result, where the recall result at least includes a reply text and a playback control image of at least one target video; A sending module, configured to send the recall result to the video playback client based on a streaming graphic protocol.
17. An intelligent interaction device, characterized in that, The device includes: An acquisition module, configured to, in response to receiving a voice acquisition instruction, acquire voice data and send the voice data to the server; A receiving module, configured to receive the target text recognized by the server for the voice data and display the target text; The receiving module is further configured to, when the intent type corresponding to the target text is a video recommendation type, receive a recall result sent by the server, where the recall result is streamingly generated by the server through incremental reasoning based on the intent recognition result and the target text, and the recall result at least includes a reply text and a playback control image of at least one target video; A parsing and display module, configured to parse and display the recall result.
18. An electronic device, characterized in that, The electronic device includes: a memory, configured to store computer-executable instructions or a computer program; a processor, configured to, when executing the computer-executable instructions or the computer program stored in the memory, implement the method according to any one of claims 1 to 9, or any one of claims 10 to 15.
19. A computer-readable storage medium stores computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or the computer program are executed by the processor, the method according to any one of claims 1 to 9, or any one of claims 10 to 15 is implemented.
20. A computer program product, comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or the computer program are executed by the processor, the method according to any one of claims 1 to 9, or any one of claims 10 to 15 is implemented.
Citation Information
Cited By
Text-to-voice real-time streaming conversion method, system and device, medium and program product
CN121565135A
Unified streaming processing method, system and device for multi-mode AI interactive content, medium and program product
CN121705057A