Human-computer interaction method and device, electronic equipment and storage medium

By acquiring the current video frame and its historical frame sequence from the video stream, and using a large model to generate proactive dialogue content, the problem of human-computer interaction relying on explicit input in existing technologies is solved, achieving a more intelligent and interactive interactive experience.

CN120872141APending Publication Date: 2025-10-31BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510898167.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In existing human-computer interaction models, the system's response behavior mainly relies on the user's explicit input, lacking proactive perception and response capabilities, resulting in high complexity and insufficient interactivity in the interaction process.

Method used

By acquiring the current video frame and its historical frame sequence from the video stream, proactive dialogue content is generated using a large model, including a combination of classification models and large models. This allows for the rapid identification and output of proactive dialogue content, reducing user operation complexity and enhancing interactivity.

Benefits of technology

This system enables the system to proactively identify and execute interactive operations that match the user's intent without requiring explicit user input, thereby improving the intelligence level of human-computer interaction and the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872141A_ABST
    Figure CN120872141A_ABST
Patent Text Reader

Abstract

The invention provides a man-machine interaction method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, augmented reality, large models, natural language processing and the like. According to the specific implementation scheme, on the basis of a current video frame in a video stream and a historical video frame sequence of the current video frame, under the condition that it is determined that an active dialogue needs to be initiated on the current video frame, active dialogue content is generated through a large model according to the current video frame and the historical video frame sequence, and the active dialogue content is output. Therefore, in the man-machine interaction process, the active dialogue is actively initiated, the identification and execution of the interaction operation matched with the user intention are improved, the operation complexity of the user is reduced, the interaction of man-machine interaction is enhanced, the intelligent level of interaction is effectively improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the fields of computer vision, augmented reality, large models, and natural language processing, and especially to a human-computer interaction method, device, electronic device, and storage medium. Background Technology

[0002] With the development of artificial intelligence and computer vision technologies, more and more intelligent systems are gaining the ability to proactively perceive their environment. However, in current mainstream human-computer interaction models, the system's response behavior still primarily relies on explicit user input. Typically, the system only responds based on the corresponding video stream content after the user initiates a dialogue via voice or text. Summary of the Invention

[0003] This disclosure provides a human-computer interaction method, apparatus, electronic device, and storage medium.

[0004] According to one aspect of this disclosure, a human-computer interaction method is provided, the method comprising: acquiring a current video frame in a video stream and acquiring a sequence of historical video frames preceding the current video frame; determining, based on the current video frame and the sequence of historical video frames, that an active dialogue needs to be initiated on the current video frame; generating active dialogue content based on the current video frame and the sequence of historical video frames using a large model; and outputting the active dialogue content.

[0005] According to another aspect of this disclosure, a human-computer interaction device is provided, comprising: an acquisition module, configured to acquire a current video frame in a video stream and acquire a sequence of historical video frames preceding the current video frame; a determination module, configured to determine, based on the current video frame and the sequence of historical video frames, that an active dialogue needs to be initiated on the current video frame; a generation module, configured to generate active dialogue content using a large model, based on the current video frame and the sequence of historical video frames; and an output module, configured to output the active dialogue content.

[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the human-computer interaction method proposed in this disclosure.

[0007] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing a computer to perform the human-computer interaction method proposed in this disclosure above.

[0008] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the resource recommendation method proposed in this disclosure.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0011] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;

[0012] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;

[0013] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure;

[0014] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure;

[0015] Figure 5 This is a schematic diagram according to the fifth embodiment of the present disclosure;

[0016] Figure 6 This is a schematic block diagram of an electronic device 600 used to implement embodiments of the present disclosure. Detailed Implementation

[0017] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0018] Figure 1 The diagram is based on the first embodiment of this disclosure. It should be noted that the human-computer interaction method of this disclosure can be applied to a human-computer interaction device, which can be configured in an intelligent system, and the intelligent system can be configured in an electronic device.

[0019] Among them, electronic devices can be any device with computing capabilities, such as personal computers (PCs), mobile terminals, servers, etc. Mobile terminals can be, for example, in-vehicle devices, mobile phones, tablets, personal digital assistants, wearable devices, smart speakers, servers, server clusters, and other hardware devices with various operating systems, touch screens and / or displays.

[0020] It should be noted that in the following embodiments, the execution subject is an electronic device as an example for illustrative purposes.

[0021] like Figure 1 As shown, the human-computer interaction method may include the following steps:

[0022] Step 101: Obtain the current video frame in the video stream and obtain the sequence of historical video frames preceding the current video frame.

[0023] In some embodiments, after the electronic device acquires the current video frame in the video stream, it can acquire multiple historical video frames that are consecutive within a preset time period before the current video frame, and sort the multiple historical video frames in order from front to back according to the acquisition time to obtain a historical video frame sequence.

[0024] The preset duration is a duration set in advance according to actual needs. For example, the preset duration can be 5 seconds, 6 seconds, 9 seconds, etc. This embodiment does not specifically limit the value of the preset duration.

[0025] Step 102: Based on the current video frame and the historical video frame sequence, determine whether an active dialogue needs to be initiated on the current video frame.

[0026] In this embodiment, it can be determined whether an active dialogue needs to be initiated on the current video frame based on the current video frame and the historical video frame sequence. If it is determined that an active dialogue needs to be initiated on the current video frame, step 103 is executed. That is, it is determined whether to speak actively on the current video frame based on the current video frame and the historical video frame sequence, and if it is determined that to speak actively on the current video frame, step 103 is executed.

[0027] Step 103: Using a large model, generate proactive dialogue content based on the current video frame and the historical video frame sequence.

[0028] In some embodiments, the current video frame and a sequence of historical video frames can be input into a large model to generate proactive dialogue content.

[0029] Step 104: Output the content of the proactive dialogue.

[0030] In some embodiments, the active dialogue content may be displayed on the interactive interface, and / or the active dialogue content may be output by voice. This embodiment does not specifically limit the way the active dialogue content is output.

[0031] For example, in a scenario where you're teaching someone to draw, based on the current video frame and a sequence of historical video frames, you might determine that the user needs to speak actively on the current video frame. Correspondingly, the current video frame and the sequence of historical video frames can be input into a large model. The large model can then analyze the current video frame and the sequence of historical video frames and generate active dialogue content based on the analysis results. Suppose the analysis result indicates that the user has finished drawing the house, the corresponding active dialogue content could be "Now it's time to draw the tree."

[0032] For example, in a video surveillance scenario, if it is determined that the user needs to speak on the current video frame based on the current video frame and the historical video frame sequence, the current video frame and the historical video frame sequence can be input into a large model. Correspondingly, the large model can analyze the current video frame and the historical frame sequence and generate proactive dialogue content based on the analysis results. Suppose the analysis result indicates that the user walks to the elevator entrance, and the elevator corresponding to the elevator entrance is under maintenance. Correspondingly, the generated proactive dialogue content could be "The elevator is under maintenance".

[0033] The human-computer interaction method provided in this disclosure, based on the current video frame and the historical video frame sequence in the video stream, determines that an active dialogue needs to be initiated on the current video frame. Then, using a large model, it generates and outputs the active dialogue content based on the current video frame and the historical video frame sequence. Therefore, by actively initiating dialogue during human-computer interaction, it improves the recognition and execution of interactive operations that match the user's intent, reduces the complexity of user operations, enhances the interactivity of human-computer interaction, effectively improves the intelligence level of the interaction, and enhances the user experience.

[0034] In some embodiments, to quickly and easily determine whether an active dialogue needs to be initiated in the current video frame, a trained classification model can be used, combined with the current video frame and a sequence of historical video frames, to determine whether an active dialogue needs to be initiated in the current video frame. The following section will combine... Figure 2 The process is described exemplarily.

[0035] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure.

[0036] like Figure 2 As shown, this human-computer interaction method includes the following steps:

[0037] Step 201: Obtain the current video frame in the video stream and obtain the sequence of historical video frames preceding the current video frame.

[0038] In some embodiments, to quickly obtain the sequence of historical video frames preceding the current video frame, a buffer corresponding to the video stream can be acquired, and multiple historical video frames with consecutive times preceding the current video frame can be retrieved from the buffer. These historical video frames are then sorted according to their acquisition time from front to back to obtain the historical video frame sequence. Thus, by using the buffer corresponding to the video stream, historical video frames preceding the current video frame can be obtained conveniently and quickly without reloading or processing the entire video stream, thereby saving computational resources and time costs.

[0039] In some embodiments, to facilitate the subsequent processing of video frames following the current video frame in the video stream, the current video frame can be easily and quickly obtained, and correspondingly, the current video frame can be stored in the buffer.

[0040] In some embodiments, when the storage space of the buffer corresponding to the video stream is limited, in order to store the current video frame completely in the buffer, one possible implementation is to store the current video frame in the buffer only when the remaining storage space in the buffer is greater than a storage space threshold. Therefore, storing the current video frame to be cached in the buffer when there is sufficient remaining space avoids data loss due to insufficient remaining space in the buffer, ensuring the integrity of the current video frame stored in the buffer.

[0041] In other embodiments, one possible way to store the current video frame in the buffer is as follows: if the remaining storage space in the buffer is less than or equal to a storage space threshold, delete the oldest historical video frame stored in the buffer and store the current video frame in the buffer. Thus, when the remaining storage space in the buffer is insufficient, deleting the oldest historical video frame ensures that there is sufficient storage space in the buffer to store the current video frame, guaranteeing the integrity of the current video frame stored in the buffer.

[0042] In some embodiments, in order to save storage space in the cache while reducing the impact on the generation of proactive dialogue, correspondingly, when the number of video frames in the cache exceeds a preset threshold, the earliest stored historical video frames in the cache are merged into a video summary frame.

[0043] In some embodiments, the earliest stored historical video frames in the buffer can be merged to obtain a single video summary frame. This saves buffer storage space while preserving the key content of the earliest stored historical video frames.

[0044] Step 202: Input the current video frame and the sequence of historical video frames into the trained classification model to obtain the classification result output by the classification model.

[0045] It should be noted that the above classification model was trained on the initial classification model based on the sample videos and the sample video frames marked as initiating active dialogue.

[0046] In some embodiments, to further improve the accuracy of the classification results, the first question input by the user regarding the video stream can also be obtained. Correspondingly, the first question, the current video frame, and the sequence of historical video frames are input into a trained classification model to obtain the classification result. Thus, by inputting the first question input by the user regarding the video stream into the classification model, the model can perform contextual understanding based on the first question, the current video frame, and the visual content in the sequence of historical video frames, thereby accurately classifying whether an active dialogue needs to be initiated in the current video frame, improving the accuracy of the classification results obtained by the model.

[0047] In some embodiments, the classification model described above may be obtained by training an initial classification model based on sample videos, a first question for the sample videos, and sample video frames in the sample videos that are marked as requiring active dialogue.

[0048] Step 203: If the classification result indicates that an active dialogue needs to be initiated on the current video frame, then determine that an active dialogue needs to be initiated on the current video frame.

[0049] In some embodiments, 0 and 1 can be used to represent the classification result. Correspondingly, if the classification result is 1, it means that an active dialogue needs to be initiated on the current video frame. Correspondingly, if the classification result is 0, it means that an active dialogue does not need to be initiated on the current video frame.

[0050] Step 204: Using a large model, generate proactive dialogue content based on the current video frame and the historical video frame sequence.

[0051] In some embodiments, current video frames and historical video frame sequences can be input into a large model to generate proactive dialogue content.

[0052] In some embodiments, corresponding prompts can be generated based on the current video frame and a sequence of historical video frames. These prompts instruct the generation of proactive dialogue content corresponding to an active conversation with the user on the current video frame, based on the current and historical video frame sequences. The prompts are then input into a large model to obtain the proactive dialogue content generated by the model. Thus, by guiding the large model with prompts, it is possible to accurately understand the background of the target and accurately generate proactive dialogue content.

[0053] In some embodiments, a possible implementation of generating proactive dialogue content using a large model based on the current video frame and a sequence of historical video frames is as follows: A first visual vector sequence of the current video frame and a second visual vector sequence of historical video frames are concatenated to obtain a third visual vector sequence; this third visual vector sequence is then input into the large model to obtain the proactive dialogue content. Therefore, by inputting the visual vector sequence of the video frames into the large model, the amount of input data is reduced, thereby improving the processing speed and efficiency of the large model and contributing to the efficiency of obtaining proactive dialogue content.

[0054] In some embodiments, the current video frame can be divided into blocks to obtain an image block sequence of the current video frame, and the image blocks in the image block sequence of the current video frame can be represented by vectors to obtain a first visual vector sequence. That is, the first visual vector sequence is obtained by representing the image block sequence of the current video frame by vectors.

[0055] In other embodiments, for any historical video frame in the historical video frame sequence, the historical video frame can be divided into blocks to obtain a vector representation of the image block sequence of the historical video frame, thus obtaining a second visual vector sequence of the historical video frame. That is, the second visual vector sequence is obtained by vector representation of the image block sequence of the historical video frame.

[0056] In some embodiments, one possible way to input a third-visual vector sequence into a large model to obtain active dialogue content is to input the third-visual vector sequence into the large model when its length is less than or equal to the maximum input length of the large model, thereby obtaining the active dialogue content. This allows the large model to accurately understand and process third-visual vector sequences shorter than the maximum input length, effectively avoiding truncation and information loss caused by exceeding the input length limit, thus improving the accuracy of the active dialogue content obtained by the large model.

[0057] The maximum input length refers to the maximum length of input data that a large model can process.

[0058] In some embodiments, another possible implementation of inputting the third visual vector sequence into a large model to obtain active dialogue content is as follows: If the length of the third visual vector sequence is greater than the maximum input length of the large model, the visual vector sequences of the first historical video frame in the third visual vector sequence are merged and compressed to obtain a fourth visual vector sequence. The length of the fourth visual vector sequence is less than or equal to the maximum input length, and the first historical video frame is at least one historical video frame from the historical video frame sequence with an earlier acquisition time. The fourth visual vector sequence is then input into the large model to obtain the active dialogue content. Therefore, when the length of the third visual vector sequence is determined to be greater than the maximum input length of the large model, by merging and compressing the third visual vector sequence, the processed fourth visual vector sequence is made smaller than the maximum input length of the large model. This allows the large model to accurately understand and process the fourth visual vector sequence, avoiding the possibility of the large model truncating the input data, which could lead to errors in the large model's understanding and reasoning of the context, thus improving the accuracy of the active dialogue content generated by the large model.

[0059] Step 205: Output the content of the proactive dialogue.

[0060] In this embodiment, after acquiring the current video frame and the sequence of historical video frames preceding it in the video stream, a trained classification model quickly and conveniently determines whether an active dialogue needs to be initiated on the current video frame based on the current frame and the historical video frame sequence. If it is determined that an active dialogue needs to be initiated on the current video frame, the large model generates and outputs the active dialogue content based on the current video frame and the historical video frame sequence. Thus, by quickly and conveniently determining whether an active dialogue will be triggered on the current video frame through a classification model, and then proactively initiating the dialogue after it has been triggered, the level of intelligence in the interaction is effectively improved.

[0061] Based on the above embodiments, in order to facilitate the subsequent processing of video frames after the current video frame in the video stream, the active dialogue content associated with the current video frame can be easily and quickly obtained from the buffer area, and the current video frame and the active dialogue content can be stored together in the buffer area.

[0062] In some embodiments, during interaction with the user based on the video stream, the user may ask at least one round of questions on historical video frames preceding the current video frame. Correspondingly, the electronic device provides answers to the corresponding rounds of questions. Therefore, when it is determined that an active dialogue needs to be initiated on the current video frame, in order to accurately determine the content of the active dialogue, a second historical video frame associated with historical dialogue text in the historical video frame sequence can be identified, and the historical dialogue text associated with the second historical video frame can be obtained. Furthermore, the active dialogue content is generated using a large model based on the current video frame, the historical video frame sequence, and the historical dialogue text. To clearly understand this process, the following section combines... Figure 3 The process is described further with examples.

[0063] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure.

[0064] like Figure 3 As shown, the method may include:

[0065] Step 301: Obtain the current video frame in the video stream and obtain the sequence of historical video frames preceding the current video frame.

[0066] Step 302: Based on the current video frame and the historical video frame sequence, determine whether an active dialogue needs to be initiated on the current video frame.

[0067] It should be noted that for a detailed description of steps 301 and 302, please refer to the relevant descriptions in other embodiments, which will not be repeated here.

[0068] Step 303: Determine the second historical video frame in the historical video frame sequence that is associated with historical dialogue text, and obtain the historical dialogue text associated with the second historical video frame.

[0069] In some embodiments, a second historical video frame associated with historical dialogue text may be determined from a buffer corresponding to the video stream, and the historical dialogue text associated with the second historical video frame may be obtained from the buffer.

[0070] The historical dialogue text includes the question and the answer to the question. The answer to the question is generated by the large language model based on the question, the historical video frame, and the sequence of historical video frames preceding the historical video frame.

[0071] Step 304: Input the current video frame, historical video frame sequence, and historical dialogue text into the large model to generate proactive dialogue content.

[0072] In some embodiments, prompt words can be generated based on the current video frame, a sequence of historical video frames, and historical dialogue text. These prompt words are used to generate the proactive dialogue content required to initiate a dialogue on the current video frame, based on the current video frame, the sequence of historical video frames, and the historical dialogue text. Correspondingly, the prompt words are input into a large model to obtain the proactive dialogue content output by the large model. Thus, by guiding the large model with prompt words, the large model can accurately understand the background of the target and accurately generate proactive dialogue content.

[0073] In other embodiments, the first visual vector sequence of the current video frame, the second visual vector sequence of historical video frames in the historical video frame sequence, and the text vector sequence of historical dialogue text can be concatenated to obtain a first concatenated vector sequence. This first concatenated vector sequence is then input into a large model to obtain the active dialogue content. Therefore, by concatenating the vector sequences of text and video frames and inputting the resulting first concatenated vector sequence into the large model, the amount of input data to the large model can be reduced, thereby improving the data processing speed and efficiency of the large model and contributing to improving the efficiency of obtaining active dialogue content.

[0074] In some embodiments, the current video frame can be divided into blocks to obtain an image block sequence of the current video frame, and the image blocks in the image block sequence of the current video frame can be represented by vectors to obtain a first visual vector sequence. That is, the first visual vector sequence is obtained by representing the image block sequence of the current video frame by vectors.

[0075] In other embodiments, for any historical video frame in the historical video frame sequence, the historical video frame can be divided into blocks to obtain a vector representation of the image block sequence of the historical video frame, thus obtaining a second visual vector sequence of the historical video frame. That is, the second visual vector sequence is obtained by vector representation of the image block sequence of the historical video frame.

[0076] In some embodiments, the historical dialogue text can be segmented to obtain a segmented sequence of the historical dialogue text, and the segmented sequence can be vectorized to obtain a text vector sequence of the historical dialogue text.

[0077] In some embodiments, one possible way to input the first concatenated vector sequence into the large model to obtain the active dialogue content is as follows: if the length of the first concatenated vector sequence is less than or equal to the maximum input length of the large model, then the first concatenated vector sequence is input into the large model to obtain the active dialogue content. This allows the large model to accurately understand and process the first concatenated vector sequence whose length is less than the maximum input length, effectively avoiding truncation and information loss problems caused by exceeding the input length limit, thereby improving the accuracy of the active dialogue content obtained by the large model.

[0078] In other embodiments, another possible way to input the first concatenated vector sequence into the large model to obtain the active dialogue content is as follows: If the length of the first concatenated vector sequence is greater than the maximum input length of the large model, the visual vector sequence of the third historical video frame in the first concatenated vector sequence is merged and compressed to obtain a second concatenated vector sequence. The length of the second concatenated vector sequence is less than or equal to the maximum input length, and the third historical video frame is at least one historical video frame from the historical video frame sequence with an earlier acquisition time. The second concatenated vector sequence is then input into the large model to obtain the active dialogue content. Therefore, when the length of the first concatenated vector sequence is determined to be greater than the maximum input length of the large model, by merging and compressing the first concatenated vector sequence, the processed second concatenated vector sequence is made smaller than the maximum input length of the large model. This allows the large model to accurately understand and process the second concatenated vector sequence, avoiding the situation where the input data is truncated by the large model, which could lead to errors in the large model's understanding and reasoning of the context, thus improving the accuracy of the active dialogue content generated by the large model.

[0079] Step 305: Output the content of the proactive dialogue.

[0080] In this embodiment, when it is determined that a dialogue needs to be initiated on the current video frame, a second historical video frame associated with historical dialogue text is identified in the historical video frame sequence, and the historical dialogue text associated with the second historical video frame is obtained. The current video frame, the historical video frame sequence, and the historical dialogue text are then input into the large model to generate proactive dialogue content. This allows the large model to better understand the context of the proactive dialogue content to be generated based on the input historical dialogue text, which helps to improve the accuracy of the proactive dialogue content generated by the large model.

[0081] Based on any of the above embodiments, the target question input by the user on the current video frame can also be obtained; the historical video frame sequence, the current video frame, and the target question are input into a large model, and the answer information for the target question and the target historical video frame on which the answer information is based are obtained through the large model. Thus, the large model can accurately obtain the answer to the target question input by the user on the current video frame, satisfying the user's need to actively ask questions on the current video frame.

[0082] For example, in a video surveillance scenario, when a user views the current video frame, they might say, "Did I just pick up a package?" This "Did I just pick up a package?" becomes the target question the user inputs on the current video frame. Correspondingly, the historical video frame sequence, the current video frame, and the target question can be input into a large model. The model then obtains the answer to the target question and the target historical video frame upon which the answer is based. Assuming the answer is: "You just picked up a package," and providing the target historical video frame upon which this answer is based, it can be understood that the video content in this target historical video frame indicates that the user has already picked up the package.

[0083] In some embodiments, in order to facilitate the subsequent processing of video frames after the current video frame in the video stream, the target question associated with the current video frame and the answer information of the target question can be obtained conveniently and quickly. Correspondingly, the current video frame, the target question, and the corresponding answer information can also be associated and stored in the buffer area corresponding to the video stream.

[0084] To facilitate a clear understanding of this disclosure, the following will be combined with... Figure 4 The method of this embodiment will be further described exemplarily.

[0085] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure.

[0086] like Figure 4 As shown, the method may include:

[0087] Step 401: Obtain the current video frame in the video stream.

[0088] Step 402: Obtain the sequence of historical video frames preceding the current video frame from the buffer corresponding to the video stream.

[0089] Step 403: Using a trained classification model, determine whether an active dialogue needs to be initiated on the current video frame based on the current video frame and the historical video frame sequence.

[0090] In some embodiments, the user's first question input to the video stream can also be obtained, and correspondingly, a trained classification model can be used to determine whether an active dialogue needs to be initiated on the current video frame based on the first question, the current video frame, and the historical video frame sequence.

[0091] The trained classification model determines the specific description of the active dialogue to be initiated on the current video frame based on the first question, the current video frame, and the historical video frame sequence. See the relevant descriptions in other embodiments for details, which will not be repeated here.

[0092] Step 404: For a historical video frame in the historical video frame sequence, if it is determined that there is historical dialogue text associated with the historical video frame in the buffer, obtain the historical dialogue text associated with the historical video frame.

[0093] Step 405: Concatenate the first visual vector sequence of the current video frame, the second visual vector sequence of the historical video frames in the historical video frame sequence, and the text vector sequence of the historical dialogue text to obtain the first concatenated vector sequence.

[0094] Step 406: Input the first concatenated vector sequence into the large model to obtain the active dialogue content.

[0095] For a detailed description of steps 405 and 406, please refer to the relevant descriptions in other embodiments, which will not be repeated here.

[0096] Step 407: Output the content of the proactive dialogue.

[0097] In this embodiment, when the classification model determines that an active dialogue needs to be initiated on the current video frame based on the current video frame and the historical video frame sequence in the video stream, the large model generates and outputs the active dialogue content based on the current video frame, historical dialogue text, and the historical video frame sequence. Thus, initiating active dialogue during human-computer interaction improves the recognition and execution of interactive operations that match the user's intent, reduces the complexity of user operations, enhances the interactivity of human-computer interaction, effectively improves the intelligence level of the interaction, and enhances the user experience.

[0098] To implement the above embodiments, this disclosure also provides a human-computer interaction device.

[0099] Figure 5 This is a schematic diagram according to the fifth embodiment of the present disclosure.

[0100] like Figure 5 As shown, the human-computer interaction device 50 may include: an acquisition module 501, a determination module 502, a generation module 503, and an output module 504, wherein:

[0101] The acquisition module 501 is used to acquire the current video frame in the video stream and to acquire the sequence of historical video frames preceding the current video frame.

[0102] The determination module 502 is used to determine, based on the current video frame and the historical video frame sequence, whether an active dialogue needs to be initiated on the current video frame.

[0103] The generation module 503 is used to generate proactive dialogue content based on the current video frame and the sequence of historical video frames using a large model.

[0104] Output module 504 is used to output the content of the active dialogue.

[0105] As one possible implementation of this disclosure, the determining module 502 includes:

[0106] The acquisition unit is used to input the current video frame and the sequence of historical video frames into the trained classification model and obtain the classification result output by the classification model.

[0107] The determination unit is used to determine whether an active dialogue needs to be initiated in the current video frame when the classification result indicates that an active dialogue needs to be initiated in the current video frame.

[0108] As one possible implementation of this disclosure, the apparatus further includes:

[0109] The question retrieval module is used to retrieve the first question the user inputs in response to the video stream;

[0110] The acquisition unit is specifically used to: input the first question, the current video frame, and the sequence of historical video frames into the trained classification model, and obtain the classification result output by the classification model.

[0111] As one possible implementation of this disclosure, the generation module 503 includes:

[0112] The splicing unit is used to splice the first visual vector sequence of the current video frame and the second visual vector sequence of the historical video frames in the historical video frame sequence to obtain the third visual vector sequence.

[0113] The first generation unit is used to input the third visual vector sequence into the large model to obtain the active dialogue content.

[0114] As one possible implementation of this disclosure, the first generation unit is specifically used to: input the third visual vector sequence into the large model to obtain active dialogue content when the length of the third visual vector sequence is less than or the maximum input length of the large model.

[0115] As one possible implementation of this disclosure, the generation unit is specifically used for: when the length of the third visual vector sequence is greater than the maximum input length of the large model, merging and compressing the visual vector sequences of the first historical video frames in the third visual vector sequence to obtain a fourth visual vector sequence, wherein the length of the fourth visual vector sequence is less than or equal to the maximum input length, and wherein the first historical video frame is at least one historical video frame with an earlier acquisition time in the historical video frame sequence; and inputting the fourth visual vector sequence into the large model to obtain active dialogue content.

[0116] As one possible implementation of this disclosure, the apparatus may further include:

[0117] The first processing module is used to determine the second historical video frame in the historical video frame sequence that is associated with historical dialogue text, and to obtain the historical dialogue text associated with the second historical video frame.

[0118] The generation module 503 includes:

[0119] The second generation unit is used to input the current video frame, historical video frame sequence, and historical dialogue text into the large model to generate proactive dialogue content.

[0120] As one possible implementation of this disclosure, the second generation unit includes:

[0121] The splicing subunit is used to splice the first visual vector sequence of the current video frame, the second visual vector sequence of the historical video frames in the historical video frame sequence, and the text vector sequence of the historical dialogue text to obtain the first spliced ​​vector sequence.

[0122] The generated sub-unit is used to input the first concatenated vector sequence into the large model to obtain the active dialogue content.

[0123] As one possible implementation of this disclosure, the generation subunit is specifically used for:

[0124] If the length of the first concatenated vector sequence is less than or equal to the maximum input length of the large model, the first concatenated vector sequence is input into the large model to obtain the active dialogue content.

[0125] As one possible implementation of this disclosure, the generation subunit is specifically used for: when the length of the first spliced ​​vector sequence is greater than the maximum input length of the large model, merging and compressing the visual vector sequence of the third historical video frame in the first spliced ​​vector sequence to obtain a second spliced ​​vector sequence, wherein the length of the second spliced ​​vector sequence is less than or equal to the maximum input length, and wherein the third historical video frame is at least one historical video frame with an earlier acquisition time in the historical video frame sequence; inputting the second spliced ​​vector sequence into the large model to obtain active dialogue content.

[0126] As one possible implementation of this disclosure, the apparatus may further include:

[0127] The second processing module is used to obtain the target question input by the user on the current video frame; input the historical video frame sequence, the current video frame and the target question into the large model, and obtain the answer information of the target question and the target historical video frame on which the answer information is based through the large model.

[0128] As one possible implementation of this disclosure, the acquisition module 501 is specifically used for: acquiring a buffer corresponding to the video stream; acquiring multiple historical video frames that are consecutive in time before the current video frame from the buffer; and sorting the multiple historical video frames in order from front to back according to the acquisition time to obtain a historical video frame sequence.

[0129] As one possible implementation of this disclosure, the device further includes a storage module for storing the current video frame in a buffer.

[0130] As one possible implementation of this disclosure, the storage module is specifically used to: store the current video frame in the buffer when the remaining storage space in the buffer is greater than the storage space threshold.

[0131] As one possible implementation of this disclosure, the storage module is specifically used to: delete the oldest historical video frame stored in the cache and store the current video frame in the cache when the remaining storage space in the cache is less than or equal to the storage space threshold.

[0132] As one possible implementation of this disclosure, the device may further include: a merging module, used to merge multiple historical video frames stored earliest in the buffer into a video summary frame when the number of video frames in the buffer exceeds a preset threshold.

[0133] As one possible implementation of this disclosure, the device may further include: an associated storage module, used to associate and store the current video frame and the active dialogue content in a buffer.

[0134] It should be noted that the foregoing explanation of the human-computer interaction method also applies to the human-computer interaction device of this embodiment, and will not be repeated here.

[0135] The human-computer interaction device of this disclosure, based on the current video frame in the video stream and the historical video frame sequence above the current video frame, determines that an active dialogue needs to be initiated on the current video frame. Then, using a large model, it generates and outputs the active dialogue content based on the current video frame and the historical video frame sequence. Thus, during human-computer interaction, proactive dialogue is initiated, improving the recognition and execution of interactive operations that match the user's intent, reducing the complexity of user operations, enhancing the interactivity of human-computer interaction, effectively improving the intelligence level of the interaction, and improving the user experience.

[0136] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision, and disclosure of users' personal information are all carried out with the consent of the users, and all comply with the provisions of relevant laws and regulations, and do not violate public order and good morals.

[0137] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0138] Figure 6 This is a schematic block diagram of an electronic device 600 for implementing embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0139] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0140] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0141] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as human-computer interaction methods. For example, in some embodiments, the human-computer interaction method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the human-computer interaction method or the training method of the resource representation generation model described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform human-computer interaction methods by any other suitable means (e.g., by means of firmware).

[0142] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0143] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0144] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0145] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0146] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0147] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0148] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0149] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A human-computer interaction method, comprising: Obtain the current video frame in the video stream, and obtain the sequence of historical video frames preceding the current video frame; Based on the current video frame and the historical video frame sequence, it is determined that an active dialogue needs to be initiated on the current video frame; Using a large model, proactive dialogue content is generated based on the current video frame and the historical video frame sequence; Output the content of the proactive dialogue.

2. The method according to claim 1, wherein, The step of determining whether an active dialogue needs to be initiated on the current video frame based on the current video frame and the historical video frame sequence includes: Input the current video frame and the sequence of historical video frames into a trained classification model to obtain the classification result output by the classification model; If the classification result indicates that an active dialogue needs to be initiated on the current video frame, then it is determined that an active dialogue needs to be initiated on the current video frame.

3. The method according to claim 2, wherein, The method further includes: Obtain the first question input by the user regarding the video stream; The step of inputting the current video frame and the historical video frame sequence into a trained classification model to obtain the classification result output by the classification model includes: The first question, the current video frame, and the sequence of historical video frames are input into a trained classification model to obtain the classification result output by the classification model.

4. The method according to claim 1, wherein, The process of generating proactive dialogue content using a large model, based on the current video frame and the historical video frame sequence, includes: The first visual vector sequence of the current video frame and the second visual vector sequence of the historical video frames in the historical video frame sequence are concatenated to obtain the third visual vector sequence. The third visual vector sequence is input into the large model to obtain the active dialogue content.

5. The method according to claim 4, wherein, The step of inputting the third visual vector sequence into the large model to obtain the active dialogue content includes: If the length of the third visual vector sequence is less than or equal to the maximum input length of the large model, the third visual vector sequence is input into the large model to obtain the active dialogue content.

6. The method according to claim 4, wherein, The step of inputting the third visual vector sequence into the large model to obtain the active dialogue content includes: When the length of the third visual vector sequence is greater than the maximum input length of the large model, the visual vector sequences of the first historical video frame in the third visual vector sequence are merged and compressed to obtain a fourth visual vector sequence, wherein the length of the fourth visual vector sequence is less than or equal to the maximum input length, and wherein the first historical video frame is at least one historical video frame with an earlier acquisition time in the historical video frame sequence. The fourth visual vector sequence is input into the large model to obtain the active dialogue content.

7. The method according to claim 1, wherein, The method further includes: Identify the second historical video frame in the historical video frame sequence that is associated with historical dialogue text, and obtain the historical dialogue text associated with the second historical video frame; The step of generating proactive dialogue content using a large model based on the current video frame and the historical video frame sequence includes: The current video frame, the historical video frame sequence, and the historical dialogue text are input into the large model to generate proactive dialogue content.

8. The method according to claim 7, wherein, The step of inputting the current video frame, the historical video frame sequence, and the historical dialogue text into the large model to generate proactive dialogue content includes: The first visual vector sequence of the current video frame, the second visual vector sequence of the historical video frames in the historical video frame sequence, and the text vector sequence of the historical dialogue text are concatenated to obtain the first concatenated vector sequence. The first concatenated vector sequence is input into the large model to obtain the active dialogue content.

9. The method according to claim 8, wherein, The step of inputting the first concatenated vector sequence into the large model to obtain the active dialogue content includes: If the length of the first concatenated vector sequence is less than or equal to the maximum input length of the large model, the first concatenated vector sequence is input into the large model to obtain the active dialogue content.

10. The method according to claim 8, wherein, The step of inputting the first concatenated vector sequence into the large model to obtain the active dialogue content includes: If the length of the first spliced ​​vector sequence is greater than the maximum input length of the large model, the visual vector sequence of the third historical video frame in the first spliced ​​vector sequence is merged and compressed to obtain a second spliced ​​vector sequence, wherein the length of the second spliced ​​vector sequence is less than or equal to the maximum input length, and the third historical video frame is at least one historical video frame with an earlier acquisition time in the historical video frame sequence. The second concatenated vector sequence is input into the large model to obtain the active dialogue content.

11. The method according to claim 1, wherein, The method further includes: Obtain the target question input by the user on the current video frame; The historical video frame sequence, the current video frame, and the target question are input into the large model, and the answer information for the target question and the target historical video frame on which the answer information is based are obtained through the large model.

12. The method according to any one of claims 1-11, wherein, The step of obtaining the historical video frame sequence preceding the current video frame includes: Obtain the buffer corresponding to the video stream; Obtain multiple consecutive historical video frames preceding the current video frame from the buffer. The historical video frames are sorted in chronological order of acquisition time to obtain the historical video frame sequence.

13. The method according to claim 12, wherein, The method further includes: The current video frame is stored in the buffer.

14. The method according to claim 13, wherein, The step of storing the current video frame into the buffer includes: If the remaining storage space in the buffer is greater than the storage space threshold, the current video frame is stored in the buffer.

15. The method according to claim 13, wherein, The step of storing the current video frame into the buffer includes: If the remaining storage space in the buffer is less than or equal to the storage space threshold, the earliest historical video frame stored in the buffer is deleted, and the current video frame is stored in the buffer.

16. The method according to claim 12, wherein, The method further includes: If the number of video frames in the buffer exceeds a preset threshold, the earliest stored historical video frames in the buffer will be merged into a video summary frame.

17. The method according to claim 13, wherein, The method further includes: The current video frame and the active dialogue content are stored together in the cache area.

18. A human-computer interaction device, comprising: The acquisition module is used to acquire the current video frame in the video stream and acquire the sequence of historical video frames preceding the current video frame; The determining module is used to determine, based on the current video frame and the historical video frame sequence, whether an active dialogue needs to be initiated on the current video frame. The generation module is used to generate proactive dialogue content based on the current video frame and the historical video frame sequence using a large model; The output module is used to output the content of the active dialogue.

19. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-17.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-17.

21. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-17.

Citation Information

Patent Citations

  • Active dialogue generation method and device, chat robot and storage medium

    CN117763118A

  • Active dialogue robot system based on large language model

    CN119719298A

  • Active interaction method, electronic device and readable storage medium

    US20220019847A1

  • Video question answering method, electronic device and storage medium

    US20230121838A1