Interaction method, system and related device
By recognizing the image content captured by the camera through electronic devices or cloud servers, the intelligent assistant can proactively interact, solving the problem of users needing to actively initiate dialogue and improving the user experience and the diversity of application scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-06-09
- Publication Date
- 2026-05-15
AI Technical Summary
Smart assistants require users to initiate conversations, lacking proactive interaction capabilities, resulting in a poor user experience.
Electronic devices enter an unintentional, proactive interaction mode, capturing images through cameras, recognizing image content, and initiating chat conversations, or processing image data through cloud servers to provide interactive data.
It enhances user engagement with the smart assistant, improves the user experience, and increases the diversity and intelligence of application scenarios.
Smart Images

Figure CN122053547A_ABST
Abstract
Description
[0001] This application is a divisional application. The original application has the application number 202510770028.1 and the original application date is June 9, 2025. The entire contents of the original application are incorporated herein by reference. Technical Field
[0002] This application relates to the field of terminals, and more particularly to an interaction method, system and related device. Background Technology
[0003] Today, with the development of multimodal large language models (MLLMs) using deep neural networks and the generation of massive amounts of text and image data, powerful multimodal intelligent assistants have emerged. These assistants can receive dialogue data such as text, voice, or images from the user and generate corresponding dialogue responses based on that data, thus enabling them to converse with the user. However, intelligent assistants require the user to actively initiate the dialogue and passively respond to the user's input. If users don't know what to talk about with the intelligent assistant, it will reduce their willingness to use it, resulting in a poor user experience. Summary of the Invention
[0004] This application provides an interaction method, system, and related apparatus that achieves the effect of actively outputting interactive data based on a video stream.
[0005] Firstly, this application provides an interaction method applied to an electronic device. The method includes: triggering the electronic device's chat application to enter a no-intention proactive interaction mode based on a user's first operation on the electronic device or pre-set settings on the electronic device. The no-intention proactive interaction mode is an interaction mode where the chat application proactively initiates a chat conversation with the user on a specific topic. In the no-intention proactive interaction mode, an image captured by the electronic device's camera is acquired. Based on the recognition of a first content in the image, a chat conversation is initiated with the user regarding that first content. Thus, when the electronic device is in the no-intention proactive interaction mode, it can proactively recognize the image captured by the camera and, based on the recognition of the first content in the image, proactively initiate a chat conversation with the user regarding that first content. Even when the user does not know what to input, it can proactively chat with the user, increasing the user's interest in using the chat application and providing a better user experience.
[0006] In one possible implementation, the image comes from real-time footage captured by a camera, or from pre-generated video or photo files. This allows electronic devices to initiate conversations with users based on real-time camera footage or images from video / photo files, leading to more diverse application scenarios.
[0007] In one possible implementation, the first operation is the user's activation of a first function, which has an unintentional, proactive interaction mode. In this way, the user can decide whether to activate the first function. When the user needs to use the first function, the first operation triggers the electronic device to actively recognize an image and initiate a dialogue with the user.
[0008] In one possible implementation, a pre-set feature includes preventing the chat application from receiving user input for a predetermined period. This increases the likelihood that the user will not initiate a conversation if no data is input into the electronic device within the preset time, allowing the electronic device to proactively initiate a conversation and improving the user experience.
[0009] In one possible implementation, the method further includes displaying images on the electronic device's display interface in a non-intentional, proactive interaction mode. This allows the electronic device to proactively initiate dialogue with the user while displaying images, providing a more vivid and engaging conversational experience by combining visual content.
[0010] In one possible implementation, the display interface also shows at least one of the following: a first control for starting or exiting the unintentional active interaction mode; a second control for starting or stopping the voice input function; a third control for starting or stopping the voice broadcast function; a fourth control for starting or stopping the camera; and a fifth control for starting or stopping the text interaction function. In this way, the user can actively set the electronic device to start or stop the unintentional active interaction mode through the first control, and can stop the first function when the user does not want to use it. The user can start or stop the voice input function through the second control, enabling the voice input function only when voice input is desired, thus preventing the electronic device from receiving unnecessary voice data and saving power. The user can start or stop the voice broadcast function through the third control, which can be applied in scenarios where external sound is not suitable. The user can temporarily turn off the camera when they do not want to recognize the currently being photographed object through the fourth control. The user can turn off the text interaction function to avoid text obscuring the image.
[0011] In one possible implementation, the method further includes: in the unintentional proactive interaction mode, exiting the unintentional proactive interaction mode in response to the chat application receiving chat input from the user. In this way, the electronic device proactively exits the unintentional proactive interaction mode and replies to the user's chat input after receiving it. This satisfies both the user's need to initiate a conversation and the user's desire for the electronic device to initiate a conversation.
[0012] In one possible implementation, chat input is voice input, with the audio characteristics of the voice input matching the user's audio characteristics, or the sound source of the voice input is located behind the camera's view. This allows the electronic device to receive only the user's voice input, reducing interference from other sounds in determining whether the device exits the unintentional active interaction mode.
[0013] In one possible implementation, the image includes multiple objects, with the first being the largest object in the image, the most centrally located object, or the object closest to the user. This allows the electronic device to focus more on objects the user is more likely to look at.
[0014] In one possible implementation, the primary content includes emoticons, gestures, text, or QR codes. This allows electronic devices to recognize a variety of content, enriching the conversations with the user.
[0015] In one possible implementation, based on the recognition of first content in an image, a chat dialogue is initiated with the user regarding the first content. This includes: based on the recognition of the first content in the image, outputting a first dialogue about the first content to the user; the method further includes: in response to the first content existing in the image for a duration exceeding a threshold, outputting a second dialogue about the first content to the user, wherein the first description of the first content in the second dialogue is more detailed than the second description of the first content in the second dialogue. Thus, as the image continuously displays the first content, the likelihood of the user wanting to know more about the first content increases, and the electronic device outputs a more detailed second dialogue to meet the user's needs.
[0016] Secondly, this application provides another interaction method applied to a cloud server. The method includes: receiving an image captured by the camera of an electronic device; determining first interaction data based on the recognition of first content in the image, the first interaction data being used by the electronic device to initiate a chat conversation with a user regarding the first content; and sending the first interaction data to the electronic device. In this way, the cloud server can provide a service to multiple electronic devices to determine interaction data based on images. Even if the electronic device does not have the ability to recognize the first content, it can still actively initiate a chat conversation with the user through the cloud server, and the overhead of determining interaction data by the electronic device is also reduced.
[0017] Thirdly, this application provides an interaction method applied to a communication system including electronic devices and a cloud server. The method includes: the electronic device triggering its chat application to enter a no-intention proactive interaction mode based on a user's first operation on the device or pre-set settings on the device; the no-intention proactive interaction mode is an interaction mode where the chat application proactively initiates a chat conversation with the user on a specific topic; in the no-intention proactive interaction mode, the electronic device acquires an image captured by its camera; the electronic device sends the image to the cloud server; the cloud server determines first interaction data based on the recognition of first content in the image; the cloud server sends the first interaction data back to the electronic device; and the electronic device initiates a chat conversation with the user on the first content based on the received first interaction data. In this way, the electronic device can process the image captured by its camera through the cloud server, saving power consumption. The cloud server can provide a service to determine interaction data for multiple electronic devices, helping multiple electronic devices achieve proactive interaction functionality.
[0018] Fourthly, this application provides an interaction method applied to an electronic device. The method includes: receiving a first input; responding to the first input, displaying a first interface and sending a first video stream captured by a camera to a cloud server; wherein the first interface includes a first video stream, the first video stream includes first video frames, and the first video frames include a first subject being filmed; receiving first interactive data sent by the cloud server, the first interactive data including an introduction to the first subject being filmed; and outputting the first interactive data when the first video frame is displayed on the first interface. In this way, the electronic device can proactively output interactive data even without receiving dialogue data input from the user, providing the user with a more intelligent interactive experience.
[0019] In some examples, the first input can be input 1, input 2, input 3, or input 4.
[0020] In one possible implementation, the first video stream further includes a second video frame, which includes a second subject being filmed; the method further includes: receiving second interactive data sent by a cloud server, the second interactive data including an introduction to the second subject being filmed; and outputting the second interactive data when the second video frame is displayed on the first interface. In this way, after receiving the first input, the electronic device can continuously and proactively output interactive data based on the first video stream captured by the camera, without requiring user input of dialogue data.
[0021] In one possible implementation, before receiving the first interactive data sent by the cloud server, the method further includes: sending a first request to the cloud server, the first request instructing the cloud server to send interactive data. In this way, the electronic device proactively sends the first request, triggering the cloud service to send interactive data, which can save communication resources between the electronic device and the cloud server.
[0022] In one possible implementation, before receiving the second interactive data sent by the cloud server, the method further includes: sending a second request to the cloud server, the second request being used to instruct the cloud server to send interactive data.
[0023] In one possible implementation, the difference between the time the electronic device sends the second request to the cloud server and the time the electronic device sends the first request to the cloud server is a preset value. In this way, the electronic device can continuously acquire and output interactive data by sending requests at regular intervals.
[0024] In one possible implementation, the first video stream further includes a third video frame, which includes the first subject being captured; the method further includes: receiving third interactive data sent by a cloud server; wherein the third interactive data includes a detailed description of the first subject being captured, and the amount of data in the third interactive data is greater than the amount of data in the first interactive data; and outputting the third interactive data when the third video frame is displayed on the first interface.
[0025] In one possible implementation, all video frames between the third and first video frames include the first subject. Thus, when the subject in the video captured by the electronic device's camera remains unchanged, it typically indicates that the user prefers to learn more about the subject. The electronic device 100 can then acquire and output interactive data containing more information about the subject, providing the user with a more intelligent object recognition experience.
[0026] In one possible implementation, before receiving the first input, the method further includes: displaying a second interface; wherein the second interface includes a second video stream, and the second video stream includes a fourth video frame; while displaying the fourth video frame on the second interface, receiving the first dialogue data input by the user; and outputting first response data based on the first dialogue data and the fourth video frame. In this way, the electronic device can also respond to the first dialogue data input by the user after receiving it. It can both actively initiate a dialogue with the user and passively respond to a dialogue initiated by the user.
[0027] Fifthly, this application provides another interaction method applied to a cloud server; the method includes: receiving a first video stream sent by an electronic device; determining one or more trigger frames based on the first video stream; wherein the one or more trigger frames include the first trigger frame, and the first trigger frame includes a first subject being photographed; determining first interaction data based on the first trigger frame; wherein the first interaction data includes introductory content of the first subject being photographed; and sending the first interaction data to the electronic device. In this way, the cloud server can proactively identify trigger frames in the video stream provided by the electronic device and provide the electronic device with interaction data for the electronic device to initiate proactive interaction.
[0028] In one possible implementation, one or more trigger frames further include a second trigger frame, which includes a second subject being captured; the method further includes: determining second interactive data based on the second trigger frame; wherein the second interactive data includes introductory content of the second subject being captured; and sending the second interactive data to the electronic device. In this way, the cloud server can continuously identify trigger frames in the video stream of the electronic device and continuously provide the electronic device with interactive data that can be used to initiate proactive interactions.
[0029] In one possible implementation, determining the first interaction data based on the first trigger frame specifically includes: receiving a first request sent by an electronic device; responding to the first request, determining the first trigger frame from one or more trigger frames; wherein the acquisition time of the first trigger frame among the one or more trigger frames is closest to the time of receiving the first request; and determining the first interaction data based on the first trigger frame. In this way, the cloud server determines the interaction data only after receiving the request from the electronic device, which reduces the workload of the cloud server and the communication resources required for communication between the electronic device and the cloud server, making it easier for the cloud server to provide the service of determining interaction data to more electronic devices.
[0030] In one possible implementation, determining the second interaction data based on the second trigger frame specifically includes: receiving a second request sent by an electronic device; responding to the second request, determining a second trigger frame from one or more trigger frames; wherein the acquisition time of the second trigger frame among the one or more trigger frames is closest to the time of receiving the second request; and determining the second interaction data based on the second trigger frame. In this way, the cloud server can determine the second interaction data based on the more recent second trigger frame, increasing the likelihood that the electronic device will output the second interaction data when displaying video frames including the second captured object.
[0031] In one possible implementation, the subject in the trigger frame is the same as the subject in the previous m video frames of the trigger frame and different from the subject in the (m+1)th video frame before the trigger frame, and / or, the subject in the trigger frame is different from the subject in the previous (m+1)th video frames of the trigger frame and the same as the subject in the previous (m+1)th video frames of the trigger frame.
[0032] In one possible implementation, the method further includes: determining a segmentation mask and a confidence score for video frames in the first video stream based on the first video stream; wherein the segmentation mask can be used to indicate the location of the subject in the video frame, and the confidence score can be used to represent the probability that the subject exists in the video frame, and the first video stream includes a fifth video frame; if the confidence score of the fifth video frame is greater than a first score threshold, determining the subject of the fifth video frame based on the segmentation mask of the fifth video frame; if the subject of the fifth video frame is the same as the subject in the previous m video frames and different from the subject in the (m+1)th video frame before the fifth video frame, and / or, the subject in the fifth video frame is different from the subjects in the previous (m+1)th video frames and the subject in the (m+1)th video frames before the trigger frame, determining the fifth video frame as the trigger frame.
[0033] In one possible implementation, the first interaction data is determined based on the first trigger frame, specifically including: determining first image information for describing the first photographed object based on the first trigger frame; and generating the first interaction data based on the first image information, the first trigger frame, and the multimodal dialogue big model.
[0034] In one possible implementation, after sending the first interactive data to the electronic device, the method further includes: if it is determined that the n video frames following the first trigger frame are the same as the subject being filmed in the first trigger frame, determining first comprehensive information based on the sixth video frame among the n video frames; wherein the first comprehensive information includes detailed information about the first subject being filmed; generating third interactive data based on historical communication data, the sixth video frame, the first comprehensive information, and a multimodal dialogue model, wherein the amount of data in the third interactive data is greater than the amount of data in the first interactive data; and sending the third interactive data to the electronic device. In this way, when the n video frames of the electronic device are the same as the subject being filmed in the first trigger frame, the cloud server believes that the likelihood of the user wanting more detailed information about the first subject is increased, and the cloud server can generate more interactive data about the first subject being filmed, providing the user with a detailed introduction to the first subject being filmed.
[0035] In one possible implementation, the first comprehensive information is determined based on the sixth video frame out of y video frames. Specifically, this includes determining the first comprehensive information based on the sixth video frame and related information of the sixth video frame. The related information includes one or more of the following: semantic information, facial expression information of the subject being photographed, action information of the subject being photographed, QR code information, and text information of the video frame.
[0036] Sixthly, this application provides a communication system including an electronic device and a cloud server. The electronic device is used to execute the interaction method in any possible implementation of the first aspect, and the cloud server is used to execute the interaction method in any possible implementation of the second aspect. Alternatively, the electronic device is used to execute the interaction method in any possible implementation of the fourth aspect, and the cloud server is used to execute the interaction method in any possible implementation of the fifth aspect.
[0037] In a seventh aspect, this application provides an electronic device, including: one or more processors, one or more memories, and a camera; the camera, the one or more memories being coupled to the one or more processors, the one or more memories being used to store a computer program, and when the one or more processors are executing the computer program, executing the interaction method in any possible implementation of the first aspect or any possible implementation of the fourth aspect.
[0038] Eighthly, this application provides a cloud server, including: one or more processors and one or more memories; the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer programs. When the one or more processors execute the computer programs, they execute the interaction methods in any possible implementation of the second aspect or any possible implementation of the fifth aspect.
[0039] Ninthly, embodiments of this application provide a computer storage medium for storing a computer program that, when executed by a processor, implements the interaction method in any of the possible implementations of the first aspect, the second aspect, the fourth aspect, or the fifth aspect.
[0040] In a tenth aspect, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the interaction method in any of the possible implementations of the first aspect, the second aspect, the fourth aspect, or the fifth aspect.
[0041] In one aspect, this application provides a chip system including a processing circuit and an interface circuit. The interface circuit is used to receive code instructions and transmit them to the processing circuit. The processing circuit is used to run the code instructions to execute the interaction method in any of the possible implementations of the first aspect, the second aspect, the fourth aspect, or the fifth aspect. Attached Figure Description
[0042] Figure 1 A schematic diagram of a communication system 10 provided in an embodiment of this application;
[0043] Figures 2A-2H A set of interface schematic diagrams provided for embodiments of this application; Figure 3 A flowchart illustrating an interaction method provided in an embodiment of this application; Figure 4 A schematic diagram of a module provided in an embodiment of this application; Figure 5 Another schematic diagram of a module provided in an embodiment of this application; Figure 6 A flowchart illustrating the process of determining a trigger frame is provided in an embodiment of this application. Figure 7 A schematic diagram of a segmentation mask provided in an embodiment of this application; Figure 8 A schematic diagram of another segmentation mask provided in an embodiment of this application; Figure 9 An interactive schematic diagram provided for an embodiment of this application; Figure 10 A schematic diagram illustrating video frames and their specified information stored in a cloud server 200, provided as an embodiment of this application; Figure 11 A flowchart illustrating an interaction method provided in an embodiment of this application; Figures 12A-12H Another set of interface schematic diagrams provided for embodiments of this application; Figure 13 This is a schematic diagram of the structure of an electronic device 100 provided in an embodiment of this application; Figure 14 This is a schematic diagram of the structure of a cloud server 200 provided in an embodiment of this application. Detailed Implementation
[0044] The technical solutions in the embodiments of this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; the word "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0045] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0046] In one possible implementation, the smart assistant application of electronic device 100 can acquire image data via a camera. After receiving dialogue data input by the user, the smart assistant application of electronic device 100 can process the image data based on that dialogue data. When the image data and dialogue data correspond, electronic device 100 can output response data. For example, in a cooking scenario, the smart assistant application of electronic device 100 can receive data instructing the application to assist the user in cooking. The smart assistant application of electronic device 100 can gradually output response data guiding the user in cooking based on the image data. For instance, when the image data identifies the user's arrival at the kitchen, the electronic device 100 can output response data prompting the user on how to prepare the ingredients. After the image data identifies the user preparing the ingredients, the electronic device 100 can output response data prompting the user on how to stir-fry the ingredients, and so on. In this way, the smart assistant application can proactively output response data at appropriate times based on the image data to help the user complete the tasks instructed by the dialogue data.
[0047] Although the electronic device 100 can actively output response data during the image data acquisition process, this response data is obtained based on dialogue data input by the user. This constitutes an active interaction process with user input and cannot achieve an active interaction process without user input. When the user does not know what to input, the electronic device 100 cannot output response data, resulting in a poor user experience.
[0048] In one possible implementation, the smart assistant application of electronic device 100 can acquire image data through a camera. The electronic device can process the image data using a visual semantic parsing model to determine the image semantics corresponding to the image data. When the image semantics satisfy the active interaction conditions, electronic device 100 can output interactive data. For example, active interaction conditions can include, but are not limited to, image semantic indications of a specified event (e.g., returning home, waking up), a specified user state (e.g., happy, frustrated), or a specified scene (e.g., scenic spot). In this way, electronic device 100 can initiate an active interaction process, improving the user experience. However, the active interaction conditions that the smart assistant application can set are limited; it can only initiate dialogue when the active interaction conditions are met, lacking versatility and having a narrow application scenario.
[0049] This application provides an interaction method. Based on a user's first operation on the electronic device 100 or pre-set parameters on the electronic device 100, the electronic device 100's chat application is triggered to enter a no-intention proactive interaction mode. This no-intention proactive interaction mode is an interaction mode where the chat application proactively initiates a chat conversation with the user about a specific topic. In this mode, the electronic device 100 acquires an image captured by its camera. Based on the recognition of a first content in the image, the electronic device 100 initiates a chat conversation with the user about that first content. Thus, in this no-intention proactive interaction mode, the electronic device 100 can proactively recognize images captured by its camera and, based on the recognition of a first content in the image, proactively initiate a chat conversation with the user about that first content. Even when the user doesn't know what to input, it can proactively chat with the user, increasing the user's interest in using the chat application and providing a better user experience.
[0050] In one possible implementation, upon receiving input 1, the electronic device 100 can display a first interface in response to input 1. The first interface includes a first video stream captured by the electronic device 100 via a camera. The electronic device 100 can send the first video stream to a cloud server 200. Upon receiving the first video stream, the cloud server 200 can determine first interactive data based on the first video stream. The cloud server 200 can then send the first interactive data to the electronic device 100. Upon receiving the first interactive data, the electronic device 100 can output the first interactive data. In this way, the electronic device 100 can proactively output interactive data even without receiving user-input dialogue data, providing the user with a more intelligent interactive experience.
[0051] In some examples, after receiving the first video stream, the cloud server 200 can determine one or more trigger frames in the first video stream, where the one or more trigger frames include the first trigger frame, and the first trigger frame includes the first subject being filmed. The first interactive data includes an introduction to the first subject being filmed.
[0052] For example, when a user is visiting a museum or botanical garden, the electronic device 100 can proactively explain to the user the objects captured by its camera during the visit, using the interaction method provided in this application embodiment. In these scenarios, users tend to remain quiet and do not want the electronic device 100 to answer their questions. Instead, they want the electronic device 100 to capture real-time footage, allowing it to proactively explain objects that might interest them, thus enabling them to explore their surroundings. It should be noted that the electronic device 100 does not receive user-input dialogue data, and the smart assistant application cannot passively follow instructions from the dialogue data to process video frames. Therefore, for the electronic device 100, the above usage scenario can be understood as a proactive interaction process without user input.
[0053] In one possible implementation, when the electronic device 100 captures a video stream through a camera, it can implement the interaction method provided in this application embodiment, outputting interactive data to the user based on the video stream acquired by the camera. The electronic device 100 can also receive dialogue data input by the user, and output response data to the user based on the dialogue data and the video frames indicated by the dialogue data. In this way, the electronic device 100 can both proactively output interactive data to the user and passively respond to the user's needs, resulting in more diverse application scenarios and providing users with more intelligent services.
[0054] In some examples, the electronic device 100 includes an active interaction mode and a passive interaction mode. In active interaction mode, the electronic device 100 can output interactive data to the user based on the video stream acquired by the camera. In passive interaction mode, the electronic device 100 can receive dialogue data input by the user and output response data to the user. The electronic device 100 can also receive user settings input to switch between active and passive interaction modes.
[0055] The following describes a communication system 10 provided in an embodiment of this application.
[0056] For example, such as Figure 1 As shown, the communication system 10 may include an electronic device 100 and a cloud server 200. The electronic device 100 can acquire video streams via a camera. The electronic device 100 can send the video streams of the electronic device 100 to the cloud server 200 via a communication network. For example, the communication network can be a wired network and / or a wireless network.
[0057] In some examples, electronic device 100 can send a video stream to cloud server 200 at a specified frame rate. For example, the specified frame rate could be the frame rate at which the camera of electronic device 100 captures the video stream, or it could be a frame rate preset by electronic device 100 or cloud server 200. Optionally, electronic device 100 can determine the specified frame rate based on the network speed of the communication network. The faster the network speed, the higher the value of the specified frame rate.
[0058] The cloud server 200 can receive video streams sent by one or more electronic devices (including electronic device 100) via a communication network. The cloud server 200 can identify trigger frames in the video stream. Based on the trigger frames, the cloud server 200 can determine interactive data and send the interactive data to the electronic device 100. The electronic device 100 can output the interactive data after receiving it from the cloud server 200.
[0059] In some examples, the communication links used by electronic device 100 and cloud server 200 to transmit video frames and interactive data are different. For example, electronic device 100 can send video frames to cloud server 200 via communication link 1. Cloud server 200 can send interactive data to electronic device 100 via communication link 2.
[0060] In this system, the smart assistant application of electronic device 100 can send proactive interaction requests to cloud server 200 via communication link 2 at preset intervals. Upon receiving a proactive interaction request, cloud server 200 determines the interaction data based on the trigger frame whose acquisition time is closest to the time cloud server 200 receives the proactive interaction request, and then sends this interaction data to electronic device 100. The acquisition time of the trigger frame can be the time when the camera of electronic device 100 captures the trigger frame, the time when electronic device 100 sends the trigger frame, the time when cloud server 200 receives the trigger frame, or the time when cloud server 200 determines the trigger frame from the video stream, etc. Thus, since not all video frames are trigger frames, cloud server 200 does not need to maintain communication link 2 with electronic device 100 continuously. Electronic device 100 sends proactive interaction requests to cloud server 200 at regular intervals to obtain interaction data, reducing the communication resources used. Cloud server 200 can then provide the service of determining interaction data to more electronic devices.
[0061] In this application embodiment, the electronic device 100 may be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, as well as cellular phone, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence (AI) device, wearable device (e.g., smart glasses with camera, smartwatch, etc.), in-vehicle device, smart home device and / or smart city device. This application embodiment does not impose any special restrictions on the specific type of electronic device.
[0062] The following section presents a set of interface diagrams provided in the embodiments of this application.
[0063] For example, such as Figure 2AAs shown, the electronic device 100 can display a desktop 200. The desktop 200 can include multiple application icons (e.g., a smart assistant application icon 201, etc.). The smart assistant application icon 201 can be used to trigger the electronic device to display the interface of the smart assistant application, which can be used to wake up the smart assistant. Optionally, a status bar is also displayed at the top of the desktop 200, which may include information such as a communication signal strength indicator, battery level, and time. Optionally, a tray (dock) area may be present at the bottom of the desktop 200, which may include one or more tray icons (e.g., a dialer application icon, a messaging application icon, a contacts application icon, a camera application icon, etc.).
[0064] After receiving input from a user to wake up the smart assistant, the electronic device can respond to that input by displaying something like... Figure 2B The dialogue interface 210 is shown. The input for waking up the smart assistant can be a specific voice command, such as a preset voice command or a voice command input by the user in advance. Here, the voice command could be "Hey Celia". It should be noted that the smart assistant is not limited to being woken up by a voice command; the electronic device can also receive input from the user pressing and holding the power button to wake up the smart assistant. Alternatively, the electronic device can receive input from the user targeting a specific control in the smart assistant application's interface to wake up the smart assistant, etc. This embodiment does not limit this aspect.
[0065] like Figure 2B As shown, the dialogue interface 210 includes a dialogue display area 211. The dialogue display area 211 includes a chat page 212. The chat page 212 is used to display historical communication data between the user and the electronic device 100. Optionally, the dialogue interface 210 may also include a function bar 213. The function bar 213 may include one or more function controls. These one or more function controls may include, but are not limited to, a keyboard control 214, an add control 216, a voice control 215, etc. The keyboard control 214 can be used to trigger the electronic device 100 to display a keyboard area, which can be used by the user to type text data. The add control 216 can be used to trigger the electronic device 100 to display an add area, which can be used to send images, documents, etc., stored by the electronic device 100 to the smart assistant application. The voice control 215 can be used to trigger the electronic device 100 to stop collecting the user's voice signal. Here, the voice control 215 can also be used to indicate that the electronic device 100 is collecting the user's voice signal. The electronic device 100 can convert the collected voice signal into text data and display the text data on the chat page 212.
[0066] Optionally, after receiving input from the user to wake up the smart assistant, the electronic device 100 may respond to the input by displaying, as shown below. Figure 2B Simultaneously with the dialogue interface 210 shown, input prompts are output to guide the user through dialogue data. For example, the electronic device 100 can verbally announce the input prompts. For instance, the input prompt could be: "I'm here." Thus, the user can input dialogue data into the electronic device 100 after hearing the input prompts.
[0067] like Figure 2C As shown, the electronic device 100 can display a dialog box 217 including the specified dialogue data after receiving it. In this embodiment, the specified dialogue data can be used to trigger the electronic device 100 to output interactive data based on the video stream captured by the camera. For example, the specified dialogue data may include the voice input "Open Xiaoyi to see the world". It should be noted that "Open Xiaoyi to see the world" is only an example of the specified dialogue data and should not be construed as limiting the specific content of the specified dialogue data.
[0068] When the electronic device 100 receives the specified dialogue data, it can respond to the specified dialogue data by displaying, for example... Figure 2D The object recognition interface 220 is shown. The electronic device 100 can also respond to specified dialogue data, determine and output interactive data based on the video stream acquired by the camera.
[0069] like Figure 2D As shown, the object recognition interface 220 can be used to display the video stream captured by the camera. Here, the electronic device 100 displays video frames 221 from the video stream captured by the camera in the object recognition interface 220. It can be understood that the electronic device 100 can output interactive data while displaying the object recognition interface 220.
[0070] In some examples, the object recognition interface 220 may also include settings 222. Settings 222 may include text prompts for "active interaction function" and a switch control. The text prompts can be used to inform the user of the function of settings 222. The switch control may include an on state and an off state, and can be used to trigger the electronic device 100 to receive and respond to user input (e.g., a click) to toggle the state of the switch control. When the switch control is in the on state, the electronic device 100 enables the active interaction function and can determine and output interaction data based on the video stream captured by the camera. When the switch control is in the off state, the electronic device 100 disables the active interaction function and cannot determine and output interaction data based on the video stream captured by the camera. Optionally, when the switch control is in the off state, the electronic device 100 disables the active interaction function, but can receive user dialogue data and determine and output response data based on the dialogue data and the video frames indicated by the dialogue data. Figure 2DIn the object recognition interface 220 shown, the switch control for setting item 222 is in the "on" state. This allows users to quickly enable / disable the active interaction function through setting item 222, improving the user experience.
[0071] In some examples, the object recognition interface 220 may also include, but is not limited to, one or more of the following: control 223, control 224, camera 225, control 226, and control 227. Control 223 can be used to switch the camera used by the electronic device 100 to capture the video stream (e.g., switch between a front-facing camera and a rear-facing camera). Control 224 can be used to turn the microphone on / off. Control 225 can be used to turn the camera on / off. Control 226 can be used to trigger the electronic device 100 to cancel the display of the object recognition interface 220. For example, after the electronic device 100 cancels the display of the object recognition interface 220, it may display something like... Figure 2A The desktop 200 shown. Control 227 can be used to turn the subtitles on / off. When the electronic device 100 does not turn on the subtitles, the electronic device 100 can verbally announce interactive data and / or response data. When the electronic device 100 turns on the subtitles, the electronic device 100 can display interactive data and / or response data on the object recognition interface 220. Optionally, when the electronic device 100 turns on the subtitles, the electronic device 100 can also verbally announce interactive data and / or response data. Optionally, the object recognition interface 220 may also include control 228 (…). Figure 2D (Not shown in the image). Control 228 can be used to turn the voice broadcast function on / off.
[0072] In other examples, electronic device 100 may provide a designated control for setting active interaction functionality. This designated control can be used to turn the active interaction functionality on / off. If electronic device 100 enables the active interaction functionality, it can determine and output interaction data based on the video stream captured by the camera. If electronic device 100 disables the active interaction functionality, it cannot determine and output interaction data based on the video stream captured by the camera. Optionally, if electronic device 100 disables the active interaction functionality, it can receive user dialogue data and determine and output response data based on the dialogue data and the video frames indicated by the dialogue data.
[0073] For example, after the electronic device 100 activates the active interaction function, upon receiving the aforementioned specified dialogue data, it can respond to the specified dialogue data by displaying, as shown below. Figure 2D The object recognition interface 220 shown determines and outputs interactive data based on the video stream captured by the camera. After the active interaction function is disabled, when the electronic device 100 receives the specified dialogue data, it can respond to the specified dialogue data by displaying, as shown below. Figure 2DThe object recognition interface 220 shown receives user dialogue data, determines and outputs response data based on the dialogue data and the video frames indicated by the dialogue data. It should be noted that when the electronic device 100 disables its interactive function, the switch control of setting item 222 is in the off state.
[0074] For example, after receiving a video stream sent by an electronic device 100, the cloud server 200 can determine the interaction data 232 based on the video stream. The video stream includes... Figure 2D The video frame 221 is shown. The cloud server 200 can send interactive data 232 to the electronic device 100. After receiving the interactive data 232 sent by the cloud server 200, the electronic device 100 outputs the interactive data 232.
[0075] For example, electronic device 100 can display such as in object recognition interface 220 Figure 2E When video frame 231 is displayed, the interactive data 232 is played. For example, interactive data 232 may include: "This potted green plant is growing so well! It's beautiful to have in the office and it purifies the air. It makes people feel good just looking at it. What are some tips for taking care of it?" Interactive data 232 includes descriptions of the "potted green plant" in video frame 221. The subject in video frame 231 is the same as the subject in video frame 221. Thus, the electronic device 100 can output interactive data 232 while displaying the same subject. Alternatively, the subject in video frame 231 may be different from the subject in video frame 221. In this case, when the user moves the camera of the electronic device 100 away from the subject in video frame 221, the electronic device 100 has already determined interactive data 232, and can still output interactive data 232 even if the electronic device 100 does not display the subject in video frame 221.
[0076] For example, after displaying the object recognition interface 220, the electronic device 100 can continuously send the video stream captured by the camera to the cloud server 200. After determining the interaction data 242 based on the video stream from the electronic device 100, the cloud server 200 can send the interaction data 242 to the electronic device 100. After receiving the interaction data 242 sent by the cloud server 200, the electronic device 100 can output the interaction data 242.
[0077] For example, electronic device 100 can display such as in object recognition interface 220 Figure 2F When video frame 241 is shown, the interactive data 242 is broadcast. For example, interactive data 242 may include: "Getting ready to see 'Step By'? The view from seat 6 in row 5 of the IMX theater should be pretty good. It starts at 4:45 PM, remember to arrive early!" Interactive data 242 includes information about the "movie ticket" shown in video frame 241.
[0078] For example, after determining the interaction data 252 based on the video stream of the electronic device 100, the cloud server 200 can send the interaction data 252 to the electronic device 100. After receiving the interaction data 252 sent by the cloud server 200, the electronic device 100 can output the interaction data 252.
[0079] For example, electronic device 100 can display such as in object recognition interface 220 Figure 2G When video frame 251 is shown, the interactive data 252 is played. For example, interactive data 252 may include: "This white XX brand sedan is really nice, it looks very sporty. Would you like to know more about this sedan?" Interactive data 252 includes information about the "vehicle" being filmed in video frame 251.
[0080] In this way, the electronic device 100 can actively output interactive data during the video recording process, giving users the experience of listening to the introduction of the subject in the video while recording it.
[0081] Optionally, the electronic device 100 can receive user input... Figure 2A After inputting the smart assistant application icon 201 as shown, in response to that input, the following will be displayed: Figure 2H The smart assistant interface shown is 260. (As shown in the image...) Figure 2H As shown, the smart assistant interface 260 may include one or more controls provided by the smart assistant application. These one or more controls can be used to trigger the smart assistant application to perform corresponding operations. These one or more controls may include control 261. Upon receiving user input to control 261, the electronic device 100 can respond to the input by displaying, as shown... Figure 2D The object recognition interface 220 shown outputs interactive data based on the video stream captured by the camera. In this way, the electronic device 100 can provide users with the function of outputting interactive data based on the video stream through a smart assistant application, helping users to actively interact when it is inconvenient to input dialogue data.
[0082] The following is a flowchart illustrating an interaction method provided in an embodiment of this application.
[0083] For example, such as Figure 3 As shown, the interaction method provided in this application embodiment includes the following steps: S301. Electronic device 100 receives input 1.
[0084] S302. Electronic device 100 responds to input 1 by displaying a first interface; wherein the first interface includes a first video stream captured by a camera.
[0085] S303. Electronic device 100 sends the first video stream to cloud server 200.
[0086] Upon receiving input 1, the electronic device 100 responds by displaying a first interface and sending a first video stream captured by a camera to the cloud server 200. Input 1 can be used to trigger the electronic device 100 to execute the interaction method provided in this embodiment. For example, input 1 can be a response to... Figure 2C The electronic device 100 shown inputs specified dialogue data. For example, if... Figure 2D The switch control for setting item 222 shown is in the off state. Input 1 can be used to... Figure 2D The input for setting item 222 shown is (for example, clicking). For example, inputting 1 could be a... Figure 2H The input of control 261 shown, etc.
[0087] In some examples, electronic device 100 has a smart assistant application installed. The smart assistant application of electronic device 100 can receive user input 1, display a first interface provided by the smart assistant application, and send a first video stream captured by a camera to cloud server 200. The smart assistant application can also receive first interactive data sent by cloud server 200 and output that first interactive data.
[0088] It should be noted that, not limited to smart assistant applications, other applications of the electronic device 100 (e.g., chat applications, browser applications, short video applications, etc.) can also provide smart assistant functions. After enabling the smart assistant function, other applications can receive user input 1, display a first interface, and determine interactive data from the first video stream captured by the camera. In the following embodiments, the interaction method provided in the embodiments of this application will be described using a smart assistant application as an example.
[0089] Understandably, after receiving input 1, electronic device 100 can continuously send the first video stream to cloud server 200 until it receives input from the user instructing it to stop executing the interaction method. For example, the input instructing electronic device 100 to stop executing the interaction method could be... Figure 2D , Figure 2E , Figure 2F or Figure 2G The input for control 226 is shown. This allows the user to turn the object recognition interface on / off independently.
[0090] S304. Cloud server 200 can determine one or more trigger frames in the first video stream based on the received first video stream; wherein, the one or more trigger frames include the first trigger frame, which includes the first shooting object.
[0091] S305. Cloud Server 200 determines the first interaction data based on the first trigger frame.
[0092] The cloud server 200, upon receiving the first video stream from the electronic device 100, can determine one or more trigger frames based on the first video stream. The subject in each trigger frame is the same as the subject in the preceding m video frames and different from the subject in the (m+1)th video frame before the trigger frame. Alternatively, the subjects in the preceding (m+1) video frames are all the same, and the subject in this trigger frame is different from the subjects in the preceding (m+1) video frames, where m is greater than or equal to 0.
[0093] It should be noted that if the subjects in the first (m+1) video frames of the first video stream are all the same, the cloud server 200 can mark the (m+1)th video frame of the first video stream as the trigger frame. Thus, when the electronic device 100 points its camera at a subject, since the electronic device 100 is not capturing any other objects, but the subject is also a newly appearing subject, the (m+1)th video frame is marked as the trigger frame.
[0094] In some examples, the subject in the trigger frame is the same as the subject in the preceding m video frames of the trigger frame, and different from the subject in the (m+1)th video frame before the trigger frame. Specifically, the trigger frame includes subject 1, and the (m+1)th video frame before the trigger frame does not include subject 1. And / or, the subjects in the preceding (m+1)th video frames are all the same, and the subject in the trigger frame is different from the subjects in the preceding (m+1)th video frames of the trigger frame. Specifically, the preceding (m+1)th video frames include subject 2, and the trigger frame does not include subject 2, where m is greater than or equal to 0.
[0095] In this way, when the subject 1 appears in (m+1) consecutive video frames in the first video stream of electronic device 100, the likelihood of the camera of electronic device 100 randomly scanning over the subject 1 decreases. Cloud server 200 tends to believe that the user wants to know about the subject 1, and can determine the (m+1)th video frame as the trigger frame. After the subject 2 appears in (m+1) consecutive video frames in the first video stream of electronic device 100, and the subject 2 disappears in the (m+2)th video frame, cloud server 200 determines that the subject 2 that the user has been paying attention to has disappeared, and can determine the (m+2)th video frame as the trigger frame. Cloud server 200 can combine the (m+2)th video frame with one or more previous video frames that include the subject 2 to determine the interactive data used to prompt the user that the subject 2 has disappeared. This avoids the situation where, if electronic device 100 sets active interaction conditions, it repeatedly outputs the same content to the user after the same active condition is met, providing the user with a more intelligent and personalized active interaction experience.
[0096] In some examples, the cloud server 200 can mark a trigger frame as either a trigger frame for adding a new shooting object or a trigger frame for removing a shooting object. When the cloud server 200 determines that a trigger frame is for adding a shooting object, it can determine interactive data based on that trigger frame, or it can determine interactive data based on that trigger frame and one or more video frames following that trigger frame that have the same shooting object as the trigger frame. When the cloud server 200 determines that a trigger frame is for removing a shooting object, it can determine interactive data based on that trigger frame, or it can determine interactive data based on that trigger frame and one or more video frames preceding that trigger frame that have a different shooting object than the trigger frame. Thus, when a new shooting object appears in the video stream, the electronic device 100 can display interactive data to introduce the newly added shooting object, helping the user understand the newly appeared shooting object. When an old shooting object disappears from the video stream, the electronic device 100 can display interactive data to prompt the user that the old shooting object has disappeared, helping the user understand the change in the old shooting object.
[0097] In some examples, cloud server 200 can determine first image information describing the first captured object based on the first trigger frame. Cloud server 200 can generate first interaction data based on the first image information, the first trigger frame, and the multimodal dialogue model.
[0098] In some examples, the first trigger frame includes multiple subjects, the first subject being the largest subject in the video frame, or the most central subject, or the subject closest to the camera (i.e., the user).
[0099] S306. Cloud server 200 sends the first interactive data to electronic device 100.
[0100] After determining the first interaction data, the cloud server 200 can send the first interaction data to the electronic device 100.
[0101] S307. Electronic device 100 can output first interactive data after receiving first interactive data.
[0102] After receiving the first interactive data, the electronic device 100 can broadcast the first interactive data through a speaker. Alternatively, the electronic device 100 can establish a communication connection with a headset device (e.g., wired connection, Wi-Fi connection, Bluetooth connection, etc.), and the electronic device 100 can broadcast the first interactive data through the headset device.
[0103] Optionally, the electronic device 100 may also display the first interactive data on a screen. This helps the user accurately receive the first interactive data.
[0104] In other examples, after receiving user input to mute the microphone, electronic device 100 can display the first interactive data on the screen.
[0105] In some examples, the first video stream includes a first video frame, and the first video frame includes a first subject being captured. The electronic device 100 can output first interactive data when displaying the first video frame on a first interface. For example, the first trigger frame could be... Figure 2D The video frame 221 shown can be the first video frame. Figure 2E The video frame shown is 231.
[0106] In other examples, since it takes time for the cloud server 200 to determine the first interactive data based on the first trigger frame, the electronic device 100 may receive and output the first interactive data while displaying a second video frame that includes the second subject. In this way, the electronic device 100 can continue to output the first interactive data after the subject in its camera changes.
[0107] Optionally, after receiving the first interactive data, the electronic device 100 can detect whether the shooting object of the video frame displayed on the first interface of the electronic device 100 is the same as the shooting object described in the first interactive data. The electronic device 100 can output the first interactive data when it detects that the shooting object of the video frame displayed on the first interface is the same as the shooting object described in the first interactive data. The electronic device 100 can also not output the first interactive data when it detects that the shooting object of the video frame displayed on the first interface is different from the shooting object described in the first interactive data. In this way, if the user removes the electronic device 100 before it outputs the first interactive data describing the first shooting object, and the shooting object changes, it can indicate that the user prefers not to know about the first shooting object, and the electronic device 100 can choose not to output the first interactive data.
[0108] For example, the electronic device 100 can identify the name of the subject in the video frame displayed on the first interface based on an image recognition algorithm. If the subject is an animal or plant, the name can include its genus (e.g., *Canis*, *Felis*, etc.) or species (e.g., *Xiaosi Dog*, *Corgi*, etc.). If the subject is a person, the name can include, but is not limited to, one or more of the person's name, age, etc. The electronic device 100 can determine that the subject in the video frame displayed on the first interface is the same as the subject described in the first interactive data when the first interactive data includes all or part of the identified name. The electronic device 100 can determine that the subject in the video frame displayed on the first interface is different from the subject described in the first interactive data when the first interactive data does not include the identified name. Furthermore, the electronic device 100 can also identify whether the subject has changed using object re-identification algorithms, face recognition algorithms, etc., etc., and this application embodiment does not limit this.
[0109] In one possible implementation, electronic device 100 can send a first request to cloud server 200. Upon receiving the first request from electronic device 100, cloud server 200 can determine first interaction data based on a first trigger frame and send the first interaction data to electronic device 100. Specifically, the acquisition time of the first trigger frame (one or more trigger frames) is closest to the time when cloud server 200 receives the first request.
[0110] In some examples, the one or more trigger frames may also include a second trigger frame. The electronic device 100 may also send a second request to the cloud server 200. Upon receiving the second request from the electronic device 100, the cloud server 200 may determine second interaction data based on the second trigger frame and send the second interaction data to the electronic device 100. Optionally, the difference between the time the electronic device 100 sends the second request and the time it sends the first request is a preset value. In this way, the electronic device 100 may send proactive interaction requests (e.g., first request, second request, etc.) to the cloud server 200 at preset intervals, triggering the cloud server 200 to send interaction data to the electronic device 100.
[0111] In other examples, electronic device 100 can determine the frequency of sending proactive interaction requests within a preset transmission duration after the current moment based on the frequency of receiving interactive data within a preset historical duration prior to the current moment. Specifically, the higher the frequency of electronic device 100 receiving interactive data within the preset historical duration, the higher the frequency of sending proactive interaction requests within the preset transmission duration. Conversely, the lower the frequency of electronic device 100 receiving interactive data within the preset historical duration, the lower the frequency of sending proactive interaction requests within the preset transmission duration. Thus, since cloud server 200 can send interactive data based on each proactive interaction request sent by electronic device 100, indicating a large number of trigger frames in the first video stream, electronic device 100 can increase the frequency of sending proactive interaction requests to obtain more timely interactive data.
[0112] It should be noted that the frequency of the electronic device 100 sending proactive interaction requests in the next period is not limited to being determined based on the frequency of interactive data received by the electronic device 100 in the past period. The electronic device 100 can also determine the frequency of sending proactive interaction requests based on the network speed of the communication network. Generally, the faster the communication network speed, the higher the frequency of the electronic device 100 sending proactive interaction requests; the slower the communication network speed, the lower the frequency of the electronic device 100 sending proactive interaction requests, and so on. This application embodiment does not limit this.
[0113] In one possible implementation, the cloud server 200, upon detecting that the subject in any of the n video frames following the first trigger frame is the same as the subject in the first trigger frame, determines third interactive data based on any one of the n video frames or the first trigger frame and sends the third interactive data to the electronic device 100. The amount of the third interactive data is greater than the amount of the first interactive data. The third interactive data includes detailed information about the first subject. Thus, when the subject in the video captured by the camera of the electronic device 100 remains unchanged, it usually indicates that the user prefers to know more about the subject. The electronic device 100 can then acquire and output interactive data including more detailed information about the subject, providing the user with a more intelligent object recognition experience.
[0114] For example, the first video stream also includes a third video frame, which includes the first subject being captured. The electronic device 100 can output third interactive data when displaying the third video frame on the first interface. All video frames between the third and first video frames include the first subject being captured.
[0115] In some examples, the electronic device 100 can display a second interface before receiving input 1. This second interface includes a second video stream, which in turn includes a fourth video frame. While displaying the second interface, the electronic device 100 receives first dialogue data input by the user. Based on the first dialogue data and the fourth video frame indicated by the first dialogue data, the electronic device 100 can output first response data. In this way, the electronic device 100 can first display a second interface for receiving dialogue data from the user and outputting response data, and after receiving input 1, display a first interface for outputting interactive data based on the video stream captured by the camera. This provides the user with multiple interaction modes, is suitable for more diverse scenarios, and meets the user's various object recognition needs.
[0116] In some examples, the electronic device 100 is set to active interaction mode by default. When the electronic device 100 receives dialogue data input by the user in active mode, it can switch from active interaction mode to passive interaction mode. The descriptions of active and passive interaction modes can be found in the above embodiments and will not be repeated here.
[0117] Optionally, the electronic device 100 can only be set to active interaction mode after its active interaction function is enabled. This allows users to decide whether to use active interaction mode, which better suits their usage habits.
[0118] In some examples, electronic device 100 can receive user voice input via a microphone. Electronic device 100 can identify whether the sound signal collected by the microphone belongs to the user. For example, electronic device 100 can use voiceprint recognition technology to identify whether the collected sound signal belongs to the user. When it is determined that the collected sound signal belongs to the user, electronic device 100 can receive the user's input dialogue data via the microphone, switching from active interaction mode to passive interaction mode. For example, electronic device 100 can determine whether the received voice input is the user's voice input by detecting whether the audio features of the voice input match the user's audio features; if the audio features of the voice input match the user's audio features, the voice input is determined to be the user's voice input. In this way, electronic device 100 only receives user-input dialogue data, avoiding situations where electronic device 100 cannot enter active interaction mode due to noisy ambient sounds.
[0119] In some examples, electronic device 100 can detect the direction of the sound source of the voice input. Electronic device 100 can determine that the voice input is user voice input when it detects that the direction of the sound source of the voice input is behind the camera of electronic device 100.
[0120] In some examples, after outputting response data in passive interaction mode, electronic device 100 can switch to active interaction mode. Alternatively, if electronic device 100 does not receive any dialogue data from the user within a preset waiting time in passive interaction mode, it can switch to active interaction mode. In this way, electronic device 100 can automatically switch to active interaction mode, providing users with a more proactive object recognition service.
[0121] It should be noted that, in addition to using voiceprint recognition technology to identify whether a sound signal belongs to a user, the electronic device 100 can also use noise reduction algorithms to remove noise from the environment other than the user, etc. This application embodiment does not limit this.
[0122] In other examples, the first interface of the electronic device 100 includes a long-press voice input control. The electronic device 100 can receive user-input dialogue data via a microphone while receiving user input (e.g., touch) to the long-press voice input control. It is understood that when the long-press voice input control does not receive user input, the electronic device 100 cannot receive user-input dialogue data via the microphone. This allows the user to decide whether to input dialogue data.
[0123] The following is a schematic diagram of a module provided by an embodiment of this application.
[0124] For example, such as Figure 4 As shown, the electronic device 100 may include, but is not limited to, a smart assistant application 11, a camera 12, and a display screen 13. The cloud server 200 may include, but is not limited to, a trigger frame recognition module 21, an image search module 22, and a multimodal dialogue model 23. The electronic device 100 and the cloud server 200 can transmit data to each other via a communication network.
[0125] The camera 12 can be used to capture video streams. The display screen 13 can be used to display video streams.
[0126] The smart assistant application 11 can provide users with proactive interaction functions. The smart assistant application 11 can notify the camera 12 to capture a video stream and acquire the video stream captured by the camera 12. The smart assistant application 11 can also notify the camera 12 to send the captured video stream to the display screen 13. Optionally, the smart assistant application 11 can send the video stream to the display screen 13 after receiving it from the camera 12.
[0127] The smart assistant application 11 can send the video stream provided by the camera 12 to the cloud server 200 through the communication network.
[0128] The trigger frame recognition module 21 can be used to determine trigger frames from the video stream sent by the electronic device 100. The trigger frame recognition module 21 can send the trigger frames to the image search module 22 and the multimodal dialogue model 23. For a detailed description of how the trigger frame recognition module 21 determines the trigger frames, please refer to [link to relevant documentation]. Figure 6 The illustrated embodiment.
[0129] The image search module 22 can be used to determine the image information of the trigger frame based on the trigger frame. This image information includes a description of the object being photographed in the trigger frame. The image search module 22 can send the image information of the trigger frame to the multimodal dialogue model 23. For example, when the object is a building, the description may include, but is not limited to, one or more of the following: the building's name, address, function, and historical background. When the object is a plant or animal, the description may include, but is not limited to, one or more of the following: the plant or animal's name, distribution area, and economic value. When the object is a person, the description may include, but is not limited to, one or more of the following: the person's name, age, hairstyle, clothing, and makeup. When the object is an item, the description may include, but is not limited to, one or more of the following: the item's name, use, and size, and so on.
[0130] The multimodal dialogue model 23 can be used to generate interactive data based on the trigger frame and its image information. In this way, interactive data that better reflects human speaking habits can be obtained through the multimodal dialogue model 23. Optionally, the trigger frame recognition module 21, the image search module 22, and the multimodal dialogue model 23 of the cloud server 200 can be collectively referred to as the proactive interaction service module.
[0131] After receiving the interaction data, the cloud server 200 can send the interaction data to the electronic device 100. The smart assistant application 11 can then output this interaction data. For example, the electronic device 100 may also include a speaker, allowing the smart assistant application 11 to broadcast the interaction data. Alternatively, the electronic device 100 may establish a communication connection with a headset device, allowing the smart assistant application 11 to send the interaction data to the headset device, which can then broadcast the interaction data. Optionally, the smart assistant application 11 can also display the interaction data on the display screen 13.
[0132] In some examples, the electronic device 100 may also include a microphone 14. The electronic device 100 may receive dialogue data input by the user through the microphone 14 and determine response data based on the dialogue data. Optionally, the electronic device 100 may receive voice commands from the user through the microphone 14, and in response to the voice commands, open the smart assistant application 11, or execute the interaction method provided in the embodiments of this application.
[0133] In some examples, the electronic device 100 may also include a voiceprint recognition module 15. The voiceprint recognition module 15 can be used to identify whether the sound signal received by the microphone 14 belongs to the user. When the voiceprint recognition module 15 determines that the sound signal received by the microphone 14 belongs to the user, it can send the sound signal collected by the microphone 14 to the smart assistant application 11, or notify the microphone 14 to send the collected sound signal to the smart assistant application 11. The smart assistant application 11 can process the sound signal to obtain input data after receiving it. Alternatively, the voiceprint recognition module 15 can process the sound signal to obtain input data and send the input data to the smart assistant application 11.
[0134] The voiceprint recognition module 15 can prevent the voice signal received by the microphone 14 from being sent to the smart assistant application 11 when it determines that the voice signal does not belong to the user. In this way, the electronic device 100 can accurately receive the dialogue data input by the user and will not receive the dialogue data input by other users.
[0135] Optionally, the electronic device 100 may perform the above steps via the voiceprint recognition module 15 during the execution of the interaction method. When the electronic device 100 is not executing the interaction method, it does not use the voiceprint recognition module 15. Thus, in other application scenarios (e.g., recording scenarios), the electronic device 100 can acquire nearby sound signals via the microphone 14.
[0136] It should be noted that the above modules are merely examples. The electronic device 100 and the cloud server 200 can be split or combined with all or part of the modules, and this application embodiment does not limit this.
[0137] This application provides another interaction method. The smart assistant application of the electronic device 100 can display a first interface in response to receiving input 1. The first interface includes a first video stream captured by the electronic device 100 through a camera. The electronic device 100 can send the first video stream to a cloud server 200. The electronic device 100 can send a first request to the cloud server 200. After receiving the first request, the cloud server 200 can determine first interaction data based on the first video stream. The cloud server 200 can send the first interaction data to the electronic device 100. After receiving the first interaction data, the electronic device 100 can output the first interaction data. In this way, since the cloud server 200 only determines the interaction data based on a portion of the video frames in the first video stream and does not continuously send interaction data to the electronic device 100, the electronic device 100 obtains interaction data by sending proactive interaction requests (e.g., the first request) at preset intervals. This saves communication resources between the electronic device 100 and the cloud server 200, and also allows proactive output of interaction data even when no user dialogue data is received, providing a more intelligent interactive experience for the user.
[0138] In some examples, cloud server 200 can identify one or more trigger frames in the first video stream, including a first trigger frame that includes a first subject. Upon receiving a first request, cloud server 200 can determine first interactive data based on the first trigger frame, which includes a description of the first subject. Specifically, the acquisition time of the first trigger frame is closest to the time cloud server 200 receives the first request. This closest proximity between the acquisition time and the receipt time of the first request indicates that electronic device 100 is most likely displaying a video frame including the first subject. Cloud server 200's determination of the first interactive data based on the first trigger frame helps electronic device 100 output first interactive data describing the first subject when displaying a video frame including the first subject, providing users with a more intelligent object recognition service in conjunction with visual processing.
[0139] The following is a schematic diagram of a module provided by an embodiment of this application.
[0140] For example, such as Figure 5 As shown, the electronic device 100 may include, but is not limited to, a smart assistant application 11, a camera 12, and a display screen 13. The cloud server 200 may include, but is not limited to, a trigger frame recognition module 21, an image search module 22, a multimodal dialogue model 23, a video frame buffer 24, and a trigger frame selection module 25. The electronic device 100 and the cloud server 200 can transmit data to each other via a communication network.
[0141] The camera 12 can be used to capture video streams. The display screen 13 can be used to display the video streams. The smart assistant application 11 can provide users with proactive interactive functions. The smart assistant application 11 can notify the camera 12 to capture video streams and obtain the video streams captured by the camera 12. The smart assistant application 11 can also notify the camera 12 to send the captured video streams to the display screen 13. Optionally, the smart assistant application 11 can send the video streams to the display screen 13 after receiving the video streams captured by the camera 12. The smart assistant application 11 can send the video streams provided by the camera 12 to the cloud server 200 through a communication network.
[0142] After receiving the video stream sent by the electronic device 100, the cloud server 200 can store the video stream in the video frame buffer 24.
[0143] The trigger frame identification module 21 can obtain video frames from the video stream from the video frame buffer 24 and obtain trigger frame indication information based on the video frames. The trigger frame indication information can be used to indicate whether a video frame is a trigger frame. The trigger frame identification module 21 can send the trigger frame indication information to the video frame buffer 24. The video frame buffer 24 can mark trigger frames based on the trigger frame indication information. For a detailed description of how the trigger frame identification module 21 determines trigger frames, please refer to [link to relevant documentation]. Figure 6 The illustrated embodiment.
[0144] The smart assistant application 11 of electronic device 100 can also send an active interaction request to the trigger frame selection module 25 of cloud server 200 via a communication network. After receiving the active interaction request sent by electronic device 100, trigger frame selection module 25 can obtain the trigger frame from the video stream of electronic device 100 from video frame buffer 24. After obtaining the trigger frame, trigger frame selection module 25 can send the trigger frame to image search module 22 and multimodal dialogue big model 23.
[0145] The image search module 22 can be used to determine the image information of the trigger frame based on the trigger frame. The image information includes a description of the subject being photographed in the trigger frame. The image search module 22 can send the image information of the trigger frame to the multimodal dialogue model 23. The multimodal dialogue model 23 can be used to generate interactive data based on the trigger frame and its image information. Optionally, the trigger frame recognition module 21, image search module 22, multimodal dialogue model 23, video frame buffer 24, and trigger frame selection module 25 of the cloud server 200 can be collectively referred to as the active interaction service module.
[0146] After receiving the interaction data, the cloud server 200 can send the interaction data to the electronic device 100. The smart assistant application 11 can then output this interaction data.
[0147] In some examples, the electronic device 100 may also include a microphone 14. The electronic device 100 can receive user-input dialogue data via the microphone 14 and determine response data based on that dialogue data. A detailed description of the microphone 14 can be found in [link to relevant documentation]. Figure 4 The embodiments shown are not described in detail here.
[0148] In some examples, the electronic device 100 may also include a voiceprint recognition module 15. The voiceprint recognition module 15 can be used to identify whether the sound signal received by the microphone 14 belongs to the user. A detailed description of the voiceprint recognition module 15 can be found in [link to relevant documentation]. Figure 4 The embodiments shown are not described in detail here.
[0149] Thus, the electronic device 100 and the cloud server 200 can implement the interaction method provided in this application embodiment through the aforementioned multiple modules. It should be noted that the aforementioned multiple modules are merely examples, and the electronic device 100 and the cloud server 200 can be split / combined with all or part of the modules; this application embodiment does not limit this.
[0150] In one possible implementation, the cloud server 200 can determine the segmentation mask and confidence score of video frames in the first video stream based on the first video stream. The segmentation mask can be used to indicate the location of the subject in the video frame, and the confidence score can be used to represent the probability that the subject exists in the video frame. The first video stream includes a fifth video frame. If the confidence score of the fifth video frame is greater than a first score threshold, the subject of the fifth video frame can be determined based on the segmentation mask of the fifth video frame. If the subject of the fifth video frame is the same as the subjects in the previous m video frames and different from the subjects in the (m+1)th video frame before the fifth video frame, the fifth video frame is determined to be a trigger frame; or, if the subject of the fifth video frame is different from all the subjects in the previous (m+1) video frames and the subjects in the previous (m+1) video frames are the same, the fifth video frame is determined to be a trigger frame, where m is greater than or equal to 0. In this way, the cloud server 200 can quickly determine whether a video frame is a trigger frame by using the segmentation mask and confidence score of the video frame.
[0151] In some examples, if the subject in the fifth video frame is the same as the subject in the preceding m video frames, and different from the subject in the (m+1)th video frame before the fifth video frame, where the fifth video frame includes subject 1 and the (m+1)th video frame before the fifth video frame does not include subject 1; and / or, the subjects in the preceding (m+1)th video frames are all the same, and the subject in the fifth video frame is different from the subjects in the preceding (m+1)th video frames, where the preceding (m+1)th video frames include subject 2 and the fifth video frame does not include subject 2, then the fifth video frame is determined to be the trigger frame, where m is greater than or equal to 0.
[0152] The following is a flowchart illustrating the process of determining a trigger frame, as provided in an embodiment of this application.
[0153] For example, such as Figure 6 As shown, the process by which the cloud server 200 determines the trigger frame may include the following steps: S601. Cloud server 200 acquires video stream from electronic device 100.
[0154] S602. Cloud server 200 determines the segmentation mask and confidence score of video frames in the video stream, wherein the segmentation mask is used to indicate the location of the shooting object in the video frame, and the confidence score is used to indicate the probability that the shooting object exists in the video frame.
[0155] In some examples, the cloud server 200 can determine the segmentation mask and confidence score of video frames based on a saliency segmentation model. In this way, the saliency segmentation model can be trained on manually labeled saliency segmentation data, which segments according to human attention mechanisms, reducing the impact of background clutter on trigger frame recognition.
[0156] It should be noted that, not limited to the saliency segmentation model, the cloud server 200 can also determine the location of the shooting object in the video frame based on other models (e.g., object detection models), and this application embodiment does not limit this.
[0157] S603. Cloud Server 200 detects whether the confidence score of video frames is higher than the first score threshold.
[0158] The higher the confidence score, the greater the probability that a subject is present in the video frame, and the more accurate the location of the subject indicated by the segmentation mask of the video frame. After determining the confidence score of the video frame, the cloud server 200 can detect whether the confidence score of the video frame is higher than a first score threshold. If the cloud server 200 detects that the confidence score of the video frame is higher than the first score threshold, it can execute step S605. If the cloud server 200 detects that the confidence score of the video frame is lower than the first score threshold, it can execute step S604.
[0159] S604. Cloud server 200 indicates that this video frame is not a trigger frame.
[0160] If the cloud server 200 determines that the confidence score of a video frame is lower than a first score threshold, it can mark the video frame as not a trigger frame. In some examples, the cloud server 200 can set a trigger frame identifier for each video frame, using this identifier to indicate whether the video frame is a trigger frame. For example, when the trigger frame identifier of a video frame is "0", it means that the video frame is not a trigger frame. When the trigger frame identifier of a video frame is "1", it means that the video frame is a trigger frame.
[0161] The S605.Cloud Server 200 uses a segmentation mask based on video frames to determine salient regions in the video frames that include connected components. Connected components are used to indicate the shooting objects in the video frames, and salient regions are the smallest rectangular regions that include the shooting objects.
[0162] For example, cloud server 200 can retrieve a video frame at time t from the video stream of electronic device 100 from video frame buffer 24. This time t can be the time when the camera of electronic device 100 captures the video frame. For example, the video frame includes... Given 10 pixels, this video frame can be represented as a three-dimensional matrix F. t Among them, the three-dimensional matrix F t It can be represented as:
[0163] Where h represents the height of the video frame, w represents the width of the video frame, and c represents the number of channels in the video frame. This three-dimensional matrix contains a total of... There are 10 elements, each of which is an unsigned integer with a value between 0 and 255.
[0164] Cloud server 200 processes video frames based on a saliency segmentation algorithm to obtain a segmentation mask M for the video frames. t and confidence score S t For example, the confidence score S t It can be a floating-point number in the range of 0 to 1. Where, the segmentation mask M... t It can be represented as:
[0165] Among them, the segmentation mask M t is a two-dimensional matrix, and the height and width of the segmentation mask M t are the same as those of the video frame. The segmentation mask M t may include elements. For example, the value of each element in the segmentation mask M t is 0 or 1. If the value of the element at the coordinates (i, j) in M t is 1, it means that the pixel at the coordinates (i, j) in the video frame F t belongs to the foreground region. If the value of the element at the coordinates (i, j) in M t is 0, it means that the pixel at the coordinates (i, j) in the video frame F t belongs to the background region. Among them, 0 ≤ i < w, 0 ≤ j < h. Among them, the foreground region is the region where the object being photographed is located in the video frame. When a pixel belongs to the foreground region, it can indicate that the pixel belongs to the object being photographed. When a pixel belongs to the background region, that is, does not belong to the foreground region, it can indicate that the pixel does not belong to the object being photographed.
[0166] After the cloud server 200 determines the segmentation mask M t and the confidence score S t of the video frame, if the cloud server 200 determines that the confidence score S t of the video frame is greater than the first score threshold, it can determine the connected components in the segmentation mask M t . Among them, the cloud server 200 can divide adjacent elements with a value of 1 into the same connected component. The cloud server 200 can thereby determine one or more connected components in the segmentation mask M t . After that, the cloud server 200 can determine the salient region based on one or more connected components in the segmentation mask M t . The salient region is the region where the smallest rectangle enclosing the connected component is located. Among them, the connected component can be represented as C t = {C t,r |0 ≤ r ≤ u}, where r represents the number of the connected component and u represents the number of connected components. Among them, the salient region can be represented as B t = {B t,e = (x t,1 , y t,1 , x t,2 , y t,2 )|0 ≤ e ≤ l}, where e represents the number of the salient region and l represents the number of salient regions. Among them, u and l have the same value. Among them, x t,1 , y t,1 , x t,1 , y t,1These represent the x-coordinate of the top-left element, the y-coordinate of the top-left element, the x-coordinate of the bottom-right element, and the y-coordinate of the bottom-right element, respectively.
[0167] For example, such as Figure 7 As shown, the cloud server 200 is based on video frame F t The resulting segmentation mask M t It includes 100 elements. Among them, cloud server 200 can be based on this segmentation mask M. t Two connected components C were identified. t,1 and C t,2 Among them, C t,1 The salient region of the connected component is B. t,1 C t,2 The salient region of the connected component is B. t,2 .
[0168] S606. Cloud Server 200 detects, based on the salient regions of video frames, whether the subject in a video frame is the same as the subject in the m preceding video frames and different from the subject in the (m+1)th preceding video frame; or, whether the subjects in the (m+1) preceding video frames are all the same and whether the subject in this video frame is different from the subjects in the (m+1) preceding video frames.
[0169] Specifically, the cloud server 200 can execute step S607 if the subject in the video frame is the same as the subject in the m preceding video frames and different from the subject in the (m+1)th preceding video frame, or if the subjects in the (m+1)th preceding video frames are all the same and the subject in this video frame is different from the subject in the (m+1)th preceding video frames. Otherwise, the cloud server 200 can execute step S604. For example, the cloud server 200 can execute step S604 if the subject in the video frame is different from the subject in at least one of the (m+1)th preceding video frames and different from the subject in at least two of the (m+1)th preceding video frames, or if the subject in this video frame is the same as the subject in the (m+1)th preceding video frames.
[0170] In some examples, the value of m aligns with human attention mechanisms, reflecting the degree of user attention to a particular subject. For instance, if a subject appears in a video frame and is included in m consecutive video frames, it typically indicates that the subject is of interest to the user.
[0171] In some examples, the cloud server 200 can use an object re-identification algorithm to detect whether the subject in a video frame is the same as the subject in the m preceding video frames and different from the subject in the (m+1)th preceding video frame. It can also detect whether the subjects in the (m+1) preceding video frames are all the same and whether the subject in this video frame is different from the subjects in the (m+1) preceding video frames.
[0172] It should be noted that object re-identification algorithms can be used to re-identify the same subject in a video stream, identifying the same subject across multiple video frames. The cloud server 200 can assign the same re-identification number to the same subject in different video frames based on the object re-identification algorithm; this re-identification number can be understood as the subject's identifier. For example, the object re-identification algorithm could be a Kalman filter algorithm or other algorithms based on the subject's features. The cloud server 200 can assign different re-identification numbers to different salient regions containing different subjects. It is understood that if a salient region with a certain re-identification number does not appear in a video frame, the cloud server 200 will not use that re-identification number for a period of time.
[0173] In some examples, the object re-identification algorithm can assign a re-identification number to a salient region when a new salient region appears. It should be noted that the electronic device 100 may capture the same object even when it is moving.
[0174] For example, such as Figure 8 As shown, cloud server 200 determines the segmentation mask for the video frame at time (t-1) and the segmentation mask for the video frame at time t. The salient regions in the segmentation mask for the video frame at time (t-1) include B. t-1,1 The salient regions in the segmentation mask of the video frame at time t include B. t,1 With B t,2 Cloud server 200 can determine the significant region B at time t. t,1 The subjects in the image are compared with the subjects B in the salient region at time (t-1). t-1,1 If they are identical, assign the same re-identification number to the two salient regions. For example, salient region B. t-1,1 The re-identification number can be represented as N t-1,1 =2. Significant region B t,1 The re-identification number can be represented as N t,1 =2. The fact that the re-identification numbers of the two salient regions are the same indicates that the subjects captured in the two salient regions are the same.
[0175] After identifying the salient regions of a video frame and their re-identification numbers, the cloud server 200 can determine whether a video frame is a trigger frame. Furthermore, based on the salient regions and their re-identification numbers, the cloud server 200 can determine if a video frame is a trigger frame when the subject within the video frame changes.
[0176] The change in the subject of the video shot can include, but is not limited to, the appearance or disappearance of the subject. The appearance of the subject can mean that the (m+1)th video frame before this video frame does not include all or part of the subject in this video frame, and the subject in the first m video frames is the same as that in this video frame. The disappearance of the subject can mean that the subject in the first (m+1)th video frames is the same, and the subject in this video frame does not include all or part of the subject in the first (m+1)th video frames. For example, the (m+1)th video frame includes a subject within a saliency region with a re-identification number value of a first value, and the first and second values are different in the first m video frames and in the video frame itself.
[0177] Optionally, the cloud server 200 can sort the re-identification numbers of multiple salient regions in the video frame. The cloud server 200 can retain only the o largest salient regions in the video frame, where o is greater than or equal to 1. In this way, the cloud server 200 can ignore some smaller subjects in the video frame and focus on larger subjects, making it more likely to identify subjects that the user is more interested in.
[0178] For example, the cloud server 200 can take the symmetrical difference of the re-identification number sets of two adjacent video frames, that is, detect whether there are significant regions with different re-identification numbers in two adjacent video frames, and obtain the set N of non-repeating numbers. sym The set of re-identified numbers includes the re-identified numbers of all salient regions in the video frame, or the re-identified numbers of the o salient regions with the largest area. The set of non-repeating numbers N... sym This includes the re-identification numbers of salient regions where the re-identification numbers differ between the two video frames. When this N... sym When N is an empty set, it indicates that the subjects captured in the two video frames are the same. sym When the set is not empty, it indicates that the subjects captured by the two video frames are different. Similarly, the cloud server 200 can use this to determine whether the subjects captured by the two video frames are the same, and whether the video frame is a trigger frame.
[0179] It is understandable that if the subject of multiple consecutive video frames following the trigger frame is the same as the subject of the trigger frame, then these multiple video frames are not considered trigger frames because the subject of these multiple frames has not changed.
[0180] This allows for accurate identification of trigger frames in the video stream and avoids the electronic device 100 frequently outputting interactive data introducing the same subject.
[0181] It should be noted that the determination of whether the subjects in two video frames are the same is not limited to the above process. The electronic device 100 can also determine whether the subjects in two video frames are the same through other methods. For example, the electronic device 100 uses inertial sensors (e.g., inertial measurement units, accelerometers, gyroscopes, etc.). When the electronic device 100 detects that the distance moved by the electronic device 100 within a preset silent period is less than a preset distance through the inertial sensor, it can detect the similarity of pixels in multiple video frames within the preset silent period. When the similarity between pixels in two adjacent video frames is greater than a similarity threshold, it indicates that the subjects in the two video frames are the same. When the similarity between pixels in two adjacent video frames is less than a similarity threshold, it indicates that the subjects in the two video frames are different. Optionally, the electronic device 100 can detect the similarity of pixels in multiple video frames within the preset silent period through a cloud server 200. For another example, the electronic device 100 can use deep learning algorithms to identify the features of the subjects in the video frames, and determine whether the subjects in the video frames are the same by comparing whether the features of the subjects in two video frames are similar, etc. This application embodiment does not limit this.
[0182] S607. Cloud Server 200 marks this video frame as a trigger frame.
[0183] Specifically, the cloud server 200 can mark a video frame as a trigger frame if it determines that the subject in a video frame is the same as the subject in the m preceding video frames, but different from the subject in the (m+1)th preceding video frame; or, the subjects in the (m+1)th preceding video frames are all the same, but the subject in this video frame is different from the subjects in the (m+1)th preceding video frames. In some examples, Figure 6 The steps shown are by Figure 5 The trigger frame recognition module 21 shown is executed. Video frames and their specified information can be stored in the video frame buffer 24.
[0184] In some examples, after receiving input 1, electronic device 100 can send an active interaction request to cloud server 200 at preset intervals. Upon receiving the active interaction request, cloud server 200 can determine the trigger frame in the video stream of electronic device 100.
[0185] If the cloud server 200 determines that there are no trigger frames in the video stream of the electronic device 100, the cloud server 200 may not send interactive data to the electronic device 100. If the electronic device 100 does not receive interactive data from the cloud server 200 after sending an active interaction request, the electronic device 100 may send an active interaction request to the cloud server 200 again after a preset interval.
[0186] If the cloud server 200 determines that there is a trigger frame in the video stream of the electronic device 100, the cloud server 200 can determine the interaction data based on the trigger frame and send the interaction data to the electronic device 100. After receiving the interaction data, the electronic device 100 can output the interaction data. After outputting the interaction data, the electronic device 100 can send an active interaction request to the cloud server 200 at a preset interval, and so on. In this way, after the user opens the object recognition function provided by the electronic device 100, the electronic device 100 can obtain interaction data from the cloud server 200 through active interaction requests in active interaction mode.
[0187] Optionally, if the cloud server 200 determines that there is no trigger frame in the video stream of the electronic device 100, the cloud server 200 may send a no-interaction data response to the electronic device 100. This no-interaction data response can be used to notify the electronic device 100 that the cloud server 200 has not determined any interaction data, and the electronic device 100 can determine from the no-interaction data response that the cloud server 200 has successfully received the proactive interaction request.
[0188] For example, such as Figure 9 As shown, the electronic device 100 includes a smart assistant application 11. The smart assistant application 11 of the electronic device 100 can send proactive interaction requests (e.g., first request, second request, etc.) to the cloud server 200 at preset intervals of time Δt.
[0189] In this process, electronic device 100 sends an active interaction request to cloud server 200 at time t1. Upon receiving the active interaction request, cloud server 200 determines that the video stream of electronic device 100 has no trigger frames and does not send interaction data to electronic device 100.
[0190] Electronic device 100 sends another active interaction request to cloud server 200 at time t2. A preset interval Δt separates time t2 from time t1. Upon receiving this active interaction request, cloud server 200 determines that the video stream from electronic device 100 contains a trigger frame, determines the interaction data based on the trigger frame, and sends the interaction data to electronic device 100. After receiving the interaction data, electronic device 100 can output the interaction data. For example, electronic device 100 completes the output of the interaction data at time t3. The time difference between time t3 and time t2 is the response duration.
[0191] Electronic device 100 can send an active interaction request to cloud server 200 again at time t4, and so on. The time interval between time t4 and time t3 is a preset time interval Δt.
[0192] In other examples, electronic device 100 can determine the frequency of sending proactive interaction requests within a preset transmission duration after the current moment based on the frequency of receiving interactive data within a preset historical duration prior to the current moment. For details, please refer to [link to relevant documentation]. Figure 3 The illustrated embodiment.
[0193] In this way, the electronic device 100 can continuously acquire interactive data based on the video stream, providing users with intelligent services that allow them to identify objects while shooting, thus enhancing the user experience.
[0194] The following describes a video frame and its specified information stored on a cloud server 200, as provided in an embodiment of this application.
[0195] For example, such as Figure 10 As shown, cloud server 200 stores video frames and their specified information. The specified information for the video frames may include, but is not limited to, the video frame number and the trigger frame identifier.
[0196] In this context, a larger video frame number indicates a later time that the electronic device 100 acquired the video frame, or a later time that the video frame was uploaded to the cloud server 200. The cloud server 200 can determine the video frame number based on the time the camera of the electronic device 100 acquired the video frame, or it can determine the video frame number based on the order of the video frames in the video stream, and so on. The trigger frame identifier can be used to indicate whether a video frame is a trigger frame; a detailed description of the trigger frame identifier can be found in the above embodiments.
[0197] Optionally, the specified information for the video frame may also include the location information of salient regions. This location information can indicate the position of salient regions within the video frame. The cloud server 200 can then directly determine the location of the subject within the video frame based on the location information of the salient regions, and determine the interactive data for that video frame based on the image within the salient regions. In this way, the cloud server 200 can generate interactive data that better matches the subject based on the image within the salient regions, reducing interference from images outside the salient regions of the video frame.
[0198] Optionally, the specified information for the video frame may also include a used identifier. This used identifier can indicate whether the cloud server 200 has already generated interactive data based on the trigger frame. When the cloud server 200 has generated interactive data based on the trigger frame, the used identifier indicates that the trigger frame has been used by the cloud server 200; for example, the value of the used identifier can be 1. When the cloud server 200 has not generated interactive data based on the trigger frame, the used identifier indicates that the trigger frame has not been used by the cloud server 200; for example, the value of the used identifier can be 0. When the cloud server 200 generates interactive data based on a trigger frame that has not been used by the cloud server 200, it can combine this trigger frame with one or more trigger frames that have previously been used by the cloud server 200 to generate interactive data. In this way, the cloud server 200 can refer to the context to obtain interactive data that is more suitable for the current application scenario.
[0199] In some examples, after determining the trigger frame, the cloud server 200 can use the smallest rectangle that includes multiple salient regions within the trigger frame as the salient region of that trigger frame. For example, the cloud server 200 can use the smallest rectangle that includes o salient regions within the trigger frame as the salient region of that trigger frame. In this way, the cloud server 200 can merge the salient regions of multiple subjects within the trigger frame, requiring only a search of image information from one image, thus generating the corresponding interactive data more quickly.
[0200] It should be noted that when the cloud server 200 receives a new video frame from the first video stream of the electronic device 100, the cloud server 200 can process multiple video frames in the first video stream sequentially according to the order of the video frame numbers from smallest to largest. For example, the cloud server 200 can store the new video frame in the first video stream in the video frame buffer 24, and the trigger frame recognition module 21 can read the latest multiple video frames in the first video stream and determine the trigger frame based on these multiple video frames.
[0201] In some examples, after receiving an active interaction request, the trigger frame selection module 25 can determine the trigger frame with the largest video frame number based on the trigger frame identifier of the video frame, in descending order of video frame number. Thus, the larger the video frame number, the closer the video frame acquisition time is to the time the active interaction request is received, resulting in a more timely video frame.
[0202] In some examples, the trigger frame selection module 25 can detect whether the video frame with the largest of the k video frame numbers is the trigger frame. Here, k is greater than m. This allows for the identification of usable trigger frames while minimizing the resources used by the cloud server 200.
[0203] In other examples, cloud server 200 can delete video frames preceding the trigger frame after it is identified. This reduces the storage space occupied by video frames on cloud server 200.
[0204] In one possible implementation, after determining the first trigger frame, the cloud server 200 can detect whether the subject in the n video frames following the first trigger frame is the same as the first subject in the first trigger frame, where n is greater than or equal to 1. If the cloud server 200 determines that the subject in the n video frames following the first trigger frame is the same, the cloud server 200 can determine first comprehensive information based on the sixth video frame. The sixth video frame can be any of the n video frames, or the first trigger frame. The first comprehensive information may include detailed information about the first subject. The cloud server 200 can determine third interactive data based on historical communication data, the sixth video frame, the first comprehensive information, and a multimodal dialogue model. The amount of data in the third interactive data is greater than the amount of data in the first interactive data. The cloud server 200 can send the third interactive data to the electronic device 100. The electronic device 100 can output the third interactive data. In this way, when the cloud server 200 determines that the subject of multiple consecutive video frames after the trigger frame is the same as the subject of the trigger frame, it indicates that the camera of the electronic device 100 has been pointing at the subject of the trigger frame. The user may be inclined to obtain more information about the subject. The cloud server 200 determines that interactive data including more descriptions of the subject can better meet the user's needs.
[0205] Understandably, if the cloud server 200 determines that the subject of the p-th video frame following the first trigger frame is different from the first subject, where p is less than or equal to n, the cloud server 200 can continue to determine the trigger frame based on the first video stream provided by the electronic device 100.
[0206] In some examples, n is greater than m. This allows the cloud server 200 to more accurately determine the first subject that the electronic device 100 is continuously capturing in the first trigger frame.
[0207] The following is a flowchart illustrating an interaction method provided in an embodiment of this application.
[0208] For example, such as Figure 11 As shown in the figure, the flow of an interaction method provided in this application embodiment includes the following steps: S1101. After determining the first trigger frame, if it is determined that the shooting object of the first trigger frame is the same as the shooting object of the n video frames after the first trigger frame, the first comprehensive information is determined based on the sixth video frame.
[0209] The cloud server 200, after determining the first trigger frame, can detect whether the subject in the next n consecutive video frames is the same as the subject in the first trigger frame. If the cloud server 200 determines that the subject in the next n consecutive video frames is the same as the subject in the first trigger frame, it can determine the first comprehensive information based on the relevant information of the sixth video frame. This relevant information may include, but is not limited to, one or more of the following: semantic information of the video frame; facial expression information of the subject (e.g., laughing, crying); action information of the subject (e.g., waving, dancing); QR code information; text information of the video frame (i.e., text appearing in the video frame); name of the subject; identity information of the subject (e.g., doctor, firefighter); structural information of the subject; appearance information of the subject (e.g., shape, size, material, color); and purpose information of the subject (e.g., for warmth, lighting).
[0210] In some examples, the cloud server 200 can determine relevant information about the sixth video frame using multiple models. These models may include, but are not limited to, semantic analysis models, image recognition models, and saliency segmentation models. Specifically, the semantic analysis model can be used to determine the semantic information of the sixth video frame. The saliency segmentation model can be used to identify the subject being filmed in the sixth video frame. The image recognition model can be used to determine the QR code information, text information, and the name, expression, action, identity, result, and appearance information of the subject being filmed in the sixth video frame.
[0211] S1102. Based on historical communication data, first comprehensive information, sixth video frame, and multimodal dialogue model, determine the third interaction data.
[0212] Historical communication data may include, but is not limited to, interactive data sent to electronic device 100 stored in cloud server 200, and / or, dialogue data input by the user in electronic device 100 and its corresponding response data, and / or, basic user information (e.g., age, hobbies, etc.) stored in electronic device 100. In this way, cloud server 200 can generate third-party interactive data that is more tailored to user habits and preferences based on historical communication data.
[0213] In some examples, the cloud server 200 may also determine whether to determine third interaction data based on historical communication data and the first aggregated information. For example, the cloud server 200 may determine third interaction data when it detects that the first aggregated information includes a description of an object that the user likes according to the historical communication data. The cloud server 200 may not determine third interaction data when it detects that the first aggregated information does not include a description of an object that the user likes according to the historical communication data, and / or, the first aggregated information includes a description of an object that the user dislikes according to the historical communication data.
[0214] In other examples, cloud server 200 can determine third interaction data when it detects that the first aggregated information includes descriptions of objects the user likes / dislikes as indicated by historical communication data. Cloud server 200 can also determine third interaction data when it detects that the first aggregated information does not include descriptions of objects the user likes or dislikes as indicated by historical communication data. Thus, cloud server 200 can use third interaction data to prompt the user to photograph objects containing that disliked item when the user dislikes something (e.g., an ingredient), helping the user avoid disliked objects.
[0215] In some examples, the cloud server 200 can set the electronic device 100 to gaze mode when it determines that the subject of the first trigger frame is the same as the subject of the n video frames following the first trigger frame. When the electronic device 100 is in gaze mode, the cloud server 200 can determine the comprehensive information of the video frame based on the video frame and its related information, and determine the interaction data based on historical dialogue data, the first comprehensive information, the sixth video frame, and the multimodal dialogue model. After determining that the subject of the first trigger frame is the same as the subject of the n video frames following the first trigger frame, the cloud server 200 can set the electronic device 100 to exit gaze mode if it detects that the subject of a video frame is different from the subject of the first trigger frame. When the electronic device 100 is not in gaze mode, the cloud server 200 can determine the interaction data of the video frame based on the video frame.
[0216] For example, the cloud server 200 can perform text detection on the sixth video frame and detect whether the text in the sixth video frame is located in the salient region of the sixth video frame. If the text in the sixth video frame is detected to be located in the salient region of the sixth video frame, the text information can be obtained by recognizing the text within the salient region of the sixth video frame.
[0217] The cloud server 200 can determine the name information within textual information. This includes names of people, places, items, and organizations. The cloud server 200 can then perform a search based on this name information to obtain detailed information about the object indicated by that name.
[0218] The cloud server 200 can determine whether the object indicated by the name information is a user's favorite object based on the name information, the detailed description of the object indicated by the name information, and historical communication data. If the cloud server 200 detects that the object indicated by the name information is a user's favorite object, it can determine the interaction data based on historical dialogue data, first comprehensive information, sixth video frame, and multimodal dialogue model. The first comprehensive information may include, but is not limited to, one or more of the following: text information, name information, and the detailed description of the object indicated by the name information.
[0219] In some examples, after determining the object's descriptive information, such as text information, name information, and the specific description of the object indicated by the name information, the cloud server 200 can, based on historical communication data, detect whether the object's descriptive information includes descriptions of objects the user likes. If the cloud server 200 detects that the object's descriptive information includes descriptions of objects the user likes, or includes descriptions of objects the user dislikes, it can add the object's descriptive information to the first comprehensive information. Optionally, if the cloud server 200 detects that the object's descriptive information does not include descriptions of objects the user likes or dislikes, it can choose not to add the object's descriptive information to the first comprehensive information.
[0220] For example, if the text information in the sixth video frame includes the name of a restaurant, the restaurant's menu information (e.g., cuisine, signature dishes, average price per person) can be retrieved through knowledge retrieval. If historical interaction data determines that the user prefers the cuisine or signature dishes indicated by the restaurant's menu information, the restaurant's descriptive information can be added to the first comprehensive information set, and interactive data can be generated based on this first comprehensive information set. This interactive data can be used to introduce the restaurant's menu information. In this way, it can help users quickly determine the restaurant they want to dine at while shopping.
[0221] For example, the cloud server 200 can perform QR code detection on the sixth video frame. When the cloud server 200 detects that the sixth video frame includes a QR code, it can detect whether the QR code in the sixth video frame is located in a salient area of the sixth video frame. If the QR code in the sixth video frame is detected to be located in a salient area of the sixth video frame, the cloud server 200 can identify the QR code within the salient area of the sixth video frame and obtain the link information of the QR code. Based on the link information, the sixth video frame, and the multimodal dialogue model, the cloud server 200 can generate third interaction data and send the third interaction data to the electronic device 100. The third interaction data may include the link information and prompts describing the function of the link information.
[0222] Optionally, after determining the link information, the cloud server 200 can check whether the page indicated by the link information is safe. If the cloud server 200 detects that the page indicated by the link information is safe, it can determine third-party interactive data based on that link information. If the cloud server 200 detects that the page indicated by the link information is unsafe, it can either not generate interactive data based on the link information of the QR code, or generate third-party interactive data to prompt the user that the QR code is unsafe and should not be scanned. For example, the cloud server 200 can determine whether a page is safe by checking whether it belongs to a whitelist. Pages belonging to the whitelist are safe pages, and pages not belonging to the whitelist are not safe pages.
[0223] In some examples, cloud server 200 may add the link information to the first comprehensive information when it determines that the page indicated by the link information is safe.
[0224] It should be noted that if the cloud server 200 has already sent the link information of the sixth video frame to the electronic device 100 in the first interaction data, the cloud server 200 may no longer need to recognize the QR code information.
[0225] For example, if the first trigger frame is Figure 2D The video frame shown is 221. The cloud server 200 determines that the subject of the n video frames following the first trigger frame is the same as the subject of the first trigger frame. These n video frames include... Figure 2E The video frame 231 shown and Figure 12A The video frame shown is 1201. The cloud server 200 can determine the [specific information] based on the sixth video frame. Figure 12A The link information for QR code 1202 shown. Cloud server 200 can send third-party interactive data including this link information to electronic device 100. After receiving the third-party interactive data, electronic device 100 can display the link information, or display something like... Figure 12B The link information shown corresponds to interface 1211.
[0226] Optionally, the cloud server 200 can generate third-party interactive data based on the content of the page indicated by the link information. In this way, the electronic device 100 does not need to jump to the page indicated by the link information to understand the content of the page, keeping the electronic device 100 in the object recognition interface, which can help the user continue to use the active object recognition function provided by the electronic device 100.
[0227] For example, the cloud server 200 can perform object detection on the sixth video frame. When the cloud server 200 detects that the sixth video frame contains an object, it can detect whether the object is located within a salient region of the sixth video frame. If the object is detected to be located within a salient region of the sixth video frame, the cloud server 200 can identify the object within that region and obtain a detailed description of the object. Based on the detailed description of the object, the sixth video frame, and the multimodal dialogue model, the cloud server 200 can generate third interaction data and send this third interaction data to the electronic device 100. The third interaction data may include a detailed description of the object, etc.
[0228] For example, when a person is included in the salient region of the sixth video frame, it is possible to detect whether the sixth video frame includes images of the person's face, hands, or limbs.
[0229] When the sixth video frame includes a person's facial image, the cloud server 200 can detect the person's name and facial expression information based on the facial image. For example, the cloud server 200 can determine the person's name and facial expression information by comparing the person's facial image with facial images in the database.
[0230] When the sixth video frame includes an image of a person's hand, the cloud server 200 can perform gesture recognition on the image of the person's hand to determine the person's gesture information.
[0231] When the sixth video frame includes an image of a person's body, the cloud server 200 can identify the person's clothing, age, height, body type, etc. Optionally, the cloud server 200 can also combine multiple video frames before and / or after the sixth video frame to identify the person's movements. For example, if the person is dancing, it can identify the dance style (e.g., jazz, classical, etc.), the characteristics of the dance (e.g., gentleness, etc.), and the dance's signature movements (e.g., spins, etc.). If the person is exercising, it can identify the type of exercise (e.g., swimming, skating, etc.) and the technical movement the person is performing (e.g., a bent-over turn, a standing spin, etc.). If the person is working, it can identify the person's profession (e.g., doctor, etc.) and the operation the person is performing (e.g., bandaging, etc.), and so on.
[0232] The cloud server 200 can add the detailed information of an object to the first comprehensive information when it recognizes the detailed information of the object.
[0233] Understandably, if the number of video frames captured by electronic device 100 is insufficient, and cloud server 200 cannot recognize the person's movements, cloud server 200 can continue to receive video frames captured by electronic device 100. Once it has received a sufficient number of video frames including the first subject, it will then perform motion recognition. If, during motion recognition, the video frames from electronic device 100 no longer include the first subject, cloud server 200 can cancel the motion recognition operation and instead perform an operation to determine the trigger frame based on the video frames from electronic device 100.
[0234] Optionally, the cloud server 200 can identify only the image within the salient region of the sixth video frame to obtain the first comprehensive information. In this way, the cloud server 200 does not need to determine whether text information, QR codes, objects, etc. in the sixth video frame are within the salient region.
[0235] S1103. Send the third interactive data to the electronic device 100.
[0236] In some examples, cloud server 200 can store the third interaction data determined based on historical dialogue data, first comprehensive information, sixth video frame, and a multimodal dialogue model. Upon receiving an active interaction request from electronic device 100, cloud server 200 can detect whether it stores the third interaction data determined based on historical dialogue data, first comprehensive information, sixth video frame, and multimodal dialogue model. If it determines that the third interaction data is stored, cloud server 200 can send the stored third interaction data to electronic device 100.
[0237] When the cloud server 200 determines that the third interactive data is not stored, it can determine whether the video stream of the electronic device 100 includes a trigger frame. If the video stream includes a trigger frame, the cloud server 200 can send interactive data determined based on the trigger frame to the electronic device 100. This provides users with interactive data that better meets their needs and also saves communication resources between the cloud server 200 and the electronic device 100.
[0238] Optionally, when the cloud server 200 determines that the seventh video frame after the sixth video frame is the trigger frame, regardless of whether the cloud server 200 stores the third interactive data, it can generate the fourth interactive data based on the seventh video frame and send the fourth interactive data to the electronic device 100.
[0239] Optionally, the cloud server 200 stores third interaction data, which can be deleted when the seventh video frame is determined to be the trigger frame.
[0240] In some examples, the sixth video frame is any one of the n video frames (e.g., the video frame with the latest reception time among the n video frames). When the cloud server 200 determines that the sixth video frame has corresponding first comprehensive information, it can determine the third interaction data based on historical dialogue data, the first comprehensive information, the sixth video frame, and the multimodal dialogue model. The cloud server 200 can mark the sixth video frame as a trigger frame and store the third interaction data; the sixth video frame corresponds to the third interaction data. After receiving an active interaction request from the electronic device 100, the cloud server 200 can detect whether the selected trigger frame includes the corresponding interaction data. The selected trigger frame can be one or more trigger frames whose acquisition time is closest to the time the active interaction request was received. When the cloud server 200 detects that the selected trigger frame includes the corresponding interaction data, it can send the corresponding interaction data to the electronic device 100. When the cloud server 200 detects that the selected trigger frame does not include the corresponding interaction data, it can generate interaction data based on the trigger frame and send the generated interaction data to the electronic device 100. For example, when the trigger frame selected by the cloud server 200 is the sixth video frame, the sixth video frame includes the corresponding third interactive data, and the cloud server 200 can send the third interactive data to the electronic device 100.
[0241] The cloud server 200 can mark the sixth video frame as not a trigger frame when it determines that the sixth video frame does not have the corresponding first comprehensive information. It can be understood that if the shooting objects of multiple video frames following the sixth video frame are the same as those of the sixth video frame, the cloud server 200 can mark these multiple video frames as not trigger frames.
[0242] In this way, the electronic device 100 can determine the user's habits and preferences based on historical communication data. When the comprehensive information includes content related to the user's habits and preferences, the electronic device 100 can determine the interaction data based on the comprehensive information, helping the user obtain more interesting interaction data. The electronic device 100 can use the interaction method provided in this application embodiment to determine the interaction data based solely on the visual information of video frames, realizing an active interaction process without semantic information. Alternatively, it can determine the interaction data based on both the visual and semantic information of video frames, realizing an active interaction process with semantic information, providing the user with a richer interactive experience.
[0243] In one possible implementation, the cloud server 200 determines that the state of the subject changes within n video frames. After determining the aforementioned third interactive data, the cloud server 200, without changing the subject, determines fifth interactive data based on multiple eighth video frames at preset change intervals. The subject in the eighth video frames is the same as the subject in the first trigger frame. The fifth interactive data can be used to indicate changes in the subject's state. For example, when the subject includes a light source, the state change may include, but is not limited to, light flickering or color changes. When the subject is a person, the state change may include, but is not limited to, changes in the subject's actions, expressions, clothing, gestures, or position, etc. Thus, after determining the first trigger frame, if the subject in multiple subsequent video frames is the same as the subject in the first trigger frame, and the subject's state changes, the cloud server 200 can also output interactive data to prompt the user about the change in the subject's state. For example, the electronic device 100 can proactively and continuously provide commentary on sports events, dance competitions, etc., to the user by outputting the third and fifth interactive data.
[0244] In some examples, after being set to gaze mode, the electronic device 100 can continuously capture images of the subject in the first trigger frame using an algorithm. In this way, the electronic device 100 can continuously focus on the subject of interest to the user, providing the user with more information about that subject.
[0245] In some examples, upon receiving input 2, electronic device 100 can, in response to input 2, capture a video stream via a camera and determine interactive data based on the video stream. Electronic device 100 can then output the interactive data. In this way, electronic device 100 can output interactive data directly through the video stream captured by the camera without needing to display a first interface, making it convenient for users to implement the interactive method provided in this application embodiment when it is inconvenient for them to look at a display screen, or when electronic device 100 does not have a display screen.
[0246] For example, when electronic device 100 is a smart glasses device without a display screen, electronic device 100 can, upon receiving input 2, capture a video stream via its camera in response to input 2. Electronic device 100 can determine interactive data based on the video stream; specifically, refer to the above embodiments. Electronic device 100 can broadcast the interactive data. In this way, since the shooting range of the smart glasses device's camera is roughly the same as the field of vision of the user wearing the smart glasses device, when the user is looking at an object, the camera of the smart glasses device can also capture that object. The smart glasses device can achieve the effect of broadcasting the description of the object the user is looking at, providing the user with a more intelligent object recognition experience.
[0247] In some examples, electronic device 100 can provide obstacle avoidance functionality. Electronic device 100 can identify road signs (e.g., traffic lights, zebra crossings, etc.) and obstacles (e.g., roadside trees, roadblocks, etc.) through real-time video streams captured by a camera. Electronic device 100 can proactively output interactive data to alert the user to road signs, obstacles, etc. In this way, electronic device 100 can assist the user in traveling.
[0248] In some examples, the electronic device 100 also includes a motor. The electronic device 100 can control the motor vibration based on the interactive data when outputting interactive data. In this way, the electronic device 100 can also combine tactile feedback to provide the user with a more diverse experience through multiple senses.
[0249] For example, when the electronic device 100 determines, through the interaction method provided in this application embodiment, that an object is rapidly moving towards the camera in a video frame, it outputs interactive data to prompt the user to avoid it and controls the motor to vibrate. In this way, vibration can help the user perceive the approaching danger more quickly and avoid it more rapidly.
[0250] For example, when the electronic device 100 determines that a natural disaster has occurred in a video frame using the interaction method provided in this application embodiment, it outputs interactive data to prompt the user on how to evacuate safely and controls the motor to vibrate. Optionally, the electronic device 100 can also determine a safe zone based on video frames captured by a camera and output interactive data to prompt the user to go to the safe zone. This can help users leave dangerous areas more quickly, improve user safety, and so on.
[0251] In some examples, electronic device 100 can provide multiple styles of interactive data (e.g., humorous, gentle, straightforward, etc.). After receiving input from the user setting the style of the interactive data, electronic device 100 can instruct cloud server 200 to generate interactive data in the user-defined style based on the video stream. In this way, electronic device 100 can proactively introduce the subject captured by its camera to the user in a language style that the user prefers, thus enhancing the user experience.
[0252] In some examples, electronic device 100 or cloud server 200 can generate images or animations based on interactive data. Electronic device 100 can then display these images or animations. For instance, electronic device 100 or cloud server 200 can generate animations or images using artificial intelligence-generated content (AIGC). In this way, electronic device 100 can more vividly and engagingly display part or all of the interactive data through interesting images / animations. For example, when a user is unhappy, electronic device 100 can proactively generate and display humorous images, animations, or short videos to help improve the user's mood.
[0253] In some examples, electronic device 100 can be connected to pan-tilt device 300 (not shown in the figure). After recognizing the subject, electronic device 100 can control pan-tilt device 300 to re-recognize the subject, ensuring that the subject remains centered on the display screen. This allows for real-time analysis based on the subject's actual movements, yielding interactive data more relevant to the context. Examples include broadcasting sports events or playing rock-paper-scissors.
[0254] In some examples, electronic device 100 can establish a communication connection with an electronic device equipped with a camera, and acquire the video stream captured by the camera through this communication connection. In this way, electronic device 100 can implement the interactive method provided in the embodiments of this application even when it does not have a camera, or when the camera is damaged.
[0255] It should be noted that the interaction data is not limited to sending the video stream to the cloud server 200 for processing. The electronic device 100 can also execute the process described above, whereby the cloud server 200 determines the interaction data based on the video stream, and independently determine and output the interaction data. This application embodiment does not limit this approach.
[0256] In one possible implementation, the electronic device 100 is not limited to determining the interaction data based on a first video stream acquired by a camera. The electronic device 100 can also determine the interaction data based on a stored video file / image file using the interaction method provided in this application embodiment. Optionally, the electronic device 100 can play the video file and output the interaction data corresponding to the video frame when displaying the video frame.
[0257] In some examples, after receiving input 3 from a user uploading a video file / image file, electronic device 100 can, in response to input 3, obtain one or more interactive data based on the video file / image file. Electronic device 100 can then output this one or more interactive data. In this way, electronic device 100 can proactively initiate a dialogue with the user based on the video file / image file.
[0258] For example, electronic device 100 can obtain a video file containing one or more interactive data based on a video file. In this way, electronic device 100 can play the video file containing one or more interactive data, and users can learn about some of the subjects captured in the video file by viewing the video file.
[0259] In other examples, after receiving input 4 from a user indicating that they are playing a video file (or image file), the electronic device 100, in response to that input 4, can initiate a chat conversation with the user about the first content in the video file (or image file) in a no-intent-driven interactive mode, based on the recognition of the first content. Thus, the electronic device 100 can proactively initiate a conversation with the user based on the file being played while playing a video file (or image file).
[0260] For example, electronic device 100 can receive user input... Figure 2B After input is added to control 216 as shown, in response to that input, the following is displayed: Figure 12C One or more controls as shown. For example... Figure 12C As shown, the one or more controls include one or more of the following: image control 1221, camera control, file control, etc.
[0261] The electronic device 100 can, upon receiving user input to the image control 1221, display one or more videos and / or images in response to that input. The one or more videos and / or images may include a video file 1231. The electronic device 100 can, upon receiving user input to the video file 1231, display, as shown in the image control 1221. Figure 12D The interface shown is 1230. (As shown) Figure 12D As shown, interface 1230 may include video file 1231, icon 1232, and send control 1233. Icon 1232 indicates that video file 1231 has been selected by the user.
[0262] Electronic device 100 can receive user input Figure 12D After input is received into the send control 1233 shown, in response to that input, the following is displayed: Figure 12E The dialog interface 210 is shown. (As shown) Figure 12E As shown, the dialog interface 210 includes a dialog box 1241. The dialog box 1241 indicates that the smart application assistant application has received the video file 1231.
[0263] Electronic device 100 can obtain video file 1235 (not shown in the figure) based on video file 1231. Video file 1235 includes all video content of video file 1231 and includes one or more interactive data. After obtaining video file 1235, electronic device 100 can display, as shown in the figure... Figure 12FThe dialog interface shown is 210.
[0264] like Figure 12F As shown, the dialog interface 210 may include a dialog box 1242. The dialog box 1242 may indicate that the smart assistant application has obtained the video file 1235. The dialog box 1242 may be used to trigger the electronic device 100 to play the video file 1235. Optionally, Figure 12F The dialog interface 210 shown may also include a dialog box 1243, which can be used to inform the user of the purpose of the video file 1235. For example, the dialog box 1243 may include a text-based prompt: "This is the video I've processed. Click to play the video and view the interactive content!"
[0265] Electronic device 100 can receive user input Figure 12F After input is received in dialog box 1242, the following will be displayed in response to that input: Figure 12G The video playback interface shown is 1250.
[0266] like Figure 12G As shown, the video playback interface 1250 is playing video file 1235. The electronic device 100 can output interactive data 1252 when playing video frame 1251 of video file 1235. The interactive data 1252 can include the introductory content of video frame 1251. For example, interactive data 1252 could include: "This athlete is playing golf. Want to know more about golf?". It is understood that the electronic device 100 can continue playing video file 1235 and output interactive data when playing a trigger frame in video file 1235. Optionally, the video playback interface 1250 may also include one or more of the following: subtitle control 1253, broadcast control 1254, etc. The subtitle control 1253 can be used to trigger the electronic device 100 to turn subtitles on / off. When the electronic device 100 turns on subtitles, it can display the video audio and / or interactive data in text form. The broadcast control 1254 can be used to trigger the electronic device 100 to turn on / off the voice broadcast function. When the electronic device 100 turns on the voice broadcast function, it can broadcast interactive data via voice.
[0267] In some examples, electronic device 100 displays as follows Figure 12H The video playback interface shown is 1260. (As shown...) Figure 12HAs shown, the video playback interface 1260 may include, but is not limited to, a video file 1261, settings 1262, and playback controls 1263. The playback controls 1263 can be used to trigger the electronic device 100 to play the video file 1261. The settings 1262 can be used to enable / disable the active interaction function. When the electronic device 100 enables the active interaction function, it can output interactive data based on the video file 1261 during playback. When the electronic device 100 disables the active interaction function, it cannot output interactive data based on the video file 1261 during playback. A description of settings 1262 can be found in [link to documentation]. Figure 2D The description of setting item 222 shown will not be repeated here.
[0268] Upon receiving user input to the playback control 1263, the electronic device 100 can respond to the input by playing the video file 1261 and output interactive data during the playback of the video file 1261. For example, a description of the electronic device 100 playing the video file 1261 and outputting interactive data during playback can be found in [reference needed]. Figure 12G The description of the video playback interface 1250 shown will not be repeated here.
[0269] In this way, when the electronic device 100 plays the stored video file in active interaction mode, it can also actively initiate a dialogue with the user.
[0270] Figure 13 A schematic diagram of the structure of the electronic device 100 is shown.
[0271] The following description uses electronic device 100 as an example to illustrate the embodiment. It should be understood that... Figure 13 The electronic device 100 shown is merely an example, and the electronic device 100 may have more than Figure 13 The more or fewer components shown can be combined into two or more components, or they can have different component configurations. The various components shown in the figure can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0272] For example, such as Figure 13 As shown, the electronic device 100 may include, but is not limited to, a processor 110, an internal memory 121, an antenna 1, an antenna 2, a wireless communication module 160, an audio module 170, a speaker 170A, a microphone 170C, a camera 193, etc.
[0273] Optionally, the electronic device 100 may also include one or more of the following: an external memory interface 120, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, a mobile communication module 150, a receiver 170B, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a display screen 194, and a subscriber identification module (SIM) card interface 195. The sensor module 180 may include one or more of the following: a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a proximity sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, and a bone conduction sensor 180M.
[0274] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0275] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0276] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0277] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0278] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0279] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0280] The charging management module 140 receives charging input from the charger. The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and supplies power to the processor 110, internal memory 121, external memory, display 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0281] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0282] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0283] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0284] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0285] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0286] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0287] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0288] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD). The display panel can also be manufactured using organic light-emitting diodes (OLEDs), active-matrix organic light-emitting diodes (AMOLEDs), flexible light-emitting diodes (FLEDs), miniled, microled, micro-OLEDs, quantum dot light-emitting diodes (QLEDs), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0289] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0290] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimizations on image noise, brightness, etc. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0291] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0292] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0293] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0294] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs can enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0295] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.
[0296] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0297] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0298] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0299] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.
[0300] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.
[0301] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0302] In some examples of this application, microphone 170C can be used to receive user voice input. Microphone 170C can be used to broadcast interactive data.
[0303] The headphone jack 170D is used to connect wired headphones. The pressure sensor 180A senses pressure signals and converts them into electrical signals. The gyroscope sensor 180B is used to determine the motion posture of the electronic device 100. The barometric pressure sensor 180C measures air pressure. The magnetic sensor 180D includes a Hall effect sensor. The accelerometer sensor 180E detects the magnitude of acceleration of the electronic device 100 in various directions (typically three axes). The distance sensor 180F measures distance. The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a photodetector, such as a photodiode. The LED may be an infrared LED. The ambient light sensor 180L senses ambient light intensity. The fingerprint sensor 180H collects fingerprints. The temperature sensor 180J detects temperature. The touch sensor 180K, also called a "touch panel," can be located on the display screen 194, forming a touchscreen, also called a "touchscreen." The bone conduction sensor 180M acquires vibration signals. Buttons 190 include a power button, volume buttons, etc. Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to different touch operations for different applications (such as taking photos, playing audio, etc.). Indicator 192 can be an indicator light. SIM card interface 195 is used to connect a SIM card.
[0304] Figure 14 This is a schematic diagram of the structure of the cloud server 200 provided in the embodiments of this application.
[0305] like Figure 14 As shown, the cloud server 200 may include components such as a processor 1410, a memory 1420, and a communication module 1430. These components can be connected via a bus 1440 or other means. Figure 14 Taking a bus connection as an example, bus 1440 is used to realize the connection and communication between processor 1410, memory 1420, and communication module 1430. Wherein: Processor 1410 may include one or more processing units. Processor 1410 can be used to provide computing and control capabilities to support the operation of the entire cloud server 200.
[0306] The memory 1420 can be used to store various software programs and / or multiple sets of instructions. Specifically, the memory 1420 may include high-speed random access memory, and may also include non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices.
[0307] The communication module 1430 can be used to communicate with other communication devices. Specifically, the communication module 1430 can be used for the cloud server 200 to connect to a communication network. The cloud server 200 can communicate with electronic devices 100 in the communication network through the communication module 1430. For example, the communication module 1430 may include a communication interface, which may include a wired interface and / or a wireless interface. The wireless interface may include, but is not limited to, one or more of the following: a 3G communication interface, a Long Term Evolution (LTE) (4G) communication interface, a 5G communication interface, a WLAN communication interface, and a WAN communication interface.
[0308] In this embodiment, the communication module 1430 can be used to receive video streams sent by other electronic devices (e.g., electronic device 100) and to send interactive data determined based on the video streams to other electronic devices. Optionally, the communication module 1430 can also be used to receive active interaction requests sent by other electronic devices. The processor 1410 can be used to determine interactive data based on the video streams. The memory 1420 can be used to store video streams sent by other electronic devices, as well as software or program code (e.g., multimodal dialogue large model 23, image search module 22, etc.) required for all or part of the functions of the cloud server 200 in the above method embodiments.
[0309] It should be noted that, Figure 14 The cloud server 200 shown is merely one implementation of the embodiment of this application. In actual applications, the cloud server 200 may include more or fewer components than shown, or combine certain components, or deploy different components. No limitation is made here.
[0310] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the above-described method embodiments.
[0311] This application also provides a computer program product, including a computing program, which, when run on a processor, can implement the steps in the above-described method embodiments.
[0312] This application also provides a chip system, which includes a processing circuit interface circuit. The interface circuit receives code instructions and transmits them to the processing circuit. The processing circuit executes the code instructions to enable the chip system to implement the steps of any method embodiment of this application. The chip system can be a single chip or a chip module composed of multiple chips.
[0313] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. An interaction method applied to a chat application, the chat application running on an electronic device, characterized in that, The method includes: In response to the user's first operation on the electronic device, or if the chat application does not receive chat input from the user within a preset time period, the chat application enters an unintentional active interaction mode, which is an interaction mode in which the chat application actively initiates a chat conversation with the user on a specific topic. In the unintentional proactive interaction mode, when the video stream displayed on the electronic device meets the first condition or the second condition, the chat application proactively initiates a first conversation with the user. The first condition is that m consecutive video frames in the video stream include a first shooting object, which is an expression, action, text, or QR code; The second condition is that m consecutive video frames in the video stream include a first shooting object, and the first shooting object is the object with the largest area in the m video frames, or the most central object, or the object closest to the user; The video stream originates from real-time captured footage by the camera of the electronic device or from a video file stored on the electronic device. The topic of the first dialogue is related to the first subject being filmed. m is an integer greater than 1.
2. The method according to claim 1, characterized in that, The method further includes: In the unintentional active interaction mode, in response to the chat application receiving the user's chat input, the unintentional active interaction mode is exited.
3. The method according to claim 1 or 2, characterized in that, The chat input is voice input, and the audio characteristics of the voice input match the user's audio characteristics, or the sound source direction of the voice input is the back of the camera's shooting direction.
4. The method according to any one of claims 1-3, characterized in that, The method further includes: In response to the first subject being present in the video stream for a duration exceeding a threshold, a second dialogue about the first subject is output to the user, wherein the description of the first subject in the second dialogue is more detailed than the description of the first subject in the first dialogue.
5. The method according to any one of claims 1-4, characterized in that, The electronic device also displays at least one of the following: A first control used to initiate or exit the unintentional active interaction mode; A second control used to enable or disable voice input functionality; A third control used to enable or disable the voice broadcast function; A fourth control used to turn the camera on or off; The fifth control used to start or stop text interaction functionality.
6. An electronic device, characterized in that, include: One or more processors, one or more memories, and a camera, wherein the camera, the one or more memories, and the one or more processors are coupled together, the one or more memories being used to store a computer program, and when the one or more processors execute the computer program, performing the interactive method as described in any one of claims 1-5.
7. A computer-readable storage medium, characterized in that, The system contains a computer program that, when executed by a processor, implements the interactive method as described in any one of claims 1-5.
8. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the interactive method as described in any one of claims 1-5.