Online human-computer interaction method, system, electronic device, storage medium and program product
By introducing audio and video interaction capabilities into the intelligent dialogue assistant, the problems of low efficiency and insufficient visual information perception in the existing text-based dialogue interaction methods are solved, enabling more natural and smooth multimodal interaction and improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
- Filing Date
- 2025-08-28
- Publication Date
- 2026-07-21
AI Technical Summary
Existing intelligent dialogue assistants struggle to meet user needs when processing complex image information and multi-view detection using text-based dialogue, resulting in low interaction efficiency and a lack of real-time visual information perception capabilities.
By introducing audio and video interaction capabilities, the system collects video streams and voice dialogue streams from the user side, generates prompts, guides images, describes text information, generates echo text, and plays echo voice, thus achieving multimodal interaction.
It improves the natural and smooth interaction between users and the intelligent dialogue system, simplifies user operations, and enhances the user experience, especially when uploading complex image information or performing multi-view detection.
Smart Images

Figure CN122432397A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese Patent Application No. 202511221891.8, filed on August 28, 2025, entitled "Online Human-Computer Interaction Method, System, Apparatus, Electronic Device, Storage Medium and Program Product". Technical Field
[0002] This specification relates to the field of artificial intelligence technology, and in particular to an online human-computer interaction method, system, device, electronic device, storage medium, and program product. Background Technology
[0003] Intelligent dialogue assistants, capable of interactive conversations with users, aim to help them solve problems or engage in casual conversation, and have been widely adopted in various fields such as healthcare. Currently, intelligent dialogue assistants primarily use text-based dialogue as their main interaction method. With the development of multimodal technology, some intelligent dialogue assistants have gradually introduced image-text dialogue capabilities, allowing users to upload images when asking questions, thereby combining the image content with a more accurate understanding and response to the user's inquiry. However, when the image information that users need to upload is complex, such as involving multiple images, videos, or even requiring real-time feedback, this image-text dialogue interaction method becomes insufficient to meet user needs.
[0004] Therefore, there is an urgent need to introduce a new interactive dialogue capability for intelligent dialogue assistants. Summary of the Invention
[0005] Several embodiments in this specification provide an online human-computer interaction method, system, device, electronic device, storage medium, and program product, which introduce audio and video interaction capabilities into intelligent dialogue assistants, facilitating the fulfillment of users' more complex question needs. Among them, In a first embodiment, this specification provides an online human-computer interaction method. The method includes: The system collects video streams and voice dialogue streams from the user side during online human-computer audio and video interactions; wherein the video streams contain multi-angle images of a target object, including parts of the user's body; Based on the video stream and the voice dialogue stream, prompt words are generated; the prompt words are used to guide the generation of detection and analysis text for the target object based on the image description text information of the video stream, and the detection and analysis text includes the analysis results of the user's body part status. Based on the prompt words, a response text is generated to respond to the voice dialogue stream, and the response text includes the detection and analysis text; Display the conversation text and / or play the corresponding conversation audio.
[0006] In a second embodiment, this specification also provides an online human-computer interaction method. This method includes: Displays the service interface used for human-computer interaction; In response to the audio-visual interaction startup operation triggered through the service interface, online human-computer audio-visual interaction is initiated; During the online human-computer audio and video interaction process, video streams and voice dialogue streams are collected from the user side; wherein, the video stream contains multi-angle images of the target object, and the target object includes parts of the user's body; The video stream and voice dialogue stream are sent to the server. The system receives the response text and corresponding voice message returned by the server; wherein the response text is generated based on prompt words to respond to the voice dialogue stream; the prompt words are generated based on the video stream and the voice dialogue stream to guide the generation of detection and analysis text for the target object based on the image description text information of the video stream, and the detection and analysis text includes the user's body part state analysis results; the response text includes the detection and analysis text. While the audio of the conversation is played, the text of the conversation is displayed.
[0007] In a third embodiment, this specification also provides an online human-computer interaction method. This method includes: Display a health service interface, which has audio and video interactive controls; In response to an operation on the audio-visual interaction control, an online human-computer audio-visual interaction is initiated; During the online human-computer audio and video interaction process, the user's biological video stream and voice dialogue stream are collected; wherein, the biological video stream is a continuous image sequence containing the user's body parts; Based on the biological video stream and the voice dialogue stream, prompt words are generated; The prompt words are input into a third preset model, and the third preset model is executed to output a response text; the response text is used to respond to the user's inquiry about health issues, and the health issues are extracted from the voice dialogue stream; Play the corresponding audio message to the user.
[0008] Fourth embodiment: This specification provides a service system. The system includes: The client is used to display a service interface for human-computer interaction; respond to an audio-visual interaction start operation triggered by the service interface to start online human-computer audio-visual interaction; during the online human-computer audio-visual interaction, it collects video streams and voice dialogue streams from the user side; and sends the video streams and voice dialogue streams to the server; wherein the video stream contains multi-angle images of a target object, and the target object includes parts of the user's body; The server is configured to generate prompt words based on the video stream and the voice dialogue stream; the prompt words are used to guide the generation of detection and analysis text for the target object based on the image description text information of the video stream, the detection and analysis text including the user's body part state analysis results; based on the prompt words, generate response text for responding to the voice dialogue stream, the response text including the detection and analysis text; convert the response text into response speech, and send the response text and the response speech to the client; The client is also used to play the conversation audio and display the conversation text.
[0009] Fifthly, this specification provides an online human-computer interaction device. The device includes: The acquisition module is used to acquire video streams and voice dialogue streams from the user side during online human-computer audio and video interaction; wherein, the video stream contains multi-angle images of the target object, and the target object includes parts of the user's body; The generation module is used to generate prompt words based on the video stream and the voice dialogue stream; the prompt words are used to guide the generation of detection and analysis text of the target object based on the image description text information of the video stream, and the detection and analysis text includes the user's body part state analysis results; The generation module is further configured to generate a response text for responding to the voice dialogue stream based on the prompt words, the response text including the detection and analysis text; The playback display module is used to display the echo text and / or play the echo voice corresponding to the echo text.
[0010] In a sixth embodiment, this specification also provides an online human-computer interaction device. The device includes: The display module is used to display the service interface for human-computer interaction. The startup module is used to respond to the audio and video interaction startup operation triggered through the service interface and start the online human-computer audio and video interaction; The acquisition module is used to acquire video streams and voice dialogue streams from the user side during the online human-computer audio and video interaction process; wherein, the video stream contains multi-angle images of a target object, and the target object includes parts of the user's body; The sending module is used to send the video stream and the voice dialogue stream to the server. A receiving module is configured to receive the response text and corresponding voice message returned by the server; wherein the response text is generated based on prompt words to respond to the voice dialogue stream; the prompt words are generated based on the video stream and the voice dialogue stream to guide the generation of detection and analysis text of the target object based on the image description text information of the video stream, and the detection and analysis text includes the user's body part state analysis results; the response text includes the detection and analysis text. The playback display module is used to display the playback text while playing the voice message.
[0011] In a seventh embodiment, this specification also provides an online human-computer interaction device. The device includes: The display module is used to display the health service interface, which has audio and video interactive controls. The startup module is used to respond to operations on the audio and video interaction controls and start online human-computer audio and video interaction; The acquisition module is used to acquire the user's biological video stream and voice dialogue stream during the online human-computer audio and video interaction process; wherein, the biological video stream is a continuous image sequence containing the user's body parts; The generation module is used to generate prompt words based on the biological video stream and the voice dialogue stream; An execution module is used to input the prompt words into a third preset model, execute the third preset model to output response text; the response text is used to respond to the user's inquiry about health issues, and the health issues are extracted from the voice dialogue stream; The playback module is used to play the corresponding voice message to the user.
[0012] Eighth embodiment: This specification provides an electronic device including a memory and a processor, wherein the memory stores executable program instructions, and when the processor executes the program instructions, it implements the methods provided in the first to third embodiments described above.
[0013] In the ninth embodiment, this specification provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, it causes the computer to perform the methods provided in the first to third embodiments described above.
[0014] In a tenth embodiment, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the methods provided in the first to third embodiments described above.
[0015] The solutions provided in the above embodiments of this specification generate prompt words based on the collected video stream and voice dialogue stream on the user side during online human-computer audio-visual interaction. Then, based on these prompt words, they generate response text to the voice dialogue stream and output this response text to the user. Therefore, this solution implements online human-computer audio-visual interaction functionality, allowing users to communicate more naturally and smoothly with the intelligent dialogue system (also known as an intelligent dialogue assistant) in a multimodal manner. This is especially beneficial when uploading complex image information or performing multi-view detection (such as hair state detection tasks), simplifying user operations and improving user experience. The aforementioned online human-computer audio-visual interaction is triggered by a service interface for human-computer interaction. Specifically, this service interface has audio-visual interaction controls, which users can use to trigger the online human-computer audio-visual interaction. For example, this service interface could be the health service interface provided by an intelligent dialogue assistant for health services. Furthermore, when outputting response text to the user, the response text can be displayed while playing the corresponding voice message, facilitating user understanding of the content in the voice message. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the various embodiments disclosed in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely examples of the various embodiments disclosed in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort. In the drawings: Figure 1 A schematic diagram of the technical architecture on which the implementation of each method in this specification is based, provided for an exemplary embodiment; Figure 2A and Figure 2B A schematic diagram of the structure of a service system (specifically an online human-computer interaction system, i.e., an intelligent dialogue system) provided for exemplary embodiments in this specification; Figure 3 , Figure 4 and Figure 5 A flowchart illustrating the online human-computer interaction method provided for exemplary embodiments in this specification; Figure 6 , Figure 7 and Figure 8 A schematic diagram of the structure of the online human-computer interaction device provided in the exemplary embodiments of this specification; Figure 9 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this specification. Detailed Implementation
[0017] With the rapid development of artificial intelligence and mobile internet technologies, intelligent dialogue assistants have gradually become important tools for human-computer interaction in various applications. Intelligent dialogue assistants can interact and converse with users, their mission being to help users solve problems or chat, and they have been widely used in various fields such as healthcare. Taking health-related intelligent dialogue assistants (such as health managers provided by some applications) as an example, they are dedicated to helping users solve various problems before, during, and after medical treatment, covering multiple scenarios such as health consultation, initial disease screening, medical advice, medication reminders, and rehabilitation guidance, significantly improving users' health management efficiency and quality of life. Currently, health-related intelligent dialogue assistants, or other types of intelligent dialogue assistants, primarily support text-based interaction. For example, users can describe their health status by inputting text, and the health-related intelligent dialogue assistant can use its built-in language model to generate corresponding health advice or medical guidance. With the development of multimodal technology, in recent years, more and more intelligent dialogue assistants have added the ability to engage in image-text dialogue interaction, supporting users to upload images when asking questions, using this information to understand and generate responses based on the image content, thereby improving the intuitiveness of the interaction and the accuracy of the service. This text-based interactive interface allows users to upload images by clicking the "Image Upload" control on the screen. This is effective and convenient when uploading a small number of images. However, when the image information to be uploaded is complex, such as requiring the continuous uploading of multiple images, videos, or even real-time feedback, this text-based interactive interface becomes insufficient to meet the user's actual needs. This is because: firstly, the image upload process is cumbersome, affecting interaction efficiency; secondly, it relies primarily on static images actively uploaded by the user (such as wound photos or medical reports), lacking the ability to perceive the user's real-time visual information and failing to achieve dynamic recognition and real-time feedback of the user's status.
[0018] To address the aforementioned issues, the embodiments described in this specification provide a solution that introduces audio and video interaction capabilities into intelligent dialogue assistants.
[0019] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments in this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0020] It should be noted that, for ease of description, the accompanying drawings only show the parts related to the relevant technical solutions. Unless otherwise specified, the embodiments and features described in this specification can be combined with each other. Furthermore, the terms "first," "second," and "third" used in the embodiments of this specification are for informational purposes only and do not constitute any limitation. Moreover, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in the process, method, product, or apparatus that includes the stated elements is not excluded. Furthermore, in this specification, unless explicitly stated otherwise, "receiving and transmitting data" does not necessarily mean direct receiving and transmitting; it can be indirect receiving and transmitting. For example, when A receives data sent by B, it can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, when B sends data to A, it can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.
[0021] Furthermore, it should be noted that specific terms are used to describe embodiments of this specification. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different locations in this specification do not necessarily refer to the same embodiment. Moreover, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples, without contradiction. Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is merely one possible order of execution among many steps, and does not represent the only possible order. Therefore, when the claims involve method steps, adjustments to the order of such steps, or parallel execution between steps, are also within the scope of protection of the claims.
[0022] Furthermore, it should be noted that the user data obtained in this manual (such as user video streams, voice conversation streams, etc.) is authorized by the user and does not involve user privacy.
[0023] The embodiments provided in this specification will be described below with reference to the accompanying drawings.
[0024] First, the terminology used in the embodiments of this specification will be explained. It should be understood that this explanation is for the purpose of providing a clearer understanding of the embodiments described herein and does not necessarily constitute a limitation on the embodiments of this specification.
[0025] Automatic Speech Recognition (ASR) is used to recognize human speech as text, that is, to convert human speech signals into text. It involves multiple processes, including speech signal acquisition, feature extraction, acoustic modeling, language modeling, and decoding.
[0026] Voice Activity Detection (VAD) is a signal processing technique designed to automatically identify the speech components in an audio signal and distinguish them from non-speech components (such as noise or silence). This technique is crucial for improving voice communication quality, reducing storage requirements, and optimizing the performance of speech recognition systems. Its primary focus is on "whether someone is speaking." In other words, VAD can be used to determine whether a user is speaking.
[0027] Text-to-Speech (TTS): Used to convert text information into natural-sounding speech output.
[0028] The preset models are pre-trained artificial intelligence models. In this specification, the preset models include a first preset model, a second preset model, a third preset model, etc. The first preset model is a multimodal model, such as a Vision-Language Model (VLM). A Vision-Language Model (VLM) is a multimodal model that integrates visual information (images or videos) and linguistic information (text). Therefore, a Vision-Language Model can recognize image content. The second preset model is a task model specifically designed for handling specific tasks (such as object detection). The third preset model is a text model, specifically, for example, a Language Model (LLM). A Language Model (LLM) is an artificial intelligence model and a key component of Natural Language Processing (NLP) technology. Based on a Transformer architecture, an LLM can understand and generate high-quality human language text, and is used for dialogue content generation in interactive dialogue systems. It is worth noting that the embodiments in this specification do not limit the number of parameters supported by the preset models, aiming to meet practical application needs.
[0029] The technical solutions provided in the embodiments described below are all based on Figure 1 The technical architecture implementation is shown in the figure. For example... Figure 1As shown, this technical architecture achieves low-latency, high-accuracy online human-computer audio and video interaction through an asynchronous link. The execution process of this asynchronous link includes the following: 1) A video stream (the user's view) is continuously collected from the user side using an image acquisition device (such as a camera). At a set frequency (e.g., once per second), a multimodal model in the image processor is invoked to process multiple consecutive frames of images from the video stream, thereby generating descriptive text for the images and storing it in a memory module. The image acquisition device is, for example, a camera. The memory module is also called RAM. The multimodal model can be, for example, a Visual Language Model (VLM).
[0030] In addition to multimodal models, the aforementioned image processors also include other models, such as task models specifically designed for performing a particular task. The reason for including task models in the image processor is that multimodal models typically analyze the entire image, making it difficult to extract features from local objects for refined analysis. However, if the task is specific to a particular application, such as object detection or security, these tasks often require identification and analysis of specific individuals or objects. In such cases, traditional task models are needed to extract features from corresponding local objects in the image to perform the task. For example, in a skin health scenario, after acquiring a video stream of a user's hand and collecting their voice conversations about their skin condition, a multimodal model can analyze multiple frames of the hand video stream to identify the presence of the arm and potential erythema areas on it. Furthermore, a corresponding task model can be used to perform refined analysis on the identified erythema areas, extracting specific features (such as size and density) of the erythema. Here, the specific features extracted from local objects in the image using the task model will be stored in the memory module in the form of structured parameters.
[0031] Based on the above, in the memory module, the image description text information of multiple consecutive frames in the video stream is stored in the form of text plus structured parameters. The text is the description text of the entire image (i.e., global description text), and the structured parameters are the structured description text of local objects in the image (i.e., local description text). This structured description text includes, but is not limited to, the position, size, category, color, density, and bounding box information of local objects in the image. Local objects can refer to specific regions or elements in the image.
[0032] 2) Continuously collect user voice input through a sound pickup device (such as a microphone) and call the automatic speech recognition (ASR) function to convert the received user voice dialogue stream into text.
[0033] Generally, Automatic Speech Recognition (ASR) uses a fixed silence duration (using a fixed VAD duration) to determine the end of the user's speech input. That is, VAD is a method based on silence detection to determine whether the user's current speech input has ended; for example, if a period of silence is detected, the user's current speech input can be considered to have ended. Of course, other methods can also be used to detect whether the user's current speech input has ended, such as endpoint detection and predefined time threshold detection. Endpoint detection considers not only the presence of silence but also factors such as speech rate and pitch changes to comprehensively determine whether the user's current speech input has ended. Predefined time threshold detection works by setting a fixed time window; if no new audio data is transmitted for more than this time window, the user's current speech input is automatically determined to have ended.
[0034] 3) When the user finishes speaking, such as when a voice VAD is detected, Automatic Speech Recognition (ASR) stops, and the text corresponding to the user's spoken words is concatenated with the relevant image description text information to form a prompt for the response generator. The response generator is built based on a text model. The text model is often called a text dialogue model, which can be, for example, a language model (LLM).
[0035] 4) Input the prompt word into the response generator, which will use its internal text model to generate the response text (also known as the reply text) based on the prompt word.
[0036] 5) Using the TTS function, the conversation text is converted into the corresponding conversation voice and played to the user.
[0037] The technical architecture mentioned above is based on a server-side and client-side implementation. See also... Figure 1As shown, the image acquisition device, audio acquisition device, and voice playback function in the technical architecture are implemented on the client side, while the image processor (including multimodal model and task model), ASR, VAD, stitching operation, response generator (including text model), and TTS function are implemented on the server side. The server side can be a server, server cluster, virtual server, or cloud, etc. The client side can be, but is not limited to, smartphones, smart wearable devices, tablets, laptops, desktop computers, etc. The server side provides corresponding functional services to the client side, such as intelligent dialogue, etc. Users can initiate an intelligent video dialogue through a browser, application (APP), web application H5 (HyperText Markup Language 5, the fifth generation of HTML), lightweight application (also known as a mini-program, a lightweight application), or cloud application on the client side. After the video dialogue is connected to the intelligent dialogue system on the server, an online human-computer audio and video dialogue interaction will be established between the client and the server.
[0038] thus, Figure 2A and Figure 2B An online human-computer interaction system (also known as a service system) according to an embodiment of this specification is also shown. This system includes a client 100 and a server 200. Client 100 is used to display a service interface for human-computer interaction; respond to an audio-visual interaction start operation triggered by the service interface to start online human-computer audio-visual interaction; during the online human-computer audio-visual interaction, it collects video streams and voice dialogue streams from the user side; and sends the video streams and voice dialogue streams to the server; wherein the video stream contains multi-angle images of a target object, and the target object includes parts of the user's body. Server 200 is configured to generate prompt words based on the video stream and the voice dialogue stream; the prompt words are used to guide the generation of detection and analysis text for the target object based on the image description text information of the video stream, the detection and analysis text including the user's body part state analysis results; based on the prompt words, generate response text for responding to the voice dialogue stream; the response text includes the detection and analysis text; convert the response text into response speech, and send the response text and the response speech to the client; The client 100 is also used to play the echo voice and display the echo text.
[0039] The online human-computer video interaction function described above can be integrated into intelligent dialogue assistants in any field. The intelligent dialogue assistant is deployed on the server side, and users can access the service interface provided by the intelligent dialogue assistant for human-computer interaction through a client.
[0040] The specific implementation of the functions of the server 200 and client 100 will be described in detail in the following method embodiments, and will not be elaborated here.
[0041] The technical solutions provided in this specification will be described below by way of method embodiments.
[0042] Figure 3 This diagram illustrates a flowchart of an online human-computer interaction method according to an embodiment of this specification. The execution entity of this method is the server in the aforementioned system. See also... Figure 3 As shown, this online human-computer interaction method includes the following steps: 102. Acquire the video stream and voice dialogue stream on the user side in online human-computer audio and video interaction; wherein, the video stream contains multi-angle images of the target object, and the target object includes the user's body parts; 104. Generate prompt words based on the video stream and the voice dialogue stream; the prompt words are used to guide the generation of detection and analysis text for the target object based on the image description text information of the video stream, and the detection and analysis text includes the user's body part state analysis results; 106. Based on the prompt words, generate a response text for responding to the voice dialogue stream, wherein the response text includes the detection and analysis text; 108. Display the dialogue text and / or play the dialogue voice corresponding to the dialogue text.
[0043] In section 102 above, the online human-computer audio-visual interaction is initiated by the user through a service interface displayed on the client for human-computer interaction. The initiation method can be, but is not limited to, operating corresponding audio-visual interaction controls, voice triggering, or text input triggering. The service interface is provided by a corresponding intelligent dialogue assistant; specifically, this service interface can be a traditional human-computer interaction interface that supports text or image-text dialogue. Therefore, in this specification, the intelligent dialogue assistant supports online human-computer audio-visual interaction and can be applied to intelligent dialogue systems in any field, such as health services (including medical and health services), beauty care, online education, and remote work.
[0044] For example, taking a smart health manager (a smart conversational assistant that provides health services) as an example, users access the human-computer interaction service interface provided by the smart health manager through a client. Generally, this service interface supports text-based or image-based interaction by default. Figure 2A As shown, users can trigger online human-computer audio and video interaction by operating the "Audio and Video Dialogue" control 11 on the service interface. Alternatively, as... Figure 2BAs shown, users can also trigger online human-computer audio-visual interaction by entering an audio-visual dialogue start command (such as "Please start audio-visual dialogue") in the input box 12 provided on the service interface. Of course, in other instances, the audio-visual dialogue start command can also be entered in other ways, such as voice input, such as by operating the "Voice Input" control 13 to input the voice message "Please start audio-visual dialogue".
[0045] This solution integrates online human-computer audio and video interaction functions into the intelligent dialogue assistant, enabling online human-computer audio and video interaction. This allows users to communicate with the intelligent dialogue assistant (i.e., the intelligent dialogue system) more naturally and smoothly. Especially when it is necessary to upload complex image information or perform multi-view detection (such as hair state detection tasks), it can simplify user operations and improve user experience.
[0046] After initiating online human-computer audio and video interaction, the user and the smart health manager will engage in online audio and video interaction via video call. During this interaction, the client will use different acquisition devices to capture the user's video stream and voice dialogue stream, and send them to the server for processing. For example, the user's video input may be captured through a camera to form a video stream, and the user's voice input may be captured through a microphone to form a voice dialogue stream.
[0047] Based on the above, step 102, "obtaining the video stream and voice dialogue stream on the user side in online human-computer audio and video interaction," includes: 1021. Receive the video stream and voice conversation stream sent by the client from the user side; The video stream is video data captured by the client on the user side using an image acquisition device. The image acquisition device can be, but is not limited to, a camera or other image capture device. The voice conversation stream is captured by the client using a sound pickup device, which can be, but is not limited to, a microphone.
[0048] In addition, the video stream contains multi-angle images of the target object. Multi-angle images refer to images of the target object captured from different perspectives (front, side, top, etc.). Changes in perspective can come from changes in the position of the image acquisition device (such as a camera) or from the rotation of the target object itself.
[0049] In different interaction domains, the target objects are different.
[0050] For example, in the field of health services, the target object can be a part of the user's body, such as the face, head, hands, skin, hair, nails, etc.
[0051] For example, in the field of remote work, the target object could be a target device that is malfunctioning.
[0052] For example, in the field of online education, the target audience could be related test questions; or, the target audience could be the user themselves, their body posture, and their movement trajectory (in some cases, it might also include environmental elements or devices involved in the user's interaction). For instance, if the video stream is a dynamic video stream captured of a student conducting a scientific experiment (such as chemistry, biology, physics, or medicine), then the target audience could be the student's body posture, key movement parts, operational behavior trajectory, and the experimental equipment they are interacting with; while if the video stream is a dynamic video stream captured of a user practicing dance, then the target audience could be the user themselves, their body posture, and movement trajectory (especially key body parts and movement characteristics related to dance movements).
[0053] Furthermore, during online human-computer audio and video interaction, in order to collect video data that meets the requirements, the system can guide users to shoot and provide corresponding shooting feedback in a timely manner. Specifically, when guiding users to shoot, appropriate shooting guidance prompts can be provided based on different interaction task types.
[0054] Therefore, the method provided in this embodiment may further include the following steps: S12. Determine the current interaction task type; S14. Determine the appropriate shooting guidance prompts based on the type of interactive task; S16. Output the shooting guidance prompt information to the user to guide the user to adjust the shooting posture.
[0055] The aforementioned video stream includes video clips captured by guiding the user to adjust their shooting posture.
[0056] In S12 above, the interaction task type refers to the type of characteristic interaction task between the intelligent dialogue assistant and the user during human-computer audio and video interaction. The interaction task defines the purpose and content of the interaction. Interaction tasks differ across different fields. For example, in the health service field, the interaction task could be a user health detection task, specifically, a skin condition detection task, a hair condition detection task, an eye health detection task, etc. As another example, in the remote work field, the interaction task can be, but is not limited to, device anomaly analysis tasks (such as hardware configuration, network connection, system crash, security vulnerability analysis, etc.). Furthermore, in the online education field, the interaction task can be, but is not limited to, answering test questions, analyzing experimental operation errors, analyzing student dance movements, etc.
[0057] In this embodiment, the type of interactive task can be determined by the user's own selection or by the intelligent dialogue assistant's active recommendation. This embodiment does not specifically limit the method of determining the type of interactive task.
[0058] For example, after online human-computer audio and video interaction is initiated, the type of interaction task can be determined based on the user's voice input. If the user's voice input includes "I want to check my hair loss status," then the type of interaction task can be determined as "hair status detection task." And / or, the online human-computer audio and video interaction interface can provide a text input entry (specifically, a text chat entry). Users can use this text input entry to input the content they want to interact with, such as "I want to check my hair loss status." In this case, based on this text input, the type of interaction task can be determined as "hair status detection task."
[0059] For example, after the online human-computer audio-visual interaction is initiated, a task window can pop up, displaying multiple interactive task options. The user can select any of these options, and the type of interactive task can be determined based on the selected option. Alternatively, multiple interactive task options can be displayed on one side of the online human-computer audio-visual interaction interface for the user to choose from, thus determining the type of interactive task based on the user's selection. For instance, if the user selects the "Start Hair Detection" option, the type of interactive task can be determined as "Hair Status Detection Task".
[0060] For example, in some scenarios, users may not explicitly specify the task. In such cases, the type of interaction task required can be automatically inferred based on relevant user information. For instance, based on the user's historical online human-computer interaction information (including historical online text and / or image-text dialogue interactions, and historical online audio-visual interactions), if it is determined that the user previously performed a "skin detection" or "facial detection," it can be inferred that the user may need to perform a hair detection next. For this inferred hair detection task, confirmation can be made with the user, for example, by outputting a voice message such as, "Hello, I am **Health Manager. Do you want to consult about hair problems?" Upon receiving confirmation from the user (e.g., "Yes"), the type of interaction task is ultimately determined to be a "hair detection task."
[0061] Based on the example given above, the above S12 "determine the interaction task type" can include any of the following methods: Method 1: Respond to the user-triggered task selection operation and determine the interactive task type based on the selected task item; wherein, the task selection operation may, but is not limited to, being implemented through voice, text input, or point selection. Alternatively, Method 2: Determine the type of interactive task recommended to the user based on the user's relevant information; wherein, the relevant information includes the user's historical online human-computer interaction information.
[0062] In the above S14~S16, in order to collect user-side videos that meet the requirements and provide better interactive services (such as problem consultation services) for users, appropriate shooting guidance prompts can be determined according to the type of interactive task. These shooting guidance prompts can then guide users to shoot and collect suitable image frames.
[0063] The form of the shooting guidance prompts may be, but is not limited to, visual prompts and / or voice prompts. Visual prompts may be, for example, at least one of the following: animated demonstrations, text prompts, and illustrative diagrams.
[0064] For example, taking a hair state detection task as an example, the determined shooting guidance prompts can be a simple tutorial animation. This tutorial animation explains which angles the user needs to take photos of their head (such as the top of the head, forehead, sides, and back of the head), and can be played at a location on the online human-computer audio-visual interaction interface (such as the lower right corner). For example, the tutorial animation might show guidance such as: "Please align the top of your head with the center of the screen," and "Now, please slowly rotate your head to ensure that you can capture the hairline on both sides." Guided by this tutorial animation, the user can adjust the shooting angle by adjusting their head posture or the camera position, thus taking photos of their hair from multiple angles. During the process of the user taking photos of their hair from multiple angles, the camera will collect a continuous image data segment, which becomes the video clip taken under the guidance of the user. In the guided shooting situation, the user will also be provided with timely voice feedback, such as feedback on the image quality (whether it is clear, whether it is obstructed, etc.), whether the current shooting angle meets the requirements, and whether the shooting is complete, etc.
[0065] Therefore, the aforementioned user-side video stream includes video clips captured by guiding the user to adjust their shooting posture.
[0066] Here, in online human-computer audio and video interaction, by guiding users to take pictures from different angles and providing real-time voice feedback, the comprehensiveness and accuracy of image information collection can be ensured, which is beneficial to the accuracy of subsequent response text generation.
[0067] Furthermore, for video streams continuously collected by image acquisition devices, they can be cached first, so that they can be retrieved only after a set time interval has elapsed. Figure 1 The image processor shown in the image processes the image and generates corresponding image description text information.
[0068] Therefore, the method provided in this embodiment may further include the following steps: S22. Perform image recognition analysis on the acquired video stream according to a set time interval to generate image description text information corresponding to the video stream.
[0069] S24. Store the image description text information for later use in generating the prompt words.
[0070] Specifically, it involves calling [the function] according to a set time interval. Figure 1 The image processor shown in the figure performs image recognition and analysis on the video stream continuously collected by the image acquisition device, thereby generating image description text information corresponding to the video stream and storing the image description text information in the corresponding memory.
[0071] The time interval mentioned above can be flexibly set according to the actual situation, and this embodiment does not limit it. For example, the image processor can be called once every 1 second, that is, the image processor can be called once per second. In other words, it can be understood that the image acquisition device will call the image processor once every time a 1-second video stream is collected to perform image recognition and analysis on the 1-second video stream.
[0072] The image processor includes a first preset model and a second preset model. The second preset model is used to perform overall recognition analysis on the image, and the second preset model is used to perform local recognition analysis on the image.
[0073] Based on the above, in a specific feasible technical solution, step S22, "performing image recognition analysis on the acquired video stream and generating descriptive text information corresponding to the video stream," includes: S222: Call the first preset model to perform global recognition on multiple consecutive frames of images in the video stream and generate global description text; S224. Call the second preset model to perform local object recognition in the continuous multi-frame images and generate local description text.
[0074] Furthermore, the "storing the image description text information" in step S24 above includes: S242. The global description text and the local description text corresponding to the consecutive multi-frame images are associated and stored for later use in the generation of the prompt words.
[0075] In S222 and S224 above, the first preset model and the second preset model can be respectively Figure 1The multimodal model and task model shown herein, where a task model is used to perform a specific task. The generated global descriptive text refers to the textual information content that provides a general description of the image as a whole across multiple consecutive frames. The generated local descriptive text refers to the detailed local object description information generated by extracting the object features of local objects through local object recognition in multiple consecutive frames. The local object can be, but is not limited to, a specific local region or element (such as hair, erythema, etc.) in the image, and the corresponding local descriptive text includes at least one of the following: location, size, density, color, category, and morphology of the local object.
[0076] For example, the first preset model can be called once every second (i.e., the first preset model is called once per second) to perform global recognition on multiple consecutive frames of images in the video stream collected by the image acquisition device within one second, thereby generating global descriptive text for the multiple consecutive frames of images; furthermore, the second preset model can be called to extract features from local objects identified from the multiple consecutive frames of images, thereby generating local descriptive text for the local objects, etc.
[0077] It should be noted that step S224 is not a mandatory step, but rather an optional one. Depending on the type of interaction task, step S224 can be triggered only when it is determined that local object recognition is required. That is, the method provided in this embodiment may also include the following steps: S223. When it is determined that local object recognition is required based on the type of interactive task, the above step S224 is triggered.
[0078] For details on determining the type of interactive task here, please refer to the relevant content described in the other embodiments above, which will not be repeated here.
[0079] For example, if the interaction task type is "hair state detection task", then it is determined that local object recognition is required, and the local object to be recognized is the hair region in the image.
[0080] For example, if the interaction task type is "erythema detection task on skin", then it is determined that local object recognition is required, and the local object to be recognized is the erythema element in the image.
[0081] In step S242 above, the global and local description texts of multiple consecutive frames of images can be associated and stored in corresponding memory locations for subsequent support of prompt word generation. If memory is available... Figure 1 The memory module is shown in the diagram. During associative storage, the local description text is saved in the form of structured parameters. For example, the local description text is saved as: identifier of the local object (e.g., name): **; location: **; size: **; density: **.
[0082] For user-side voice dialogue streams collected via sound pickup devices (such as microphones), Automatic Speech Recognition (ASR) is used to convert the voice signals in the dialogue stream into corresponding dialogue text. Upon detecting the end of the voice dialogue stream (e.g., detecting a period of silence indicates the current segment of the voice dialogue stream has ended, meaning the user's current voice input has ceased), ASR stops, and the system enters the prompt word generation stage. In the prompt word generation stage, image description text information semantically relevant to the dialogue text is searched from stored information. This image description text is then appended to the dialogue text to obtain the corresponding prompt word. This prompt word serves as... Figure 1 The input parameters of the response generator (specifically, the third preset model within it) are shown in the diagram so that the response generator can generate echo text based on the input prompt words to respond to the user's voice stream.
[0083] Based on the above, in a specific implementable technical solution, the above-mentioned 104 "generating prompt words based on the video stream and the voice dialogue stream" includes: 1042. Convert the voice dialogue stream into text to generate the corresponding dialogue text; 1044. Retrieve the image description text information related to the dialogue text from the stored information; 1046. Based on the dialogue text and the retrieved image description text information, generate the prompt word.
[0084] The generated prompts here may include role settings, detection and analysis task instructions, output constraints, etc., for the third preset model. These prompts guide the third preset model to use the image description text information of the video stream as context to perform corresponding detection and analysis tasks on the target object and generate matching detection and analysis text. The model then responds to the user's inquiry in the dialogue text based on the obtained detection and analysis text.
[0085] In other words, the prompt words are used to guide the generation of detection and analysis text for the target object based on the image description text information of the video stream. The detection and analysis text includes the user's body part state analysis results, such as the detection and analysis results of hair / skin conditions.
[0086] And the aforementioned 106 "Generating response text for responding to the voice dialogue stream based on the prompt words" may include: 1062. Input the prompt word into the third preset model, and execute the third preset model to output the response text.
[0087] In the above process, after converting the voice dialogue stream into corresponding dialogue text, the Retrieval Enhancement (RAG) module can be used to retrieve image description text information that is semantically related to the dialogue text from memory. Then, the dialogue text and the retrieved image description text information are concatenated to generate prompt words for inputting the third preset model.
[0088] The third preset model is a text model, specifically, for example, a language model (LLM). The response text generated based on the prompt words using this third preset model can be used to respond to user inquiries (such as hair loss) in a voice dialogue stream.
[0089] For example, in the field of healthcare services, suppose the generated prompt is: "You are an intelligent health assistant. Please generate an objective initial hair health detection and analysis result based on the image description text information of the video stream and the user's dialogue text. The output should be in Chinese and divided into three parts: summary, observation, and suggestions, with a professional and gentle tone." The image description text information of the video stream is "***", and the user's dialogue text is "Please check my hair loss condition." After inputting this prompt into a third preset model, the output response text could be: "From the image, your hairline shows a slight receding trend; multiple angle images show sparse hair on the top of the head, and some areas of scalp are visible, which may be related to genetics or stress; it is recommended to maintain a good lifestyle and consult a dermatologist for professional testing if necessary." The content in the response text, such as "slight receding hairline," "sparse hair on the top of the head," and "some areas of scalp are visible," constitutes the user's hair condition detection and analysis result.
[0090] The above description summarizes the relevant processing performed on video streams and audio dialogue streams, combined with... Figure 1 As can be seen, this embodiment actually processes video streams captured by means such as a camera and user voice input collected by means such as a microphone asynchronously, which can improve the response speed and user experience. Correspondingly, in this asynchronous method, Figure 1The image processor (specifically, the multimodal model) and response generator (specifically, the text model) shown in the diagram run asynchronously. The reason for not using a synchronous operation scheme, such as that of the multimodal model and the text model, is that synchronous operation has a longer latency, resulting in longer session response times and consequently longer user wait times. Furthermore, synchronous operation, which only combines the latest video feed, requires high synchronization between image, audio, and video, making it difficult to guarantee quality. Alternatively, the multimodal model can be used to process the video stream while simultaneously generating session text; that is, only the multimodal model is called to handle both video stream processing and session text generation. However, this approach has a slightly longer latency than the asynchronous approach, and using the multimodal model for session generation requires more alignment training data, leading to higher costs.
[0091] Furthermore, the TTS (Text-to-Speech) function can be used to convert the dialogue text into corresponding audio, which can then be sent to the user's client for playback. Alternatively, the dialogue text can also be sent to the client simultaneously, allowing the client to display the dialogue text while playing the audio, facilitating user understanding.
[0092] Therefore, a specific implementable technical solution for "displaying the echo text and / or playing the echo voice corresponding to the echo text" in the above 108 may include: 1082. Convert the echo text into echo speech; 1084. Send the echo voice and the echo text to the user-side client so that the client can play the echo voice and / or display the echo text.
[0093] As mentioned above, considering that some fields, such as healthcare services, often have specialized terminology, simply playing the audio containing these terms might lead to user misunderstanding. Therefore, this system displays the corresponding text alongside the audio, allowing users to visually understand the audio through the text.
[0094] In summary, this embodiment upgrades the traditional multimodal interaction of intelligent dialogue assistants to an audio-visual interaction mode, which can support users to conduct natural and smooth multimodal interactions, optimize the user interaction experience, and has strong practicality and interaction flexibility in practical applications. It can promote the leapfrog development of intelligent dialogue assistants and enhance their commercial value and market competitive advantage.
[0095] This manual also provides another online human-computer interaction method, in which the client in the aforementioned system is the executor. For details, please refer to... Figure 4 As shown, this online human-computer interaction method includes the following steps: 202. Display the service interface used for human-computer interaction; 204. Respond to the audio-visual interaction start operation triggered through the service interface and start online human-computer audio-visual interaction; 206. During the online human-computer audio and video interaction process, video streams and voice dialogue streams are collected from the user side; wherein, the video stream contains multi-angle images of the target object, and the target object includes parts of the user's body; 208. Send the video stream and voice dialogue stream to the server; 210. Receive the response text and corresponding voice message returned by the server; wherein the response text is generated based on prompt words to respond to the voice dialogue stream; the prompt words are generated based on the video stream and the voice dialogue stream to guide the generation of detection and analysis text of the target object based on the image description text information of the video stream, and the detection and analysis text includes the user's body part state analysis results; the response text includes the detection and analysis text. 212. While playing the echo voice, display the echo text.
[0096] For specific implementation details of the steps described above in this embodiment, please refer to the relevant content in other embodiments, which will not be repeated here. Furthermore, the method provided in this embodiment may also include some steps disclosed in other embodiments, which can also be referred to the relevant content in other embodiments, and will not be repeated here.
[0097] This specification also provides another online human-computer interaction method, in which some steps (such as 302, 304, and 306) are executed by the client, while the remaining steps (such as 308, 3010, and 3012) are executed by the server. Furthermore, the application scenario for this method is an intelligent dialogue assistant (such as a health manager) used for health services. For details, see [link to documentation]. Figure 5 As shown, this online human-computer interaction method includes the following steps: 302. Display a health service interface, wherein the health service interface has audio and video interactive controls; 304. Responding to the operation of the audio-visual interaction control, initiate online human-computer audio-visual interaction; 306. During the online human-computer audio and video interaction process, the user's biometric video stream and voice dialogue stream are collected; 308. Generate prompt words based on the biological video stream and the voice dialogue stream; 310. Input the prompt words into the third preset model, and execute the third preset model to output the response text; the response text is used to respond to the user's inquiry about health issues, and the health issues are extracted from the voice dialogue stream; 312. Play the conversation text to the user.
[0098] The health service interface mentioned above is as follows: Figure 2A and Figure 2B The service interface for human-computer interaction shown in the figure, as well as audio and video interaction controls such as Figure 2A The “Audio / Video Dialogue” control 11 is shown in the image.
[0099] Furthermore, the aforementioned bio-video stream refers to a continuous sequence of images captured by an image acquisition device that includes parts of the user's body (such as face, head, hands, neck, etc.). From this bio-video stream, image information related to the user's biometrics, such as hair condition and skin condition, can be extracted to support health check analysis.
[0100] For specific implementation details of the steps described above in this embodiment, please refer to the relevant content in other embodiments, which will not be repeated here. Furthermore, the method provided in this embodiment may also include some steps disclosed in other embodiments, which can also be referred to the relevant content in other embodiments, and will not be repeated here.
[0101] The above text combined Figures 3-5 Specific embodiments of the embodiments described herein have been described. It should be noted that other embodiments fall within the scope of the appended claims; and, in some cases, the actions or steps recited in the claims may be performed in a different order than those shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0102] The apparatus embodiments corresponding to the various method embodiments provided in this specification are described below.
[0103] Figure 6 A schematic diagram of the structure of an online human-computer interaction device provided in an exemplary embodiment of this specification is shown. Figure 6 As shown, the device includes: an acquisition module 42, a generation module 44, and a playback module 46. Among them, The acquisition module 42 is used to acquire the video stream and voice dialogue stream on the user side in online human-computer audio and video interaction; wherein the video stream contains multi-angle images of the target object, and the target object includes the user's body parts; The generation module 44 is used to generate prompt words based on the video stream and the voice dialogue stream; the prompt words are used to guide the generation of detection and analysis text of the target object based on the image description text information of the video stream, and the detection and analysis text includes the user's body part state analysis results; The generation module 44 is further configured to generate a response text for responding to the voice dialogue stream based on the prompt words, wherein the response text includes the detection and analysis text; The playback display module 44 is used to display the echo text and / or play the echo voice corresponding to the echo text.
[0104] In one embodiment, the device further includes: a recognition and analysis module and a storage module. The recognition and analysis module is used to perform image recognition and analysis on the acquired video stream at set time intervals to generate image description text information corresponding to the video stream. The storage module is used to store the image description text information for subsequent use in generating the prompt words.
[0105] In one implementation, the aforementioned recognition and analysis module, when performing image recognition and analysis on the acquired video stream to generate image description text information corresponding to the video stream, is specifically configured to: invoke a first preset model to perform global recognition on multiple consecutive frames of images in the video stream to generate global description text; and invoke a second preset model to perform local recognition on local objects in the multiple consecutive frames of images to generate local description text. Furthermore, the aforementioned storage module, when storing the image description text information, may specifically be configured to: associate and store the global description text and the local description text corresponding to the multiple frames of images.
[0106] In one embodiment, the device further includes: a determining module and a triggering module. The determining module is used to determine the current interaction task type. The triggering module is used to, when the interaction task type determines that local object recognition is required, trigger the invocation of the second preset model to perform local object recognition in the multi-frame images and generate local description text. The local description text includes at least one of the following: the position, size, density, color, and category of the local object.
[0107] In one implementation, the determining module, when used to determine the type of interactive task, is specifically used to: respond to a task selection operation triggered by a user and determine the type of interactive task based on the selected task item; or, determine the type of interactive task recommended by the user based on relevant user information; wherein, the relevant information includes the user's historical online human-computer interaction information.
[0108] In one implementation, the determining module is further configured to determine appropriate shooting guidance prompts based on the current interaction task type. The output module is further configured to output the shooting guidance prompts to the user to guide the user in adjusting their shooting posture. The video stream includes video segments captured by guiding the user to adjust their shooting posture, wherein the shooting guidance prompts include visual prompts and / or voice prompts.
[0109] In one embodiment, the generation module, when used to generate prompt words based on the video stream and the voice dialogue stream, is specifically configured to: convert the voice dialogue stream into text to generate corresponding dialogue text; retrieve image description text information related to the dialogue text from stored information; and generate the prompt words based on the retrieved image description text information and the dialogue text. Furthermore, the generation module, when used to generate response text for responding to the voice dialogue stream based on the prompt words, is specifically configured to: input the prompt words into a third preset model, and execute the third preset model to output the response text.
[0110] In one implementation, the playback module described above, when used to display the echo text and / or play the echo text to the user, is specifically configured to: convert the echo text into echo speech; and send the echo speech and the echo text to the user-side client so that the client can play the echo speech and / or display the echo text.
[0111] Figure 7 A schematic diagram of the structure of an online human-computer interaction device provided in another exemplary embodiment of this specification is shown. For example... Figure 7As shown, the device includes: a display module 52, a startup module 54, a data acquisition and transmission module 56, a receiving module 58, and a broadcast display module 510. The display module 52 displays a service interface for human-computer interaction. The startup module 54 responds to an audio-visual interaction startup operation triggered by the service interface, initiating online human-computer audio-visual interaction. The data acquisition and transmission module 56 acquires video streams and voice dialogue streams from the user side during the online human-computer audio-visual interaction; wherein the video stream contains multi-angle images of a target object, including the user's body parts; and also transmits the video streams and voice dialogue streams to the server. The receiving module 58 receives the response text and corresponding voice dialogue returned by the server; wherein the response text is generated based on prompt words to respond to the voice dialogue stream; the prompt words are generated based on the video stream and the voice dialogue stream to guide the generation of detection and analysis text for the target object based on image description text information from the video stream, the detection and analysis text including the user's body part status analysis results; the response text includes the detection and analysis text. The playback display module 510 is used to display the playback text while playing the voice message.
[0112] Figure 8 A schematic diagram of the structure of an online human-computer interaction device provided in yet another exemplary embodiment of this specification is shown. For example... Figure 8 As shown, the device includes: a display module 62, a startup module 64, a data acquisition module 66, a generation module 68, an execution module 610, and a playback module 612. The display module 62 displays a health service interface, which has audio-visual interaction controls. The startup module 64 responds to operations on the audio-visual interaction controls and initiates online human-computer audio-visual interaction. The data acquisition module 66 acquires the user's bio-video stream and voice dialogue stream during the online human-computer audio-visual interaction; wherein the bio-video stream is a continuous image sequence containing the user's body parts. The generation module 68 generates prompts based on the bio-video stream and the voice dialogue stream. The execution module 610 inputs the prompts into a third preset model, executes the third preset model, and outputs response text; the response text responds to the user's health questions, which are extracted from the voice dialogue stream. The playback module 612 plays the response text corresponding to the response voice to the user.
[0113] It should be noted that the above-mentioned devices can implement the technical solutions described in the corresponding method embodiments. The specific implementation principles of each module or unit can be found in the relevant content of the corresponding method embodiments, and will not be elaborated further here. Furthermore, for ease of description, the above devices are described by function as various modules or units. Of course, when implementing one or more of this specification, the functions of each module or unit can be implemented in one or more software and / or hardware, or a module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0114] Furthermore, embodiments of this specification also provide an electronic device. For example... Figure 9 As shown, the electronic device 700 includes a memory 71 and a processor 72.
[0115] The aforementioned memory 71 can be implemented by at least one volatile or non-volatile storage device of any type, or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Furthermore, the memory, wholly or partially, can be integrated with the processor. The memory can contain both removable and non-removable components.
[0116] The processor 72 described above may include one or more general-purpose processors and / or special-purpose processors.
[0117] Furthermore, memory 71 may contain a non-transitory computer-readable medium storing executable program instructions 712 (e.g., compiled or uncompiled program logic and / or machine code). Processor 72 is capable of executing the program instructions 712 stored in memory to implement any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Additionally, execution of program instructions 712 by processor 72 may result in processor using corresponding data 711.
[0118] For example, the program instructions 712 described above may include an operating system 7122 (e.g., an operating system kernel, device drivers, and / or other modules) installed on the electronic device 700, and one or more application programs 7121 (e.g., a browser, social media application, or game application). Similarly, the data 711 described above may include operating system data 7112 and application data 7111. The operating system data 7112 is primarily accessible to the operating system 7122, while the application data 7111 is primarily accessible to one or more application programs 7121. The application data 7111 may reside in a file system that is visible or hidden from the user of the electronic device 700.
[0119] Application 7121 can communicate with operating system 7122 through one or more application programming interfaces (APIs). These APIs facilitate application 7122 in reading and / or writing application data, transmitting or receiving information via communication components, and receiving or displaying information on the user interface. In some terms, application 7121 may be simply referred to as "app". Furthermore, application 7121 can be downloaded to the electronic device through one or more online application stores or app markets. However, application 7121 can also be installed on electronic device 700 in other ways, such as through a web browser or a physical interface on electronic device 700 (e.g., a USB port). Furthermore, such as Figure 9 As shown, the electronic device also includes: a communication component 73, a display 74, a power supply component 75, an audio component 76, a user interface 77, and other components. Figure 9 The diagram only shows some components and does not imply that the electronic device 700 includes only these components. Figure 9 The components shown. Additionally... Figure 9 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the electronic device 700. The electronic device 700 in this embodiment can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT device; it can also be a server-side device such as a conventional server, cloud server, or server array; or it can be an integrated device combining terminal and server-side devices. If the electronic device 700 in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 9 The components within the dashed box; if the electronic device 700 in this embodiment is implemented as a conventional server, cloud server, or server array, etc., then it may not include... Figure 9 The component within the dashed box.
[0120] The aforementioned communication component 73 is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component 73 can access wireless networks based on communication standards, such as 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component 73 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. Specifically, the communication component 73 includes a communication interface that enables the electronic device 700 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface may include a chipset and an antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces.
[0121] The aforementioned display 74 includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from a user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0122] The power supply component 75 provides power to various components of the device in which it resides. The power supply component 75 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component resides.
[0123] The aforementioned audio component 76 can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0124] The user interface 77 described above includes receiving user input and providing output to the user. Therefore, the user interface 77 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. The user interface 77 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, the user interface 77 may include software, circuitry, or other forms of logic capable of transmitting and receiving data from external user input / output devices. Additionally or alternatively, the electronic device 700 may support remote access from other devices via a communication interface or another physical interface (not shown). The user interface 77 may be configured to receive user input, the position and movement of which may be indicated by an indicator or cursor described herein. The user interface 77 may also be configured as a display device for rendering or displaying text fragments.
[0125] Accordingly, embodiments of this specification also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium. Furthermore, embodiments of this specification also provide a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, it causes the computer to perform actions such as... Figures 3 to 5 The method described.
[0126] This specification also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implements... Figures 3 to 5 The method described.
[0127] Those skilled in the art will recognize that the functions described in the various embodiments disclosed in this specification in one or more of the examples above can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium.
[0128] The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above descriptions are merely specific implementations of the multiple embodiments disclosed in this specification and are not intended to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent substitutions, improvements, etc., made based on the technical solutions of the multiple embodiments disclosed in this specification should be included within the protection scope of the multiple embodiments disclosed in this specification.
Claims
1. An online human-computer interaction method, characterized in that, Suitable for the client, the method includes: Display the service interface; In response to an audio / video interaction initiation operation triggered through the service interface, a video call connection is established with the server to conduct online human-computer audio / video interaction; wherein, the server deploys a preset model and uses the preset model to interact with the user; In online human-computer audio and video interaction, a video stream is captured by an image acquisition device on the client and sent to the server; During the acquisition of video streams, the image acquisition device plays and / or displays shooting guidance prompts from the server to guide the user to adjust the angle, so that the image acquisition device can acquire a video stream that meets the requirements after adjustment and send it to the server.
2. The method according to claim 1, characterized in that, The service interface displays audio and video interaction controls; as well as Responding to the audio / video interaction initiation operation triggered through the service interface, and establishing a video call connection with the server, includes: In response to the startup operation initiated through the audio and video interaction control, a video call connection is established with the server.
3. The method according to claim 1, characterized in that, The shooting guidance information includes visual prompts and / or voice prompts; The visual prompts include at least one of the following: animated demonstrations, text prompts, and illustrative prompts.
4. The method according to claim 3, characterized in that, The server displays the shooting guidance information, including: Displays an online human-computer audio-visual interaction interface; The visual prompt is displayed at a location on the online human-computer audio-visual interaction interface.
5. The method according to claim 1, characterized in that, Also includes: Display multiple interactive task options, respond to the task selection operation triggered by the user, trigger the server to determine the interactive task type based on the interactive task option selected by the user, and determine the appropriate shooting guidance prompt information according to the interactive task type; and / or Obtain the user's voice or text input; send the voice or text input to the server, whereby the server determines the interaction task type and determines the appropriate shooting guidance prompt information based on the interaction task type.
6. The method according to claim 1, characterized in that, The shooting guidance prompt information is at least one of the following prompts fed back by the server in response to the received video stream: Image quality prompts, prompts indicating whether the shooting angle meets the requirements, and prompts indicating whether the shooting is complete.
7. The method according to any one of claims 1 to 5, characterized in that, Also includes: In online human-computer audio and video interaction, the voice dialogue stream is captured by the audio pickup device on the client. The voice dialogue stream is sent to the server, whereby the server uses the preset model to process the received video stream and voice dialogue stream to generate response text. The server displays or plays the response text via voice.
8. An online human-computer interaction method, characterized in that, Suitable for the server side, the method includes: The system connects to a video call with the client to conduct online human-computer audio and video interaction with the user on the client side; the server deploys a preset model and uses the preset model to interact with the user. Receive a video stream sent by the client, wherein the video stream is captured by an image acquisition device on the client; Confirm the shooting guidance prompts; The shooting guidance information is sent to the client, so that the client plays and / or displays the shooting guidance information during the video stream acquisition process of the image acquisition device, in order to guide the user to adjust the angle so that the video stream subsequently sent to the server meets the requirements.
9. The method according to claim 8, characterized in that, Also includes: Receive the voice conversation stream sent by the client; Based on the received video stream and the received voice dialogue stream, generate the response text; The response text is sent to the client.
10. The method according to claim 9, characterized in that, Also includes: When the current voice input ends based on the received voice dialogue stream, a response text is generated.
11. An online human-computer interaction method, characterized in that, Suitable for the client, the method includes: Displays the service interface corresponding to the smart health manager; In response to the operation of video calling with the smart health manager triggered through the service interface, online human-computer audio and video interaction is initiated; During the online human-computer audio and video interaction process, a video stream is captured by the image acquisition device on the client and sent to the server corresponding to the smart health manager; the server is equipped with a preset model corresponding to the smart health manager so that the smart health manager can interact with the user using the preset model; During the video stream acquisition process, the image acquisition device plays and / or displays shooting guidance prompts from the smart health manager to guide the user to adjust the angle, so that the image acquisition device can acquire a video stream that meets the requirements after adjustment and send it to the server.
12. The method according to claim 11, characterized in that, The shooting guidance prompts include at least one of the following: The system provides guidance and prompts regarding shooting multi-angle images of body parts, image quality, whether the current shooting angle meets requirements, and whether the shooting is complete.
13. A service system, characterized in that, include: A client is used to implement the steps in the online human-computer interaction method according to any one of claims 1 to 7, or to implement the steps in the online human-computer interaction method according to claim 11 or 12. The server is used to implement the steps in the online human-computer interaction method according to any one of claims 8 to 10.
14. An electronic device, characterized in that, The method includes a memory and a processor, wherein the memory stores executable program instructions, and the processor executes the program instructions to implement the method according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed in a computer, causes the computer to perform the method described in any one of claims 1 to 12.
16. A computer program product, characterized in that, The computer program product includes a computer program or instructions that, when executed by a processor, implement the method described in any one of claims 1 to 12.