Online man-machine interaction method, system and device, electronic equipment, storage medium and program product
By introducing audio and video interaction capabilities into the intelligent dialogue assistant, collecting the user's video stream and voice stream, generating prompt words to guide image description text and playing reply voice, the problem of low efficiency of image-text dialogue interaction in the existing technology is solved, and a more natural multimodal interaction experience is achieved.
Patent Information
- Application Number
- CN202511221891.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-28
AI Technical Summary
When existing intelligent dialogue assistants process complex image information and multi-view detection, the text-and-image dialogue interaction method is difficult to meet user needs, resulting in low interaction efficiency and lack of instant visual information perception capabilities.
The system introduces audio and video interaction capabilities. By collecting the video stream and voice dialogue stream on the user side, it generates prompt words to guide image description text information, generates reply text and plays reply voice, and supports detection and analysis of multi-angle images.
It achieves more natural and smooth multimodal interaction, simplifies user operations, and improves user experience, especially when uploading complex image information or multi-view detection is required.
Smart Images

Figure CN120723950A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and in particular to an online human-computer interaction method, system, device, electronic device, storage medium, and program product. Background Art
[0002] Intelligent conversational assistants, which can interact with users and help them solve problems or chat, have been widely used in various fields such as healthcare. Currently, intelligent conversational assistants mainly use text conversations as their primary interaction method. With the development of multimodal technology, some intelligent conversational assistants have gradually introduced image-text conversation capabilities, allowing users to upload images when asking questions. This allows them to more accurately understand and respond to user questions based on the image content. However, when the image information that users need to upload is more complex, such as involving multiple images, videos, or even requiring instant feedback and interaction, this image-text conversation interaction method is difficult to meet user needs.
[0003] Therefore, there is an urgent need to introduce a new interactive dialogue capability for intelligent dialogue assistants. Summary of the Invention
[0004] The various embodiments of this specification provide an online human-computer interaction method, system, device, electronic device, storage medium, and program product, which implement the introduction of audio and video interaction capabilities into intelligent conversation assistants, helping to meet users' more complex problem needs. In the first embodiment, this specification provides an online human-computer interaction method. The method includes: Collecting video streams and voice dialogue streams from the user side during online human-computer audio and video interaction; wherein the video streams contain multi-angle images of a target object, and the target object includes a body part of the user; Generate prompt words based on the video stream and the voice dialogue stream; the prompt words are used to guide the generation of a detection analysis text of the target object based on the image description text information of the video stream, the detection analysis text including the analysis results of the user's body part status; Based on the prompt word, generating a reply text for responding to the voice dialogue flow, wherein the reply text includes the detection and analysis text; Display the reply text, and / or play the reply voice corresponding to the reply text.
[0005] In a second embodiment, this specification also provides an online human-computer interaction method. The method includes: Displaying a service interface for human-computer interaction; In response to the audio and video interaction start operation triggered by the service interface, start online human-computer audio and video interaction; During the online human-computer audio and video interaction process, a video stream and a voice dialogue stream on the user side are collected; wherein the video stream includes multi-angle images of a target object, and the target object includes a body part of the user; Sending the video stream and voice dialogue stream to the server; Receive the reply text and reply voice corresponding to the reply text returned by the server; wherein the reply text is generated based on a prompt word to respond to the voice dialogue stream; the prompt word is generated based on the video stream and the voice dialogue stream to guide the generation of a detection and analysis text of the target object based on the image description text information of the video stream, the detection and analysis text including the user's body part status analysis result; the reply text includes the detection and analysis text; While playing the reply voice, the reply text is displayed.
[0006] In a third embodiment, this specification also provides an online human-computer interaction method. The method includes: Displaying a health service interface, wherein the health service interface has audio and video interactive controls; In response to the operation of the audio and video interaction control, starting online human-computer audio and video interaction; During the online human-computer audio and video interaction process, a biological video stream and a voice dialogue stream of the user are collected; wherein the biological video stream is a continuous image sequence containing body parts of the user; generating prompt words based on the biological video stream and the voice dialogue stream; Inputting the prompt word into a third preset model, executing the third preset model to output a reply text; the reply text is used to respond to the health question consulted by the user, and the health question is extracted from the voice dialogue flow; Play the reply voice corresponding to the reply text to the user.
[0007] In a fourth embodiment, this specification provides a service system. The system includes: The client is configured to display a service interface for human-computer interaction; initiate online human-computer audio and video interaction in response to an audio and video interaction initiation operation triggered by the service interface; collect a video stream and a voice conversation stream from the user side during the online human-computer audio and video interaction; and send the video stream and the voice conversation stream to the server; wherein the video stream includes multi-angle images of a target object, wherein the target object includes a body part of the user; The server is configured to generate prompt words based on the video stream and the voice dialogue stream; the prompt words are used to guide the generation of a detection and analysis text of the target object based on the image description text information of the video stream, the detection and analysis text including the analysis results of the user's body part status; based on the prompt words, generate a reply text for responding to the voice dialogue stream, the reply text including the detection and analysis text; convert the reply text into a reply voice, and send the reply text and the reply voice to the client; The client is also used to play the reply voice and display the reply text at the same time.
[0008] In a fifth embodiment, this specification provides an online human-computer interaction device. The device includes: An acquisition module is used to acquire the video stream and voice dialogue stream of the user side in the online human-computer audio and video interaction; wherein the video stream includes multi-angle images of the target object, and the target object includes the user's body parts; a generation module, configured to generate prompt words based on the video stream and the voice dialogue stream; the prompt words are used to guide the generation of a detection analysis text of the target object based on the image description text information of the video stream, the detection analysis text including the analysis results of the user's body part status; The generating module is further configured to generate a reply text for responding to the voice dialogue flow based on the prompt word, wherein the reply text includes the detection and analysis text; The playing and displaying module is used to display the reply text and / or play the reply voice corresponding to the reply text.
[0009] In a sixth embodiment, this specification also provides an online human-computer interaction device. The device includes: A display module, used to display a service interface for human-computer interaction; A starting module, configured to respond to an audio and video interaction starting operation triggered through the service interface and start online human-computer audio and video interaction; A collection module, configured to collect a video stream and a voice conversation stream from a user during the online human-computer audio and video interaction process; wherein the video stream includes multi-angle images of a target object, and the target object includes a body part of the user; A sending module, used for sending the video stream and voice dialogue stream to the server; a receiving module, configured to receive the reply text and reply voice corresponding to the reply text returned by the server; wherein the reply text is generated based on a prompt word to respond to the voice dialogue stream; the prompt word is generated based on the video stream and the voice dialogue stream to guide the generation of a detection and analysis text of the target object based on the image description text information of the video stream, the detection and analysis text including the user's body part status analysis result; and the reply text includes the detection and analysis text; The playing and displaying module is used to play the reply voice and display the reply text at the same time.
[0010] In a seventh embodiment, this specification also provides an online human-computer interaction device. The device includes: A display module, configured to display a health service interface having audio and video interaction controls; A starting module, configured to respond to an operation on the audio and video interaction control and start online human-computer audio and video interaction; An acquisition module, configured to acquire a user's biological video stream and voice dialogue stream during the online human-computer audio and video interaction process; wherein the biological video stream is a continuous image sequence containing body parts of the user; A generation module, configured to generate prompt words based on the biological video stream and the voice dialogue stream; an execution module, configured to input the prompt word into a third preset model, execute the third preset model and output a reply text; the reply text is used to respond to the health question consulted by the user, the health question being extracted from the voice dialogue stream; The playing module is used to play the reply voice corresponding to the reply text to the user.
[0011] In an eighth embodiment, this specification provides an electronic device, comprising a memory and a processor, wherein the memory stores executable program instructions, and when the processor executes the program instructions, the methods provided in the first to third embodiments are implemented.
[0012] In a ninth embodiment, this specification provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, the computer is caused to execute the methods provided in the first to third embodiments above.
[0013] In a tenth embodiment, this specification also provides a computer program product, including a computer program / instruction, which implements the methods provided in the first to third embodiments when executed by a processor.
[0014] The solutions provided in the aforementioned embodiments of this specification enable online human-computer audio and video interaction. Prompt words are generated based on the collected user-side video and voice conversation streams. A response text is then generated based on the prompt words to respond to the voice conversation stream and output to the user. This solution implements online human-computer audio and video interaction, allowing users to engage in more natural and smooth multimodal communication with an intelligent conversation system (also known as an intelligent conversation assistant). This simplifies user operations and enhances the user experience, particularly when uploading complex image information or performing multi-view detection tasks (such as hair condition detection). This online human-computer audio and video interaction is triggered via a service interface for human-computer interaction. Specifically, this service interface includes audio and video interaction controls, which users can trigger by operating. For example, this service interface can be a health service interface provided by an intelligent conversation assistant for health services. Furthermore, when outputting the response text to the user, the response text can be displayed while the corresponding voice message is played, making it easier for the user to understand the content of the voice message. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. In the drawings: Figure 1 A schematic diagram of the technical architecture based on which each method in this specification is implemented is provided as an exemplary embodiment; Figure 2A and Figure 2B A schematic diagram of the structure of a service system (specifically, an online human-computer interaction system, i.e., an intelligent dialogue system) provided by an exemplary embodiment of this specification; Figure 3 、 Figure 4 and Figure 5 A flowchart of an online human-computer interaction method provided by an exemplary embodiment of this specification; Figure 6 、 Figure 7 and Figure 8 A schematic diagram of the structure of an online human-computer interaction device provided by an exemplary embodiment of this specification; Figure 9 This is a schematic structural diagram of an electronic device provided as an exemplary embodiment of this specification. DETAILED DESCRIPTION
[0016] With the rapid development of artificial intelligence and mobile internet technologies, intelligent conversational assistants have become essential tools for human-computer interaction in various applications. These interactive assistants can engage in conversations with users, helping them solve problems or simply chatting. They have been widely used in fields such as healthcare. For example, health-related intelligent conversational assistants (such as the health managers offered in some apps) are dedicated to helping users resolve various issues before, during, and after medical treatment. They cover a wide range of scenarios, including health consultations, initial disease screening, medical advice, medication reminders, and rehabilitation guidance, significantly improving users' health management efficiency and overall well-being. Currently, health-related intelligent conversational assistants, as well as other types of intelligent conversational assistants, primarily support text-based interaction. For example, users can enter a text description of their health status, and these intelligent conversational assistants can leverage their language models to generate relevant health recommendations or medical guidance. With the advancement of multimodal technologies, an increasing number of intelligent conversational assistants have added the ability to interact with images and text. These assistants allow users to upload images when asking questions, then interpret and generate responses based on the image content, thereby improving intuitive interaction and service accuracy. This interactive image-to-text method allows users to upload images by clicking the "Upload Image" control on the interface. This is effective and convenient when users need to upload a small number of images. However, when the image information required is complex, such as when multiple images or videos need to be uploaded continuously, or when instant feedback is required, this interactive image-to-text method fails to meet users' actual needs. This is because: firstly, the image upload process is cumbersome, affecting interaction efficiency; secondly, it relies primarily on static images (such as wound photos and test reports) uploaded by the user, lacking the ability to perceive the user's immediate visual information, and cannot provide dynamic recognition and instant feedback on the user's status.
[0017] To address the above issues, the embodiments in this specification provide a solution that introduces audio and video interaction capabilities to the intelligent conversational assistant.
[0018] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of this specification.
[0019] It should be noted that, for ease of description, only the parts related to the relevant technical solutions are shown in the accompanying drawings. In the absence of conflict, the embodiments in this specification and the features in the embodiments may be combined with each other. In addition, the words "first", "second", "third" and the like in the embodiments of this specification are only used for information distinction and do not play any limiting role. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, product or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, product or device. In the absence of further restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or device including the elements. In addition, in this specification, unless explicitly stated, "receiving and sending data" does not necessarily mean direct receiving and sending, but may be indirect receiving and sending. For example, when A receives data sent by B, it can be understood that A receives the data sent by B directly, or it can be understood that A receives the data sent by B indirectly through other entities such as C. Similarly, when B sends data to A, it can be understood that B sends the data directly to A, or it can be understood that B sends the data indirectly through other entities such as C. Here, C can be one entity, or two or more entities.
[0020] Furthermore, it should be noted that this specification uses specific words to describe the embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" refer to a certain feature, structure or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more in different places in this specification does not necessarily refer to the same embodiment. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory. Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it can be understood that the order of steps listed in the embodiments or flowcharts is only one way of executing the steps among many, and does not represent the only execution order. Therefore, when the claims involve method steps, changes and adjustments to the order of such steps, or parallelism between steps are also within the scope of protection of the claims.
[0021] Furthermore, it should be noted that the user data obtained in this manual (such as the user's video stream, voice conversation stream, etc.) is authorized by the user and does not involve user privacy.
[0022] The following describes and illustrates various embodiments provided in this specification in conjunction with the accompanying drawings.
[0023] First, the terms used in the embodiments of this specification are explained. It should be understood that this explanation is for a clearer understanding of the embodiments of this specification and does not necessarily constitute a limitation on the embodiments of this specification.
[0024] Automatic Speech Recognition (ASR): It is used to recognize human speech into text, that is, to convert human speech signals into text. It involves multiple processes including speech signal acquisition, feature extraction, acoustic model, language model and decoding.
[0025] Voice Activity Detection (VAD) is a signal processing technique designed to automatically identify speech in an audio signal and distinguish it from non-speech components (such as noise or silence). This technology is crucial for improving voice communication quality, reducing storage requirements, and optimizing the performance of speech recognition systems. It primarily focuses on the question of whether someone is currently speaking. In other words, VAD can be used to determine whether a user is speaking.
[0026] Text-to-Speech (TTS): used to convert text information into natural-sounding speech output.
[0027] A preset model is a trained artificial intelligence model. In this specification, preset models include a first preset model, a second preset model, and a third preset model. The first preset model is a multimodal model, such as a vision-language model (VLM). A vision-language model (VLM) is a multimodal model that integrates visual information (images or videos) with language information (text). Therefore, a VLM can recognize image content. The second preset model is a task model dedicated to processing specific tasks (such as object detection). The third preset model is a text model, specifically, a language model (LLM). A language model (LLM) is an artificial intelligence model that is a key component of natural language processing (NLP) technology. Based on the Transformer architecture, LLMs can understand and generate high-quality human language text and are used for conversational content generation in conversational interaction systems. It is worth noting that the embodiments of this specification do not limit the number of parameters supported by the preset models, with the goal of meeting actual application needs.
[0028] The technical solutions provided in the following embodiments of this specification are based on Figure 1 The technical architecture shown in is implemented. Figure 1As shown, in this technical architecture, low-latency, high-accuracy online human-computer audio and video interaction is achieved through an asynchronous link. The execution process of this asynchronous link includes the following: 1) An image acquisition device (such as a camera) continuously collects the user's video stream (the user's screen). The multimodal model in the image processor is called at a set frequency (e.g., once every second) to process multiple consecutive frames in the video stream. This generates image descriptions and stores them in a memory module. An example of an image acquisition device is a camera. The memory module is also called internal memory. An example of a multimodal model is a visual language model (VLM).
[0029] In addition to the multimodal model, the image processor also includes several other models, such as task models specifically designed to perform specific tasks. The inclusion of task models in the image processor is due to the fact that multimodal models typically analyze the entire image, making it difficult to extract features from local objects within the image for refined analysis. However, if the task being performed is more specific to an application, such as target detection or security, these tasks often require identification and analysis of specific individuals or objects. In this case, traditional task models are needed to extract features of the corresponding local objects in the image to perform the task. For example, in a skin health scenario, after an image acquisition device captures a video stream of a user's hand and collects a voice conversation stream in which the user asks about their skin condition, a multimodal model can be used to analyze multiple frames of the hand video stream to identify the presence of the arm and potential areas of erythema on the arm. Furthermore, the corresponding task model can be used to perform refined analysis of the identified erythema areas, extracting their specific features (such as size and density). Here, the specific features extracted from the local objects in the image using the task model will be stored in the memory module in the form of structured parameters.
[0030] In the memory model, the image description information for multiple consecutive frames in a video stream is stored in the form of text + structured parameters. The text is the description of the entire image (i.e., global description text), and the structured parameters are the structured description text of local objects in the image (i.e., local description text). This structured description text includes but is not limited to the position, size, category, color, density, and bounding box information of the local objects in the image. Local objects can refer to specific areas or elements in the image.
[0031] 2) Continuously collect user voice input through a sound pickup device (such as a microphone) and call the automatic speech recognition (ASR) function to convert the received user voice dialogue stream into text.
[0032] Automatic speech recognition (ASR) generally uses a fixed duration of silence (using a fixed VAD duration) to determine the end of a user's voice input. Specifically, VAD uses silence to detect the end of a user's voice input. For example, if a certain period of silence is detected, the user's voice input is considered to have ended. Of course, other methods can also be used to detect the end of a user's voice input, such as endpoint detection and detection based on a predefined time threshold. Endpoint detection not only considers the presence of silence but also factors such as speech rate and pitch variation to comprehensively determine whether the user's voice input has ended. Detection based on a predefined time threshold works by setting a fixed time window. If no new audio data is transmitted for longer than this window, the user's voice input is automatically determined to have ended.
[0033] 3) When the user finishes speaking, for example, if voice VAD is detected, the automatic speech recognition (ASR) stops and the text corresponding to the user's speech is concatenated with the relevant image description text information to form a prompt for the response generator. The response generator is built based on a text model, often also called a text dialogue model, which can be a language model (LLM).
[0034] 4) The prompt is fed into the response generator, which uses its internal text model to generate a response text (also called reply text) based on the prompt.
[0035] 5) Through the TTS function, the reply text is converted into the corresponding reply voice and played to the user.
[0036] The technical architecture mentioned above is based on the server and client implementation. Figure 1As shown in the technical architecture, the image acquisition device, audio acquisition device, and voice playback functions are deployed on the client, while the image processor (including multimodal models and task models), ASR, VAD, splicing operations, response generator (including text models), and TTS functions are deployed on the server. The server can be a server, server cluster, virtual server, or cloud-based system. The client can be, but is not limited to, a smartphone, smart wearable device, tablet, laptop, desktop computer, etc. The server provides corresponding functional services to the client, such as intelligent conversation. Users can initiate an intelligent video conversation through a browser, application (app), web application H5 (HyperText Markup Language 5, the fifth generation of HTML), light application (also known as mini-program, a lightweight application), or cloud application on the client. After the video conversation is connected to the intelligent conversation system on the server, an online human-computer audio and video conversation interaction is established between the client and the server.
[0037] thus, Figure 2A and Figure 2B The online human-computer interaction system (also called service system) provided by an embodiment of this specification is also shown, which includes a client 100 and a server 200. The client 100 is configured to display a service interface for human-computer interaction; initiate online human-computer audio and video interaction in response to an audio and video interaction initiation operation triggered by the service interface; collect a video stream and a voice conversation stream from the user during the online human-computer audio and video interaction; and send the video stream and the voice conversation stream to the server; wherein the video stream includes multi-angle images of a target object, wherein the target object includes a body part of the user; The server 200 is configured to generate prompt words based on the video stream and the voice conversation stream; the prompt words are used to guide the generation of a detection and analysis text of the target object based on the image description text information of the video stream, the detection and analysis text including the analysis results of the user's body part status; generate a reply text in response to the voice conversation stream based on the prompt words; the reply text including the detection and analysis text; convert the reply text into a reply voice, and send the reply text and the reply voice to the client; The client 100 is further configured to play the reply voice and display the reply text.
[0038] The online human-computer video interaction function described above can be integrated into intelligent conversation assistants in any field. The intelligent conversation assistant is deployed on the server side, and users can access the service interface for human-computer interaction provided by the intelligent conversation assistant through the client side.
[0039] The specific implementation of the functions of the server 200 and the client 100 will be described in detail in the following method embodiments and will not be described in detail here.
[0040] The technical solutions provided in this specification will be described below in the form of method embodiments.
[0041] Figure 3 The flowchart of an online human-computer interaction method provided by an embodiment of this specification is shown. The execution subject of this method is the server in the above system. Figure 3 As shown, the online human-computer interaction method includes the following steps: 102. Acquire a video stream and a voice dialogue stream on the user side in an online human-computer audio and video interaction; wherein the video stream includes multi-angle images of a target object, and the target object includes a body part of the user; 104. Generate prompt words based on the video stream and the voice dialogue stream; the prompt words are used to guide the generation of a detection and analysis text of the target object based on the image description text information of the video stream, the detection and analysis text including the analysis results of the user's body part status; 106. Generate a reply text for responding to the voice dialogue flow based on the prompt word, wherein the reply text includes the detection and analysis text; 108. Display the reply text, and / or play the reply voice corresponding to the reply text.
[0042] In the above 102, the online human-computer audio and video interaction is triggered by the user through the service interface for human-computer interaction displayed on the client, wherein the triggering method can be, but is not limited to, operating the corresponding audio and video interaction controls, voice triggering, inputting text triggering, etc. The service interface is provided by the corresponding intelligent dialogue assistant. Specifically, this service interface can be a traditional human-computer interaction interface that supports text dialogue or graphic dialogue. Therefore, in this specification, the intelligent dialogue assistant supports online human-computer audio and video interaction. The intelligent dialogue assistant can be an intelligent dialogue system applied to any field, such as health services (including medical and health services), beauty care, online education, remote office and other fields.
[0043] For example, taking the smart health manager (a smart conversation assistant that provides health services) as an example, the user enters the human-computer interaction service interface provided by the smart health manager through the client. Generally, the default interaction mode supported by this service interface is text conversation or graphic conversation. Figure 2A As shown, the user can trigger the start of online human-computer audio and video interaction by operating the "audio and video dialogue" control 11 on the service interface. Figure 2BAs shown, the user can also trigger the start of online human-computer audio and video interaction by inputting an audio and video conversation start instruction (e.g., "Please start audio and video conversation") in the input box 12 provided in the service interface. Of course, in other examples, the audio and video conversation start instruction can also be input by other means, such as voice input, such as operating the "voice input" control 13 to input the voice message "Please start audio and video conversation".
[0044] This solution integrates online human-computer audio and video interaction functions into the intelligent dialogue assistant, enabling online human-computer audio and video interaction. This allows users to communicate with the intelligent dialogue assistant (i.e., the intelligent dialogue system) more naturally and smoothly. This simplifies user operations and improves user experience, especially when it is necessary to upload complex image information or perform multi-view detection (such as hair status detection tasks).
[0045] After enabling online human-computer audio and video interaction, the user and the smart health manager will interact online via video call. During this audio and video interaction, the client will use different acquisition devices to capture the user's video stream and voice conversation stream respectively, and send them to the server for processing. For example, the camera captures the user's video input to form a video stream, and the microphone captures the user's voice input to form a voice conversation stream.
[0046] Based on the above content, the above step 102 of "obtaining the video stream and voice dialogue stream on the user side in the online human-computer audio and video interaction" includes: 1021. Receive the user-side video stream and voice conversation stream sent by the client; The video stream is user-side video data captured by the client using an image capture device. The image capture device may be, but is not limited to, a camera or other image capture device. The voice conversation stream is captured by the client using an audio pickup device, which may be, but is not limited to, a microphone.
[0047] In addition, the video stream contains multi-angle images of the target object. Multi-angle images refer to images of the target object captured from different perspectives (front, side, and top). The change in perspective can be caused by the position of the image acquisition device (such as a camera) or the rotation of the target object itself.
[0048] Among them, in different interaction fields, the target objects are different.
[0049] For example, in the field of health services, the target object may be a body part of the user, such as the face, head, hands, skin, hair, nails, etc.
[0050] For another example, in the field of remote office, the target object may be, for example, a target device that has an abnormality.
[0051] For example, in the field of online education, the target object can be a relevant test question; or the target object can also be the user, their body posture, and behavioral motion trajectory (in some cases, it may also include environmental elements or equipment with which the user interacts). For example, if the video stream is a dynamic video stream captured of a student conducting a scientific experiment (such as chemistry, biology, physics, medicine, etc.), the target object can be the student's body posture, key movement parts, operation behavior trajectory, and the experimental equipment with which they interact; and if the video stream is a dynamic video stream captured of a user practicing dancing, the target object can be the user, their body posture, and motion trajectory (especially key body parts and movement characteristics related to dance movements).
[0052] Furthermore, during online human-computer audio and video interaction, in order to collect video data that meets the requirements, users can be guided to shoot and receive corresponding shooting feedback in real time. When guiding users to shoot, appropriate shooting guidance prompts can be provided based on the type of interactive task.
[0053] In view of the above, the method provided in this embodiment may further include the following steps: S12, determining the current interactive task type; S14. Determine appropriate shooting guidance prompt information according to the interactive task type; S16: Output the shooting guidance prompt information to the user to guide the user to adjust the shooting posture.
[0054] The aforementioned collected video stream includes video clips captured by guiding the user to adjust the shooting posture.
[0055] In the above S12, the interactive task type refers to the type of feature interactive task performed between the intelligent dialogue assistant and the user during the human-computer audio and video interaction process. The interactive task defines the purpose, content, etc. of the interaction. In different fields, the interactive tasks are different. For example, in the field of health services, the interactive task may be a user health detection task, specifically, it may be a skin condition detection task, a hair condition detection task, an eye health detection task, etc. For another example, in the field of remote office, the interactive task may be, but is not limited to: device anomaly analysis tasks (such as hardware configuration, network connection, system crash, security vulnerability analysis tasks, etc.). For another example, in the field of online education, the interactive tasks may be, but are not limited to: answering test questions, analyzing experimental operation errors, analyzing student dance movements, etc.
[0056] In this embodiment, the above-mentioned interactive task type can be determined based on the user's autonomous triggering selection or can also be actively recommended by the intelligent dialogue assistant. This embodiment does not specifically limit the method for determining the interactive task type.
[0057] For example, after online human-computer audio and video interaction is initiated, the type of interaction task can be determined based on the user's voice input. For example, if the collected user voice input includes "I want to check my hair loss status," the type of interaction task can be determined to be "hair condition detection task." Alternatively, the online human-computer audio and video interaction interface can provide a text input entry (specifically, a text chat entry) through which the user can enter the desired interaction content, such as "I want to check my hair loss status." In this case, based on this text input, the type of interaction task can be determined to be "hair condition detection task."
[0058] For another example, after online human-computer audio and video interaction is started, a task window may pop up, displaying multiple interactive task options. The user can click on any of the multiple interactive task options, and the type of interactive task can be determined based on the interactive task option selected by the user. Alternatively, multiple interactive task options can be displayed on a certain side of the online human-computer audio and video interaction interface for the user to select, and the type of interactive task can be determined based on the interactive task option selected by the user. For example, if the user clicks on the "Start Hair Detection" option, the type of interactive task can be determined as the "Hair Status Detection Task."
[0059] For example, in some scenarios, the user may not explicitly state the task. In this case, the type of interactive task required can be automatically inferred based on the user's relevant information. For example, based on the user's historical online human-computer interaction information (including historical online human-computer text and / or graphic dialogue interactions, and historical online human-computer audio and video interactions), it can be determined that the user previously performed a "skin detection" or "face detection". In this case, it can be inferred that the user may need to perform a hair detection next. The inferred hair detection task can also be confirmed to the user. For example, the following voice message can be output: "Hello, I am ** Health Manager. Do you want to consult about hair issues?" Upon receiving the user's confirmation voice message (e.g., "Yes"), the interactive task type is finally determined to be a "hair detection task."
[0060] Based on the above example content, the above S12 “determining the interaction task type” may include any of the following methods: Method 1: respond to the task selection operation triggered by the user and determine the interactive task type according to the selected task item; wherein the task selection operation can be implemented by, but not limited to, voice, text input, or clicking. Or, Method 2: Determine the interactive task type recommended for the user based on the user's relevant information; wherein the relevant information includes the user's historical online human-computer interaction information.
[0061] In the above S14~S16, in order to capture user-side videos that meet the requirements and provide users with better interactive services (such as problem consultation services), appropriate shooting guidance prompt information can be determined according to the type of interactive task, so as to guide the user to shoot through the shooting guidance prompt information to capture appropriate image frames.
[0062] The form of the shooting guidance prompt information may be, but is not limited to, a visual prompt and / or a voice prompt. The visual prompt may be, for example, at least one of an animation demonstration, a text prompt, a schematic diagram prompt, and the like.
[0063] For example, in the interactive task of hair condition detection, the determined shooting guidance prompt information can be a simple tutorial animation explaining the angles from which the user should take a photo of their head (such as the top of the head, forehead, side, and back of the head). This tutorial animation can be played in a location (such as the lower right corner) of the online human-computer audio and video interaction interface. For example, the tutorial animation may display shooting guidance content such as: "Please align the top of your head with the center of the screen" and "Now, slowly rotate your head to ensure that both sides of your hairline are captured." Guided by this tutorial animation, the user can adjust the shooting angle by adjusting their head posture or the camera position, thereby capturing their hair from multiple angles. During the process of capturing their hair from multiple angles, the camera will capture a continuous segment of image data, which represents the video clip captured under the guidance of the user. Furthermore, during the user's shooting guidance, the user will also receive timely voice feedback, such as feedback on the image quality (clearness, whether there is any obstruction, etc.), whether the current shooting angle meets the requirements, and whether the shooting is complete.
[0064] Therefore, the aforementioned collected user-side video stream includes video clips captured by guiding the user to adjust the shooting posture.
[0065] Here, in online human-computer audio and video interaction, by guiding users to take images from different angles and providing instant voice feedback, the comprehensiveness and accuracy of image information collection can be ensured, which is conducive to the accuracy of subsequent reply text generation.
[0066] Furthermore, the video stream continuously collected by the image acquisition device can be cached first, so that it can be called after the set time interval is reached. Figure 1 The image processor shown in FIG is used to process the image and generate corresponding image description text information.
[0067] Therefore, the method provided in this embodiment may further include the following steps: S22. Perform image recognition analysis on the acquired video stream at set time intervals to generate image description text information corresponding to the video stream.
[0068] S24: Storing the image description text information for subsequent use in generating the prompt word.
[0069] The specific implementation is to call Figure 1 The image processor shown in the figure performs image recognition analysis on the video stream continuously collected by the image acquisition device, thereby generating image description text information corresponding to the video stream and storing the image description text information in the corresponding memory.
[0070] In the above description, the time interval can be flexibly set according to actual conditions and is not limited in this embodiment. For example, the image processor can be called once every 1 second, that is, the image processor is called once every 1 second. In other words, it can be understood that every time a video stream of 1 second is collected by the image acquisition device, the image processor is called once to perform image recognition analysis on the video stream of 1 second.
[0071] The image processor includes a first preset model and a second preset model, the second preset model is used to perform overall recognition analysis on the image, and the third preset model is used to perform local recognition analysis on the image.
[0072] Based on the above content, in a specific achievable technical solution, the above step S22 of "performing image recognition analysis on the collected video stream to generate descriptive text information corresponding to the video stream" includes: S222: Calling a first preset model to perform global recognition on a plurality of consecutive frames of images in a video stream to generate a global description text; S224: Calling a second preset model to perform local recognition on local objects in the continuous multi-frame images to generate local description text.
[0073] Furthermore, the step S24 of “storing the image description text information” includes: S242: Associate and store the global description text and the local description text corresponding to the continuous multiple frames of images, so as to prepare for subsequent generation of the prompt word.
[0074] In the above S222 and S224, the first preset model and the second preset model may be respectively Figure 1The multimodal model and task model shown in the figure are used to perform a specific task. In addition, the generated global description text refers to the text information content that provides a general description of the entire image of multiple consecutive frames. The generated local description text refers to the detailed local object description information generated by performing local recognition on local objects in multiple consecutive frames of images to extract the object features of the local objects. Among them, the local object can be, but is not limited to, a specific local area or element in the image (such as hair, erythema, etc.), and the corresponding local description text contains at least one of the position, size, density, color, category, morphology, etc. of the local object.
[0075] For example, the first preset model can be called once every 1 second (that is, the first preset model is called at a frequency of once 1 second) to perform global recognition of multiple consecutive frames of images in the video stream collected by the image acquisition device within 1 second, thereby generating global description text of the multiple consecutive frames of images; further, the second preset model can also be called to perform feature extraction on local objects identified from the multiple consecutive frames of images, thereby generating local description text of the local objects, etc.
[0076] It should be noted that step S224 is not a mandatory step but an optional step. Step S224 may be triggered only when it is determined that local object recognition is required, depending on the type of interactive task. That is, the method provided in this embodiment may further include the following steps: S223: When it is determined that local object recognition is required according to the interactive task type, the above step S224 is triggered to be executed.
[0077] Regarding the determination of the interactive task type here, reference may be made to the relevant contents described in other aforementioned embodiments, and no further details will be given here.
[0078] For example, if the interactive task type is "hair status detection task", it is determined that local object recognition is required, and the local object to be recognized is the hair area in the image.
[0079] For another example, if the interactive task type is "skin erythema detection task", it is determined that local object recognition is required, and the local object to be recognized is the erythema element in the image.
[0080] In the above S242, the global description text and the local description text of the continuous multiple frames of images can be associated and stored in the corresponding memory to prepare for the subsequent support prompt word generation. Figure 1 The memory module shown in FIG. During associative storage, the local description text is stored in the form of structured parameters. For example, the local description text is stored in the form of: local object identifier (such as name): **; location: **; size: **; density: **.
[0081] For the user-side voice conversation stream collected by a sound pickup device (such as a microphone), the automatic speech recognition (ASR) function will be used to convert the voice signal in the voice conversation stream into the corresponding conversation text. After detecting the end of the voice conversation stream (for example, when a period of silence is detected, it can be determined that the current voice conversation stream has ended, that is, the current voice input on the user side has ended), the automatic speech recognition will be stopped and the prompt word generation stage will be entered. In the prompt word generation stage, the image description text information that is more semantically relevant to the conversation text will be searched from the stored information, and the conversation text will be spliced with the searched image description text information to obtain the corresponding prompt word. The prompt word is used as Figure 1 The input parameters of the response generator (specifically the third preset model therein) shown in FIG, so that the response generator generates a reply text based on the input prompt word to respond to the voice stream on the user side.
[0082] Based on the above content, in a specific feasible technical solution, the above 104 “generating prompt words according to the video stream and the voice dialogue stream” includes: 1042. Convert the voice dialogue stream into text to generate corresponding dialogue text; 1044. Retrieve the image description text information related to the conversation text from the stored information; 1046. Generate the prompt word based on the conversation text and the retrieved image description text information.
[0083] The prompt words generated here may include role settings, detection and analysis task instructions, output constraints, etc. for the third preset model, which are used to guide the third preset model to use the image description text information of the video stream as context to perform corresponding detection and analysis tasks for the target object and generate matching detection and analysis text, so as to respond to user consultation questions expressed in the dialogue text based on the obtained detection and analysis text.
[0084] That is, the prompt word is used to guide the generation of the target object detection analysis text based on the image description text information of the video stream, wherein the detection analysis text includes the user's body part status analysis results, such as the detection analysis results of the hair / skin status.
[0085] The above-mentioned step 106 of “generating a reply text for responding to the voice dialogue flow based on the prompt word” may include: 1062. Input the prompt word into a third preset model, and execute the third preset model to output the reply text.
[0086] In the above, after converting the voice dialogue stream into the corresponding dialogue text, the retrieval enhancement (RAG) module can be used to retrieve image description text information that is semantically related to the dialogue text from the memory, and then the dialogue text and the retrieved image description text information are spliced to generate prompt words for inputting the third preset model.
[0087] The third preset model is a text model, specifically, a language model (LLM). The response text generated by the third preset model based on the prompt word can be used to respond to the user's consultation question (such as hair loss) in the voice conversation flow.
[0088] For example, in the healthcare service field, assume the generated prompt is, "You are a smart health assistant. Please generate an objective initial hair health detection and analysis result based on the video stream's image description text information and the user's conversation text. The output requirement is to be in Chinese and divided into three parts: summary, observation, and suggestions, with a professional and gentle tone." The video stream's image description text information is "***" and the user's conversation text is "Please detect my hair loss." After inputting this prompt into a third preset model and executing it, the output response text may be, for example, "From the image, your forehead hairline has a slight tendency to recede; multi-angle images show that the hair on the top of your head is sparse, and some areas of your scalp are visible, which may be related to genetics or stress; it is recommended to maintain a good sleep and rest schedule and consult a dermatologist for professional testing if necessary." In this case, the response text containing "The forehead hairline has a slight tendency to recede, the hair on the top of your head is sparse, and some areas of your scalp are visible" represents the user's hair condition detection and analysis result.
[0089] The above description of the processing content for video stream and voice dialogue stream is combined with Figure 1 It can be seen that this embodiment actually processes the video stream collected by the camera and the user voice input collected by the microphone in an asynchronous manner, which can improve the response speed of the call and the user experience. In this asynchronous manner, accordingly, Figure 1The image processor (specifically, the multimodal model) and response generator (specifically, the text model) shown in the figure operate asynchronously. The reason for not adopting a synchronous operation scheme, such as the multimodal model and text model, is that synchronous operation would result in a longer delay, which would increase the response time and cause the user to wait for a long time for the response. Furthermore, synchronous operation, which intelligently combines the current and latest video dialogue, requires high synchronization of image and audio, making it difficult to guarantee the effect. Of course, there is also an alternative solution, in which the multimodal model can be used to process the video stream while also generating the response text. This means that only the multimodal model is used to process the video stream and generate the response text. However, this solution has a slightly longer operation delay than the asynchronous operation scheme, and the multimodal model requires more aligned training data for the response, which is more expensive.
[0090] Furthermore, the reply text can be converted into the corresponding reply voice through the TTS function and sent to the user side client for playback. Of course, the reply text can also be sent to the client so that the reply text can be displayed while the reply voice is played on the client for easy understanding by the user.
[0091] Therefore, a specific technical solution for "displaying the reply text and / or playing the reply voice corresponding to the reply text" in the above 108 may include: 1082. Convert the reply text into reply voice; 1084. Send the reply voice and the reply text to the user-side client, so that the client plays the reply voice and / or displays the reply text.
[0092] Considering that some fields, such as healthcare services, often use specialized terms, simply playing back the voice response may cause users to have difficulty understanding it. Therefore, while playing back the voice response, the corresponding text is also displayed simultaneously, making it easier for users to understand the voice response through textual understanding.
[0093] In general, this embodiment upgrades the traditional multimodal interaction of the intelligent dialogue assistant to an audio and video interaction mode, which can support users to conduct multimodal interaction for natural abortion, optimize the user interaction experience, and has strong practicality and interactive flexibility in actual applications. It can promote the leapfrog development of intelligent dialogue assistants and enhance their commercial value and market competitiveness.
[0094] This specification also provides another online human-computer interaction method, the execution subject of which is the client in the above system. Figure 4 As shown, the online human-computer interaction method includes the following steps: 202. Displaying a service interface for human-computer interaction; 204. In response to the audio and video interaction start operation triggered through the service interface, start online human-computer audio and video interaction; 206. During the online human-computer audio and video interaction process, a video stream and a voice conversation stream are collected from the user side; wherein the video stream includes multi-angle images of a target object, and the target object includes a body part of the user; 208. Send the video stream and voice dialogue stream to the server; 210. Receive the reply text and reply voice corresponding to the reply text returned by the server; wherein the reply text is generated based on a prompt word to respond to the voice dialogue stream; the prompt word is generated based on the video stream and the voice dialogue stream to guide the generation of a detection and analysis text of the target object based on the image description text information of the video stream, the detection and analysis text including the user's body part status analysis result; the reply text includes the detection and analysis text; 212. While playing the reply voice, display the reply text.
[0095] The specific implementation of the above steps provided in this embodiment can refer to the relevant content in other embodiments, and will not be described in detail here. In addition, the method provided in this embodiment can also include some steps disclosed in other embodiments, and can also refer to the relevant content in other embodiments, and will not be described in detail here.
[0096] This specification also provides another online human-computer interaction method, in which some steps (such as 302, 304, and 306) are executed by the client, and the remaining steps (such as 308, 3010, and 3012) are executed by the server. In addition, the application scenario of the method is an intelligent dialogue assistant for health services (such as a health manager). Specifically, see Figure 5 As shown, the online human-computer interaction method includes the following steps: 302. Displaying a health service interface, wherein the health service interface has audio and video interactive controls; 304. In response to the operation of the audio and video interaction control, start online human-computer audio and video interaction; 306. During the online human-computer audio and video interaction process, collecting the user's biological video stream and voice dialogue stream; 308. Generate prompt words based on the biological video stream and the voice dialogue stream; 310. Input the prompt word into a third preset model, execute the third preset model and output a reply text; the reply text is used to respond to the health question asked by the user, and the health question is extracted from the voice dialogue stream; 312. Play the reply text to the user.
[0097] In the above, the health service interface is as follows Figure 2A and Figure 2B The service interface for human-computer interaction and the audio and video interaction controls shown in Figure 2A The “audio and video conversation” control 11 is shown in FIG.
[0098] Furthermore, the biometric video stream is a continuous sequence of images of a user's body parts (such as the face, head, hands, and neck) captured by an image acquisition device. Image information related to the user's biometric characteristics, such as hair condition and skin condition, can be extracted from this biometric video stream to support health checkup analysis.
[0099] The specific implementation of the above steps provided in this embodiment can refer to the relevant content in other embodiments, and will not be described in detail here. In addition, the method provided in this embodiment can also include some steps disclosed in other embodiments, and can also refer to the relevant content in other embodiments, and will not be described in detail here.
[0100] Combined with the above Figures 3-5 This specification describes specific embodiments. It should be noted that for each specific embodiment described, other embodiments are within the scope of the appended claims, and that, in some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0101] The following introduces the device embodiments corresponding to the various method embodiments provided in this specification.
[0102] Figure 6 FIG. 1 shows a schematic diagram of the structure of an online human-computer interaction device provided by an exemplary embodiment of this specification. Figure 6 As shown, the device includes: an acquisition module 42, a generation module 44, and a playback module 46. An acquisition module 42 is configured to acquire a video stream and a voice dialogue stream from a user side in an online human-computer audio and video interaction; wherein the video stream includes multi-angle images of a target object, and the target object includes a body part of the user; A generating module 44 is configured to generate prompt words based on the video stream and the voice dialogue stream; the prompt words are used to guide the generation of a detection analysis text of the target object based on the image description text information of the video stream, the detection analysis text including the analysis results of the user's body part status; The generating module 44 is further configured to generate a reply text for responding to the voice dialogue flow based on the prompt word, wherein the reply text includes the detection and analysis text; The playing and displaying module 44 is configured to display the reply text and / or play the reply voice corresponding to the reply text.
[0103] In one embodiment, the device further includes: a recognition and analysis module and a storage module. The recognition and analysis module is configured to perform image recognition analysis on the acquired video stream at set time intervals to generate textual descriptions of the images corresponding to the video stream. The storage module is configured to store the textual descriptions of the images for subsequent use in generating the prompt word.
[0104] In one embodiment, the recognition and analysis module, when used to perform image recognition analysis on a captured video stream and generate image description text information corresponding to the video stream, is specifically configured to: invoke a first preset model to perform global recognition on multiple consecutive image frames in the video stream to generate global description text; and invoke a second preset model to perform local recognition on local objects in the multiple consecutive image frames to generate local description text. Furthermore, the storage module, when used to store the image description text information, is specifically configured to: associate and store the global description text and the local description text corresponding to the multiple image frames.
[0105] In one embodiment, the device further includes a determination module and a trigger module. The determination module is configured to determine the type of the current interactive task. The trigger module is configured to, upon determining that local object recognition is required based on the interactive task type, trigger the invocation of the second preset model to perform local recognition of the local objects in the multiple image frames and generate local description text. The local description text includes at least one of the location, size, density, color, and category of the local objects.
[0106] In one embodiment, the above-mentioned determination module, when used to determine the type of interactive task, is specifically used to: respond to a task selection operation triggered by the user, and determine the type of interactive task based on the selected task item; or, determine the type of interactive task recommended for the user based on the user's relevant information; wherein the relevant information includes the user's historical online human-computer interaction information.
[0107] In one embodiment, the determination module is further configured to determine appropriate shooting guidance prompt information based on the current interactive task type. Furthermore, the output module is further configured to output the shooting guidance prompt information to the user to guide the user in adjusting the shooting posture. The video stream includes a video clip captured by guiding the user to adjust the shooting posture, wherein the shooting guidance prompt information includes a visual prompt and / or a voice prompt.
[0108] In one embodiment, when generating prompt words based on the video stream and the voice conversation stream, the generation module is specifically configured to: convert the voice conversation stream into text to generate corresponding conversation text; retrieve the image description text information related to the conversation text from stored information; and generate the prompt words based on the retrieved image description text information and the conversation text. Furthermore, when generating a response text in response to the voice conversation stream based on the prompt words, the generation module is specifically configured to: input the prompt words into a third preset model, execute the third preset model, and output the response text.
[0109] In one embodiment, the above-mentioned playback module, when used to display the reply text and / or play the reply text to the user, is specifically used to: convert the reply text into reply voice; send the reply voice and the reply text to the user-side client, so that the client plays the reply voice and / or displays the reply text.
[0110] Figure 7 FIG. 1 shows a schematic diagram of the structure of an online human-computer interaction device provided by another exemplary embodiment of this specification. Figure 7As shown, the device includes: a display module 52, a start module 54, a collection and transmission module 56, a receiving module 58, and a broadcast display module 510. The display module 52 is configured to display a service interface for human-computer interaction. The start module 54 is configured to initiate online human-computer audio and video interaction in response to an audio and video interaction start operation triggered through the service interface. The collection and transmission module 56 is configured to collect the user's video stream and voice conversation stream during the online human-computer audio and video interaction process; wherein the video stream includes multi-angle images of a target object, which includes the user's body parts; and further configured to send the video stream and voice conversation stream to the server. The receiving module 58 is configured to receive the response text and the response voice corresponding to the response text returned by the server; wherein the response text is generated based on a prompt word to respond to the voice conversation flow; the prompt word is generated based on the video stream and the voice conversation flow to guide the generation of a detection and analysis text of the target object based on the image description text information of the video stream, wherein the detection and analysis text includes the analysis results of the user's body part status; and the response text includes the detection and analysis text. The playing and displaying module 510 is configured to play the reply voice and display the reply text at the same time.
[0111] Figure 8 FIG. 1 shows a schematic diagram of the structure of an online human-computer interaction device provided by another exemplary embodiment of this specification. Figure 8 As shown, the device includes: a display module 62, a startup module 64, a collection module 66, a generation module 68, an execution module 610, and a playback module 612. The display module 62 is used to display a health service interface having audio and video interaction controls. The startup module 64 is used to initiate online human-computer audio and video interaction in response to an operation on the audio and video interaction controls. The collection module 66 is used to collect the user's biological video stream and voice dialogue stream during the online human-computer audio and video interaction process; wherein the biological video stream is a continuous sequence of images containing the user's body parts. The generation module 68 is used to generate prompt words based on the biological video stream and the voice dialogue stream. The execution module 610 is used to input the prompt words into a third preset model, execute the third preset model, and output a reply text; the reply text is used to respond to the user's health questions, and the health questions are extracted from the voice dialogue stream. The playback module 612 is used to play the reply voice corresponding to the reply text to the user.
[0112] What needs to be explained here about each of the above-mentioned devices is that: each of the devices provided above can implement the technical solutions described in the corresponding method embodiments above. The specific implementation principles of each of the above-mentioned modules or units can refer to the relevant content in the corresponding method embodiments above, and will not be described in detail here. In addition, for the convenience of description, the above devices are described by being divided into various modules or units according to their functions. Of course, when implementing one or more of the present specifications, the functions of each module or unit can be implemented in the same or more software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0113] In addition, the embodiments of this specification also provide an electronic device. Figure 9 As shown, the electronic device 700 includes: a memory 71 and a processor 72.
[0114] The memory 71 can be implemented by at least one volatile or non-volatile memory device of any type, or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. Furthermore, all or part of the memory can be integrated with the processor. The memory can include both removable and non-removable components.
[0115] The processor 72 may include one or more general-purpose processors and / or special-purpose processors.
[0116] Furthermore, the memory 71 may include a non-transitory computer-readable medium having executable program instructions 712 stored therein (e.g., compiled or non-compiled program logic and / or machine code). The processor 72 is capable of executing the program instructions 712 stored in the memory to implement any method, process, or function disclosed in this specification and / or the accompanying drawings. Furthermore, the execution of the program instructions 712 by the processor 72 may cause the processor to use corresponding data 711.
[0117] For example, the program instructions 712 may include an operating system 7122 (e.g., an operating system kernel, device drivers, and / or other modules) and one or more application programs 7121 (e.g., a browser, a social application, or a game application) installed on the electronic device 700. Similarly, the data 711 may include operating system data 7112 and application data 7111. The operating system data 7112 is primarily accessible to the operating system 7122, while the application data 7111 is primarily accessible to one or more application programs 7121. The application data 7111 may be located in a file system visible or hidden to the user of the electronic device 400.
[0118] The application 7121 can communicate with the operating system 7122 through one or more application programming interfaces (APIs). These APIs help the application 7122 read and / or write application data, transmit or receive information via communication components, receive or display information on a user interface, etc. In some terms, the application 7121 can be simply referred to as an "app". In addition, the application 7121 can be downloaded to the electronic device through one or more online application stores or application markets. However, the application 7121 can also be installed on the electronic device 400 through other means, such as through a web browser or a physical interface on the electronic device 700 (e.g., a USB port). Further, if Figure 9 As shown, the electronic device also includes: a communication component 73, a display 74, a power component 75, an audio component 76, a user interface 77 and other components. Figure 9 Only some components are shown schematically, which does not mean that the electronic device 700 only includes Figure 9 In addition, Figure 9 The components in the dotted box are optional components, not mandatory components, and the specific components depend on the product form of the electronic device 700. The electronic device 700 of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone or an IOT device, or a server device such as a conventional server, a cloud server or a server array, or an integrated device of a terminal device and a server device. If the electronic device 700 of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, etc., it can include Figure 9 If the electronic device 700 of this embodiment is implemented as a conventional server, cloud server or server array and other server devices, it may not include Figure 9 Components within the dotted box.
[0119] The communication component 73 is configured to facilitate wired or wireless communication between the device in which the communication component resides and other devices. The device in which the communication component 73 resides can access a wireless network based on a communication standard, such as a 2G, 3G, 4G / LTE, 5G, or other mobile communication network, or a combination thereof. In one exemplary embodiment, the communication component 73 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In a specific implementation, the communication component 73 includes a communication interface that enables the electronic device 700 to communicate with other devices, access networks, and transmission networks using analog or digital modulation. For example, the communication interface may include a chipset and an antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface may be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, a Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface may also support other physical layer interfaces and standard or proprietary communication protocols. The communication interface may also include multiple physical communication interfaces, such as a Wi-Fi interface, a Bluetooth interface, and a wide-area wireless interface.
[0120] The display 74 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, it may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can detect not only the boundaries of a touch or slide action, but also the duration and pressure associated with the touch or slide action.
[0121] The power supply assembly 75 provides power to various components of the device in which it is located. The power supply assembly 75 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply assembly is located.
[0122] The audio component 76 can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), which is configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals can be further stored in a memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0123] The user interface 77 described above includes both receiving user input and providing output to the user. Thus, the user interface 77 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensing panel, computer mouse, trackball, joystick, microphone, still camera, and video camera. It may also include output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, or other similar devices known or developed in the future. The user interface 77 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, or other similar devices known or developed in the future. In certain embodiments, the user interface 77 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, the electronic device 700 may support remote access from other devices, such as via a communication interface or another physical interface (not shown). The user interface 77 may be configured to receive user input, and its position and movement may be indicated by an indicator or cursor as described herein. The user interface 77 may also be configured as a display device for rendering or displaying text snippets.
[0124] Accordingly, an embodiment of the present specification also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement each step in the above method embodiment. The computer-readable storage medium includes volatile or non-volatile or a combination thereof, and may be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission medium. Furthermore, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, the computer is caused to execute the following Figures 3 to 5 Described method.
[0125] The embodiments of this specification also provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the following Figures 3 to 5 Described method.
[0126] Those skilled in the art will appreciate that, in one or more of the above examples, the functions described in the various embodiments disclosed in this specification may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0127] The specific implementation methods described above further illustrate in detail the purposes, technical solutions and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above description is only the specific implementation methods of the multiple embodiments disclosed in this specification, and is not intended to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the multiple embodiments disclosed in this specification should be included in the protection scope of the multiple embodiments disclosed in this specification.
Claims
1. An online human-computer interaction method, characterized in that: include: Acquire a video stream and a voice dialogue stream on the user side in an online human-computer audio and video interaction; wherein the video stream includes multi-angle images of a target object, and the target object includes a body part of the user; Generate prompt words based on the video stream and the voice dialogue stream; the prompt words are used to guide the generation of a detection analysis text of the target object based on the image description text information of the video stream, the detection analysis text including the analysis results of the user's body part status; Based on the prompt word, generating a reply text for responding to the voice dialogue flow, wherein the reply text includes the detection and analysis text; Display the reply text, and / or play the reply voice corresponding to the reply text.
2. The method according to claim 1, characterized in that Also includes: Performing image recognition analysis on the acquired video stream at set time intervals to generate image description text information corresponding to the video stream; The image description text information is stored for subsequent use in generating the prompt word.
3. The method according to claim 2, characterized in that Perform image recognition analysis on the collected video stream to generate image description text information corresponding to the video stream, including: Calling a first preset model to perform global recognition on a plurality of consecutive frames of images in the video stream to generate a global description text; Calling a second preset model to perform local recognition on local objects in the continuous multi-frame images to generate local description text; And, storing the image description text information, including: The global description text and the local description text corresponding to the multiple frames of images are associated and stored.
4. The method according to claim 3, characterized in that Also includes: Determine the current interactive task type; When it is determined according to the interactive task type that local object recognition is required, triggering the calling of the second preset model to perform local recognition on the local objects in the multiple frames of images and generate local description text; The local description text includes at least one of the position, size, density, color, and category of the local object.
5. The method according to claim 4, characterized in that Determine the interaction task type, including: In response to a task selection operation triggered by a user, determining the interactive task type according to the selected task item; or, Determine the type of interactive task recommended for the user based on the relevant information of the user; wherein the relevant information includes the historical online human-computer interaction information of the user.
6. The method according to any one of claims 1 to 5, characterized in that Also includes: Determine the appropriate shooting guidance prompt information based on the current interactive task type; Outputting the shooting guidance prompt information to the user to guide the user to adjust the shooting posture; The video stream includes video clips captured by guiding the user to adjust the shooting posture; The shooting guidance prompt information includes visual prompts and / or voice prompts.
7. The method according to any one of claims 2 to 5, characterized in that Generating prompt words according to the video stream and the voice dialogue stream includes: Converting the speech dialogue stream into text to generate corresponding dialogue text; Retrieving the image description text information related to the conversation text from the stored information; generating the prompt word based on the retrieved image description text information and the conversation text; And, based on the prompt word, generating a reply text for responding to the voice dialogue flow, including: The prompt word is input into a third preset model, and the third preset model is executed to output the reply text.
8. The method according to any one of claims 1 to 5, characterized in that Displaying the reply text and / or playing the reply voice corresponding to the reply text includes: Converting the reply text into reply speech; The reply voice and the reply text are sent to a user-side client, so that the client plays the reply voice and / or displays the reply text.
9. An online human-computer interaction method, characterized in that: include: Displaying a service interface for human-computer interaction; In response to the audio and video interaction start operation triggered by the service interface, start online human-computer audio and video interaction; During the online human-computer audio and video interaction process, a video stream and a voice dialogue stream on the user side are collected; wherein the video stream includes multi-angle images of a target object, and the target object includes a body part of the user; Sending the video stream and voice dialogue stream to the server; Receive a reply text and a reply voice corresponding to the reply text returned by the server; wherein the reply text is generated based on a prompt word to respond to the voice dialogue stream; the prompt word is generated based on the video stream and the voice dialogue stream to guide the generation of a detection and analysis text of the target object based on the image description text information of the video stream, the detection and analysis text including an analysis result of a user's body part status; the reply text includes the detection and analysis text; The reply text is displayed while the reply voice is played.
10. An online human-computer interaction method, characterized in that: include: Displaying a health service interface, wherein the health service interface has audio and video interactive controls; In response to the operation of the audio and video interaction control, starting online human-computer audio and video interaction; During the online human-computer audio and video interaction process, a biological video stream and a voice dialogue stream of the user are collected; wherein the biological video stream is a continuous image sequence containing body parts of the user; generating prompt words based on the biological video stream and the voice dialogue stream; Inputting the prompt word into a third preset model, executing the third preset model to output a reply text; the reply text is used to respond to the health question consulted by the user, and the health question is extracted from the voice dialogue flow; Play the reply voice corresponding to the reply text to the user.
11. A service system, characterized in that: include: The client is used to display the service interface for human-computer interaction; In response to the audio and video interaction start operation triggered by the service interface, start online human-computer audio and video interaction; During the online human-computer audio and video interaction process, a video stream and a voice conversation stream from the user side are collected; the video stream and the voice conversation stream are sent to the server; wherein the video stream contains multi-angle images of a target object, and the target object includes a body part of the user; The server is configured to generate prompt words based on the video stream and the voice dialogue stream; the prompt words are used to guide the generation of a detection and analysis text of the target object based on the image description text information of the video stream, the detection and analysis text including the analysis results of the user's body part status; based on the prompt words, generate a reply text for responding to the voice dialogue stream, the reply text including the detection and analysis text; convert the reply text into a reply voice, and send the reply text and the reply voice to the client; The client is also used to play the reply voice and display the reply text at the same time.
12. An online human-computer interaction device, characterized in that: include: An acquisition module is used to acquire the video stream and voice dialogue stream of the user side in the online human-computer audio and video interaction; wherein the video stream includes multi-angle images of the target object, and the target object includes the user's body parts; a generation module configured to generate prompt words based on the video stream and the voice dialogue stream; and generate a reply text for responding to the voice dialogue stream based on the prompt words; the prompt words are used to guide the generation of a detection and analysis text of the target object based on the image description text information of the video stream, the detection and analysis text including an analysis result of a body part status of the user; The playing and displaying module is used to display the reply text and / or play the reply voice corresponding to the reply text.
13. An online human-computer interaction device, characterized in that: include: A display module, used to display a service interface for human-computer interaction; A starting module, configured to respond to an audio and video interaction starting operation triggered through the service interface and start online human-computer audio and video interaction; A collection module, configured to collect a video stream and a voice conversation stream from a user during the online human-computer audio and video interaction process; wherein the video stream includes multi-angle images of a target object, and the target object includes a body part of the user; A sending module, used for sending the video stream and voice dialogue stream to the server; a receiving module, configured to receive a reply text and a reply voice corresponding to the reply text returned by the server; wherein the reply text is generated based on a prompt word to respond to the voice dialogue stream; the prompt word is generated based on the video stream and the voice dialogue stream to guide the generation of a detection and analysis text of the target object based on the image description text information of the video stream, the detection and analysis text including an analysis result of a user's body part status; and the reply text includes the detection and analysis text; The playing and displaying module is used to play the reply voice and display the reply text at the same time.
14. An online human-computer interaction device, characterized in that: include: A display module, configured to display a health service interface having audio and video interaction controls; A starting module, configured to respond to an operation on the audio and video interaction control and start online human-computer audio and video interaction; An acquisition module, configured to acquire a user's biological video stream and voice dialogue stream during the online human-computer audio and video interaction process; wherein the biological video stream is a continuous image sequence containing body parts of the user; A generation module, configured to generate prompt words based on the biological video stream and the voice dialogue stream; an execution module, configured to input the prompt word into a third preset model, execute the third preset model and output a reply text; the reply text is used to respond to the health question consulted by the user, the health question being extracted from the voice dialogue stream; The playing module is used to play the reply voice corresponding to the reply text to the user.
15. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores executable program instructions, and when the processor executes the program instructions, the method according to any one of claims 1 to 10 is implemented.
16. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 10.
17. A computer program product, characterized in that The computer program product includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Knowledge question and answer method, device and equipment and storage medium
CN116561276A
Image fuzziness evaluation method and device, electronic equipment and storage medium
CN117218075A
Digital human interaction method and system based on multiple modes
CN119292466A
Man-machine conversation method, device and equipment and computer readable storage medium
CN119312817A
Digital human real-time interaction system and digital human real-time interaction method
CN119440254A