Online human-computer interaction method, system, device, electronic equipment, storage medium and program product
By introducing audio and video interaction capabilities into the intelligent dialogue assistant, the shortcomings of text-based dialogue interaction methods are addressed, enabling natural multimodal interaction between users and the system, and improving interaction efficiency and accuracy, especially in complex image information and multi-view detection tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2026-03-27
AI Technical Summary
Existing intelligent dialogue assistants struggle to meet user needs when dealing with complex image information and multi-view detection. Their text-based dialogue interaction methods are cumbersome, and they lack the ability to perceive users' real-time visual information, making it impossible to achieve dynamic recognition and real-time feedback.
It introduces audio and video interaction capabilities, which generate prompt words to guide image description text information by collecting video streams and voice dialogue streams from the user side, generate echo text and play echo voice, and support multi-angle image detection and analysis to improve interaction efficiency and accuracy.
It enables natural and smooth multimodal interaction between users and the intelligent dialogue system, simplifies user operations, and improves user experience, especially when it is necessary to upload complex image information or perform multi-view detection, providing instant feedback and higher interaction accuracy.
Smart Images

Figure CN120723950B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the technical field of artificial intelligence, and in particular to an online human-computer interaction method, system, device, electronic equipment, storage medium and program product. BACKGROUND
[0002] Intelligent dialogue assistants, which can interact with users in dialogue, have been widely used in various fields such as medical health, and their mission is to help users solve problems or chat. At present, intelligent dialogue assistants mainly use text dialogue as the main interaction mode. With the development of multi-modal technology, some intelligent dialogue assistants have gradually introduced image-text dialogue capabilities, allowing users to upload images when asking questions, so as to better understand and reply to user questions in combination with image content. However, when the user needs to upload complex image information, such as multiple images, videos, or even real-time feedback interaction, this image-text dialogue interaction mode is difficult to meet the needs of users.
[0003] Therefore, it is urgent to introduce a new interactive dialogue capability for intelligent dialogue assistants. SUMMARY
[0004] The embodiments in the present specification provide an online human-computer interaction method, system, device, electronic equipment, storage medium and program product, which realizes the introduction of audio-video interaction capability in intelligent dialogue assistants, and is beneficial to meet the more complex problem needs of users. Among them,
[0005] In a first embodiment, an online human-computer interaction method is provided in the present specification. The method comprises:
[0006] Collecting a video stream and a voice dialogue stream on the user side in online human-computer audio-video interaction; wherein the video stream contains multi-angle images of a target object, and the target object includes a body part of the user;
[0007] Generating a prompt word according to the video stream and the voice dialogue stream; the prompt word is used to guide the generation of a detection analysis text of the target object based on image description text information of the video stream, and the detection analysis text contains a body part state analysis result of the user;
[0008] Generating a reply text for responding to the voice dialogue stream based on the prompt word, wherein the reply text contains the detection analysis text;
[0009] Displaying the reply text, and / or playing a reply voice corresponding to the reply text.
[0010] In a second embodiment, an online human-computer interaction method is also provided in the present specification. The method comprises:
[0011] Displaying a service interface for human-computer interaction;
[0012] in response to an audio-video interaction starting operation triggered through the service interface, starting an online man-machine audio-video interaction;
[0013] during the online man-machine audio-video interaction, collecting a video stream and a voice conversation stream on the user side; wherein the video stream contains multi-angle images of a target object, and the target object includes a body part of the user;
[0014] sending the video stream and the voice conversation stream to a server;
[0015] receiving the reply text and the reply voice corresponding to the reply text returned by the server; wherein the reply text is generated based on a prompt word, and is used to respond to the voice conversation stream; the prompt word is generated according to the video stream and the voice conversation stream, and is used to guide the generation of detection analysis text of the target object based on image description text information of the video stream, the detection analysis text contains the analysis result of the state of the body part of the user; and the reply text contains the detection analysis text;
[0016] playing the reply voice while displaying the reply text.
[0017] In a third embodiment, the present specification also provides an online man-machine interaction method. The method comprises:
[0018] displaying a health service interface, the health service interface having an audio-video interaction control;
[0019] in response to an operation on the audio-video interaction control, starting an online man-machine audio-video interaction;
[0020] during the online man-machine audio-video interaction, collecting a biological video stream and a voice conversation stream of the user; wherein the biological video stream is a continuous image sequence containing a body part of the user;
[0021] generating a prompt word based on the biological video stream and the voice conversation stream;
[0022] inputting the prompt word into a third preset model to execute the third preset model to output a reply text; the reply text is used to respond to a health problem consulted by the user, and the health problem is extracted from the voice conversation stream;
[0023] playing a reply voice corresponding to the reply text to the user.
[0024] In a fourth embodiment, the present specification provides a service system. The system comprises:
[0025] The client is configured to display a service interface for human-computer interaction, and start online human-computer audio-video interaction in response to an audio-video interaction starting operation triggered through the service interface. During the online human-computer audio-video interaction, a video stream and a voice conversation stream of a user side are collected, and the video stream and the voice conversation stream are sent to the server. The video stream contains multi-angle images of a target object, and the target object includes a body part of the user.
[0026] The server is configured to generate a prompt word based on the video stream and the voice conversation stream. The prompt word is used to guide generation of a detection analysis text of the target object based on image description text information of the video stream. The detection analysis text contains a body part state analysis result of the user. Based on the prompt word, a reply text for responding to the voice conversation stream is generated. The reply text contains the detection analysis text. The reply text is converted into a reply voice, and the reply text and the reply voice are sent to the client.
[0027] The client is further configured to play the reply voice and display the reply text.
[0028] In a fifth embodiment, an online human-computer interaction device is provided in the specification. The device includes:
[0029] A collection module is configured to collect a video stream and a voice conversation stream of a user side in online human-computer audio-video interaction. The video stream contains multi-angle images of a target object, and the target object includes a body part of the user.
[0030] A generation module is configured to generate a prompt word based on the video stream and the voice conversation stream. The prompt word is used to guide generation of a detection analysis text of the target object based on image description text information of the video stream. The detection analysis text contains a body part state analysis result of the user.
[0031] The generation module is further configured to generate a reply text for responding to the voice conversation stream based on the prompt word. The reply text contains the detection analysis text.
[0032] A playing and displaying module is configured to display the reply text and / or play a reply voice corresponding to the reply text.
[0033] In a sixth embodiment, an online human-computer interaction device is provided in the specification. The device includes:
[0034] A display module is configured to display a service interface for human-computer interaction.
[0035] A starting module is configured to start online human-computer audio-video interaction in response to an audio-video interaction starting operation triggered through the service interface.
[0036] a collection module, configured to collect a video stream and a voice conversation stream of a user side in the online man-machine audio-video interaction process, wherein the video stream comprises multi-angle images of a target object, and the target object comprises a body part of the user;
[0037] a sending module, configured to send the video stream and the voice conversation stream to a server;
[0038] a receiving module, configured to receive a reply text and a reply voice corresponding to the reply text returned by the server, wherein the reply text is generated based on a prompt word, and is used to respond to the voice conversation stream; the prompt word is generated according to the video stream and the voice conversation stream, and is used to guide generation of a detection analysis text of the target object based on image description text information of the video stream, the detection analysis text comprising a body part state analysis result of the user; and the reply text comprises the detection analysis text;
[0039] a playing and displaying module, configured to play the reply voice while displaying the reply text.
[0040] In a seventh embodiment, the present specification also provides an online man-machine interaction device. The device comprises:
[0041] a display module, configured to display a health service interface, wherein the health service interface has an audio-video interaction control;
[0042] a starting module, configured to start online man-machine audio-video interaction in response to an operation on the audio-video interaction control;
[0043] a collection module, configured to collect a biological video stream and a voice conversation stream of a user in the online man-machine audio-video interaction process, wherein the biological video stream is a continuous image sequence comprising a body part of the user;
[0044] a generation module, configured to generate a prompt word based on the biological video stream and the voice conversation stream;
[0045] an execution module, configured to input the prompt word into a third preset model, and execute the third preset model to output a reply text; the reply text is used to respond to a health problem consulted by the user, and the health problem is extracted from the voice conversation stream;
[0046] a playing module, configured to play a reply voice corresponding to the reply text to the user.
[0047] In an eighth embodiment, the present specification provides an electronic device, comprising a memory and a processor, wherein the memory stores executable program instructions, and the processor executes the program instructions to implement the method provided in the first to third embodiments.
[0048] In a ninth embodiment, the present specification provides a computer-readable storage medium, which stores a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the method provided in the first to third embodiments.
[0049] In a tenth embodiment, the present specification further provides a computer program product, comprising computer programs / instructions, which are executed by a processor to implement the method provided in the first to third embodiments.
[0050] The scheme provided in the above embodiments of the present specification generates a prompt word in online human-machine audio-video interaction according to the collected video stream and voice dialogue stream on the user side, and then generates a reply text for responding to the voice dialogue stream based on the prompt word and outputs the reply text to the user. It can be seen that the present scheme realizes the online human-machine audio-video interaction function, which enables the user to more naturally and smoothly interact and communicate with the intelligent dialogue system (also known as intelligent dialogue assistant) in multiple modalities, especially when complex image information needs to be uploaded or multi-view detection is required (such as hair state detection task), which can simplify user operation and improve user experience. The above online human-machine audio-video interaction is triggered and started through a service interface for human-machine interaction. Specifically, the service interface has an audio-video interaction control, and the user can trigger and start the online human-machine audio-video interaction by operating the audio-video interaction control. The above service interface may, for example, be a health service interface provided by an intelligent dialogue assistant for health services. Further, when the reply text is output to the user, specifically, the reply text can be displayed while playing the reply voice corresponding to the reply text, which can facilitate the user to understand the content in the reply voice. BRIEF DESCRIPTION OF DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in the present specification, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only a part of the embodiments disclosed in the present specification, and other drawings can also be obtained by those skilled in the art without creative labor. In the drawings:
[0052] Figure 1 The technical architecture diagram based on which the methods provided in the present specification are implemented is provided for an exemplary embodiment;
[0053] Figure 2A and Figure 2BA structural diagram of a service system (specifically, an online human-computer interaction system, i.e., an intelligent dialogue system) provided for an exemplary embodiment in the specification;
[0054] Figure 3 , Figure 4 and Figure 5 A flowchart of an online human-computer interaction method provided for an exemplary embodiment in the specification;
[0055] Figure 6 , Figure 7 and Figure 8 A structural diagram of an online human-computer interaction device provided for an exemplary embodiment in the specification;
[0056] Figure 9 A structural diagram of an electronic device provided for an exemplary embodiment in the specification. DETAILED DESCRIPTION
[0057] With the rapid development of artificial intelligence and mobile internet technology, intelligent dialogue assistants have gradually become an important tool for providing human-computer interaction in various applications. Intelligent dialogue assistants can interact with users in a dialogue, and their mission is to help users solve problems or chat, which has been widely used in various fields such as medical health. Taking a health intelligent dialogue assistant (such as a health manager provided on some applications) as an example, it is committed to helping users solve various problems before, during and after medical treatment, covering health consultation, disease screening, medical advice, medication reminders, rehabilitation guidance and other scenarios, which significantly improves the user's health management efficiency and life health level. Currently, health intelligent dialogue assistants or other intelligent dialogue assistants mainly support text dialogue. For example, users describe their health status by inputting text, and health intelligent dialogue assistants can use their language models to generate corresponding health advice or medical guidance. With the development of multi-modal technology, in recent years, more and more intelligent dialogue assistants have added the ability of image-text dialogue interaction, supporting users to upload images when asking questions, so as to combine image content to understand and generate replies, thereby improving the intuitiveness of interaction and the accuracy of service. This image-text dialogue interaction method allows users to upload images by clicking the "image upload" control on the interface, which is effective and easy to use when users need to upload a small number of images. However, when users need to upload complex image information, such as continuously uploading multiple images, videos or even real-time feedback interaction, the above image-text dialogue interaction method cannot meet the actual needs of users because the image-text dialogue interaction method: on the one hand, the image upload process is cumbersome, affecting the interaction efficiency; on the other hand, it mainly relies on static images uploaded by users (such as wound photos and test reports), lacks the ability to perceive real-time visual information of users, and cannot realize dynamic recognition and real-time feedback of user status.
[0058] To solve the above problems, the embodiments of the present specification provide a solution, which introduces audio and video interaction capabilities for intelligent dialogue assistants.
[0059] In order for those skilled in the art to better understand the technical solutions in the present specification, the technical solutions in the embodiments of the present specification will be clearly and completely described below with reference to the drawings in the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present specification, not all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those skilled in the art without creative work should belong to the scope of protection of the present specification.
[0060] It should be noted that, for the convenience of description, only the parts related to the technical solutions are shown in the drawings. The embodiments in the present specification and the features in the embodiments can be combined with each other without conflict. In addition, the terms "first", "second", "third" and the like in the embodiments of the present specification are only used for information differentiation, and do not have any limiting effect. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, product or equipment. Without more limitation, it does not exclude that there are other same or equivalent elements in the process, method, product or equipment including the elements. In addition, in the present specification, unless explicitly stated, "receiving and sending of data" is not necessarily direct receiving and sending, but can be indirect receiving and sending. For example, A receives the data sent by B, which can be understood as A directly receiving the data sent by B, or A indirectly receiving the data sent by B through C and other subjects; similarly, B sends data to A, which can be understood as B directly sending data to A, or B indirectly sending data to A through C and other subjects. Here, C can be one subject, or two or more subjects.
[0061] Also, it should be noted that certain terminology has been used in the specification for the purpose of providing a clear and concise description of the embodiments. As used herein, "an embodiment" or "one embodiment" or "an alternative" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of "an embodiment" or "in one embodiment" or "an alternative" or "in an alternative" are not necessarily referring to the same embodiment. Furthermore, the particular features, structures, or characteristics can be combined in any suitable manner on different embodiments or examples. Although the method steps of the embodiments are described in a particular, sequential order, it should be noted that some steps can be performed in an order other than the order described herein. Further, some steps can be performed concurrently, rather than sequentially. In addition, some steps can be performed by different components at the same time, e.g., by a processor and a controller. Furthermore, some steps can be performed by the same component or components at different times. For example, a step performed by a processor can be performed at a later time by the same processor or by a different processor. Similarly, any entities described herein can be components of the same component or components, or different components of a larger component. Further, it should be noted that the steps of the embodiments can be performed by hardware components or using instructions stored in computer-readable media. As an example, one or more of the steps of the embodiments can be performed by one or more of the processor 102, the memory 104, the communication interface 106, the display 108, the input device 110, the speaker 112, the microphone 114, the camera 116, the storage 118, the power supply 120, and / or the bus 122, individually or in any combination.
[0062] Further, it should be noted that the user data (e.g., the user's video stream, the user's voice conversation stream, etc.) obtained by the embodiments of the present disclosure is authorized by the user and does not involve the user's privacy.
[0063] The embodiments of the present disclosure are described below in conjunction with the accompanying drawings.
[0064] First, the words involved in the embodiments of the present disclosure are described. It can be understood that the description is for a clearer understanding of the embodiments of the present disclosure and does not necessarily constitute a limitation on the embodiments of the present disclosure.
[0065] Automatic Speech Recognition (ASR): is used to recognize human speech as text, i.e., to convert human speech signals into text, which involves multiple processes including speech signal acquisition, feature extraction, acoustic model, language model, and decoding, etc.
[0066] Voice Activity Detection (VAD): is a signal processing technique aimed at automatically identifying speech portions in an audio signal and distinguishing them from non-speech portions (e.g., noise or silence). This technique is very important for improving speech communication quality, reducing storage requirements, and optimizing the performance of speech recognition systems, etc., and it mainly focuses on "whether there is someone speaking now". That is, VAD can be used to determine whether the user is speaking.
[0067] Text-to-Speech (TTS): is used to convert text information into natural flow field voice output.
[0068] The preset model is a well-trained artificial intelligence model. In the specification, the preset model includes a first preset model, a second preset model, a third preset model, and the like. The first preset model is a multi-modal model, such as a vision-language model (VLM). The vision-language model (VLM) is a multi-modal model that integrates visual information (image or video) and language information (text). Thus, the vision-language model can recognize image content. The second preset model is a task model that is specifically used for processing a specific task (such as target detection). The third preset model is a text model, specifically, for example, a language model (LLM). The language model (LLM) is an artificial intelligence model that is a key component in natural language processing (NLP) technology. The LLM is based on the Transformer architecture and can understand and generate high-quality human language text. It is used in dialog interaction systems for dialog content generation. It is worth noting that the embodiments of the present specification do not limit the number of parameters supported by the preset model, and the goal is to meet the actual application requirements.
[0069] The technical solutions provided by the embodiments of the present specification are based on Figure 1 the technical architecture shown in the specification. As shown in Figure 1 , in this technical architecture, low latency, high accuracy online human-computer audio and video interaction is achieved through an asynchronous link. The execution process of the asynchronous link includes the following:
[0070] 1) continuously collect the video stream (user side picture) of the user side through an image acquisition device (such as a camera, etc.), and call the multi-modal model in the image processor to process the continuous multiple frames of images in the video stream at a set frequency (such as once every 1 second), thereby generating a description text of the image and storing it in the memory module (Memory). The image acquisition device is, for example, a camera or the like. The memory module is also called memory. The multi-modal model can be, for example, a vision-language model (VLM).
[0071] In addition to the multi-modal model, the image processor also includes some other models, such as a task model dedicated to performing a certain task. The reason for including a task model in the image processor is that since the multi-modal model is generally used for understanding and analyzing the entire image, it is difficult to extract features of local objects in the image for fine-grained analysis. If the task to be performed is biased towards a specific application, such as target detection or security, these tasks often require identification and analysis of specific objects such as a particular person or a particular item. In this case, some traditional task models are needed to extract features of the corresponding local objects in the image to perform the corresponding task. For example, in the skin health scenario, after collecting the user's hand video stream through the image acquisition device and collecting the user's voice dialogue stream asking about the skin condition, the multi-modal model can be used to identify and analyze multiple frames of images in the hand video stream to identify the presence of the arm and the potential erythema area on the arm. Further, the corresponding task model can be used to perform fine-grained analysis on the identified erythema area to extract specific features of the erythema (such as size, density, etc.). Here, the specific features extracted by the task model for the local objects in the image will be stored in the memory module in the form of structured parameters.
[0072] With the above, in the memory module, the image description text information of the continuous multiple frames of images in the video stream is saved in the form of text + structured parameters, where the text is the description text of the entire image (i.e., the global description text), and the structured parameters are the structured description text of the local objects in the image (i.e., the local description text), which includes but is not limited to the position, size, category, color, density, and bounding box of the local objects in the image. The local object can refer to a specific region or element in the image.
[0073] 2) Collect user voice input continuously through a sound pickup device (such as a microphone) and call the automatic speech recognition (ASR) function to convert the received user voice dialogue stream into text.
[0074] Among them, generally, automatic speech recognition (ASR) is to make end-of-sentence judgment on the user's input voice through the fixed silence duration configured therein (using a fixed VAD duration). That is: VAD is a way to detect whether the user's current voice input has ended based on silence. For example, when a certain duration of silence is detected, it can be considered that the user's current voice input has ended. Of course, other ways can also be used to detect whether the user's current voice input has ended, such as based on endpoint detection, based on predefined time threshold detection, etc. Among them, based on endpoint detection, not only the presence of silence is considered, but also factors such as speech rate and tone change are combined to comprehensively judge whether the user's current voice input has ended. Based on the predefined time threshold detection, the principle is: set a fixed time window, and after no new audio data transmission exceeds this time window, it will automatically judge that the user's current voice input has ended.
[0075] 3) When the user finishes speaking a sentence, such as detecting voice VAD, it will stop automatic speech recognition (ASR), and splice the text corresponding to the user's speech just now with the related image description text information to form a prompt word (Prompt) for the response generator. The response generator is constructed based on a text model. The text model is also commonly known as a text dialogue model, which can be a language model (LLM), for example.
[0076] 4) Input the prompt word (Prompt) to the response generator, and the response generator will use the text model therein to generate a conversation text (also known as a reply text) according to the prompt word.
[0077] 5) Convert the conversation text into corresponding conversation voice through the TTS function and play it to the user.
[0078] The technical architecture mentioned above is implemented based on the server and the client. For example, continue to refer to Figure 1As shown, the image acquisition device, sound pickup acquisition device, and voice playing function in the technical architecture are implemented on the client, and the image processor (including a multi-modal model and a task model), ASR, VAD, splicing operation, response generator (including a text model), and TTS function are implemented on the server. The server can be a server, a server cluster, a virtual server, or a cloud, etc. The client can be, but is not limited to, a smart phone, a smart wearable device, a tablet computer, a notebook computer, a desktop computer, etc. The server provides corresponding function services, such as intelligent dialogue, for the client. A user can initiate an intelligent video dialogue through a browser, an application (APP), a web application H5 (HyperText Markup Language 5, the fifth generation of HTML, HyperText Markup Language), a light application (also known as a mini-program, a lightweight application), or a cloud application on the client. After the video dialogue accesses the intelligent dialogue system on the server, an online human-computer audio-video dialogue interaction is established between the client and the server.
[0079] Thus, Figure 2A and Figure 2B It is also shown that an embodiment of the present specification provides an online human-computer interaction system (also referred to as a service system), which includes a client 100 and a server 200. Wherein,
[0080] The client 100 is configured to display a service interface for human-computer interaction; in response to an audio-video interaction starting operation triggered through the service interface, start an online human-computer audio-video interaction; during the online human-computer audio-video interaction, acquire a video stream and a voice dialogue stream on the user side; and send the video stream and the voice dialogue stream to the server. The video stream contains multi-angle images of a target object, and the target object includes a body part of a user.
[0081] The server 200 is configured to generate a prompt word based on the video stream and the voice dialogue stream; the prompt word is used to guide the generation of a detection analysis text of the target object based on image description text information of the video stream, and the detection analysis text contains a body part state analysis result of a user; generate a response text for responding to the voice dialogue stream based on the prompt word; the response text contains the detection analysis text; convert the response text into a response voice, and send the response text and the response voice to the client.
[0082] The client 100 is further configured to play the response voice and display the response text at the same time.
[0083] The online man-machine video interaction function described above can be integrated in an intelligent dialogue assistant in any field. The intelligent dialogue assistant is deployed on a server, and a user can enter a service interface for man-machine interaction provided by the intelligent dialogue assistant through a client.
[0084] The specific implementation of the functions of the server 200 and the client 100 will be described in detail in the method embodiments below, and will not be described in detail here.
[0085] The technical solutions provided in the specification will be described below in the form of method embodiments.
[0086] Figure 3 A flowchart of an online man-machine interaction method provided by an embodiment of the specification is shown. The execution subject of the method is the server in the system described above. Referring to FIG. 1, Figure 3 The online man-machine interaction method includes the following steps:
[0087] 102, obtaining a video stream and a voice dialogue stream on the user side in online man-machine audio-video interaction; wherein the video stream contains multi-angle images of a target object, and the target object includes a body part of a user;
[0088] 104, generating a prompt word according to the video stream and the voice dialogue stream; the prompt word is used to guide the generation of a detection analysis text of the target object based on image description text information of the video stream, and the detection analysis text contains a body part state analysis result of the user;
[0089] 106, generating a reply text for responding to the voice dialogue stream based on the prompt word, wherein the reply text contains the detection analysis text;
[0090] 108, displaying the reply text and / or playing a reply voice corresponding to the reply text.
[0091] In 102 above, the online man-machine audio-video interaction is triggered and started by a user through a service interface for man-machine interaction displayed on a client, wherein the triggering and starting mode can be but is not limited to operating a corresponding audio-video interaction control, voice triggering, input text triggering, etc. The service interface is provided by a corresponding intelligent dialogue assistant. Specifically, the service interface can be a traditional man-machine interaction interface supporting text dialogue or graphic-text dialogue. Therefore, in the specification, the intelligent dialogue assistant supports online man-machine audio-video interaction, and the intelligent dialogue assistant can be an intelligent dialogue system applied in any field, such as health services (including medical health services), beauty care, online education, remote office, etc.
[0092] Exemplarily, taking the intelligent health manager (an intelligent dialogue assistant for providing health services) as an example, a user enters a man-machine interactive service interface provided by the intelligent health manager through a client. Generally, the service interface supports a text dialogue or a picture-text dialogue by default. As shown in Figure 2A , the user can trigger to start online man-machine audio-video interaction by operating an "audio-video dialogue" control 11 on the service interface. Alternatively, as shown in Figure 2B , the user can also input an audio-video dialogue starting instruction (such as "please start audio-video dialogue") in an input box 12 provided in the service interface to trigger to start online man-machine audio-video interaction. Of course, in other examples, the audio-video dialogue starting instruction can also be input in other ways, such as voice input, for example, operating a "voice input" control 13 to input the voice of "please start audio-video dialogue".
[0093] The scheme integrates the online man-machine audio-video interaction function for the intelligent dialogue assistant, realizes online man-machine audio-video interaction, and enables the user to communicate with the intelligent dialogue assistant (i.e., the intelligent dialogue system) more naturally and smoothly, especially when complex image information needs to be uploaded or multi-view detection is required (such as hair state detection tasks), which can simplify user operation and improve user experience.
[0094] After starting the online man-machine audio-video interaction, the user and the intelligent health manager will perform online man-machine audio-video interaction in the form of video call, and during the audio-video interaction process, the client will collect the video stream and the voice dialogue stream of the user side through different collection devices and send them to the server for processing. For example, the video input of the user side is collected through a camera to form a video stream, and the voice input of the user is collected through a microphone to form a voice dialogue stream.
[0095] Based on the above content, the step 102 "obtaining the video stream and the voice dialogue stream of the user side in the online man-machine audio-video interaction" includes:
[0096] 1021. receiving the video stream and the voice dialogue stream of the user side sent by the client;
[0097] The video stream is the video data of the user side captured by the image collection device of the client. The image collection device can be but is not limited to a camera or other image capture device. The voice dialogue stream is collected by the sound pickup device of the client, and the sound pickup device can be but is not limited to a microphone.
[0098] In addition, the video stream contains multi-angle images of the target object. Multi-angle images refer to images of the target object captured from different perspectives (front, side, top view, etc.). The change of perspective can come from the change of the position of the image collection device (such as a camera) or from the rotation of the target object itself.
[0099] wherein, in different interaction fields, the target object is different.
[0100] For example, in the health service field, the target object can be a body part of the user, such as a face, a head, a hand, skin, hair, nails, etc.
[0101] For another example, in the remote office field, the target object can be, for example, an abnormal target device.
[0102] For yet another example, in the online education field, the target object can be, for example, a relevant test question; or, the target object can also be the user himself / herself and the user's body posture, behavior movement track (in some cases, can also include environmental elements or devices in the user's interaction, etc.). For example, if the video stream is a dynamic video stream collected from a student performing a scientific experiment (such as chemistry, biology, physics, medicine, etc.), the target object can be the student's body posture, key action part, operation behavior track, and experimental equipment interacting with him / her; and if the video stream is a dynamic video stream collected from a user performing dance practice, the target object can be the user himself / herself and the user's body posture and movement track (especially the key body parts and movement characteristics related to dance actions).
[0103] In addition, in the process of online human-computer audio / video interaction, in order to collect video data meeting the requirements, the user can be guided to shoot, and corresponding shooting feedback can be provided. In the process of guiding the user to shoot, the user can be provided with shooting guidance prompts adapted to different interaction task types.
[0104] Accordingly, the method provided by the embodiment can further include the following steps:
[0105] S12, determining a current interaction task type;
[0106] S14, determining a shooting guidance prompt information adapted to the interaction task type;
[0107] S16, outputting the shooting guidance prompt information to the user to guide the user to adjust a shooting pose.
[0108] The aforementioned collected video stream includes a video segment shot by guiding the user to adjust the shooting pose.
[0109] The interactive task type in S12 refers to the type of feature interaction task between the intelligent dialogue assistant and the user in the process of audio-video interaction between man and machine. The interactive task defines the purpose and content of the interaction. In different fields, the interactive task is different. For example, in the health service field, the interactive task can be a user health detection task, specifically, such as skin condition detection task, hair condition detection task, eye health detection task, etc. For another example, in the field of remote office, the interactive task can be but not limited to: device anomaly analysis task (such as hardware configuration, network connection, system crash, security vulnerability analysis task). For another example, in the field of online education, the interactive task can be but not limited to: test question answering, experiment operation error analysis, student dance action analysis, etc.
[0110] In the embodiment, the interactive task type can be determined according to user self-triggering selection or can be actively recommended by the intelligent dialogue assistant. The embodiment does not limit the determination method of the interactive task type.
[0111] For example, after the online man-machine audio-video interaction is started, the interactive task type of this time can be determined according to the voice input of the user. If the voice input of the user collected contains “I want to detect my hair loss condition”, the interactive task type of this time can be determined as “hair condition detection task”. And / or, in the interface of online man-machine audio-video interaction, a text input entrance (specifically, such as a text dialogue chat entrance) can be provided. The user can input the interactive content of this time through the text input entrance through the text input mode, such as “I want to detect my hair loss condition”. In this case, the interactive task type of this time can be determined as “hair condition detection task” according to the text input.
[0112] For another example, after the online man-machine audio-video interaction is started, a task window can be popped up, and a plurality of interactive task options can be displayed in the task window. The user can perform a point selection operation on any one of the plurality of interactive task options. The interactive task type of this time can be determined according to the interactive task option selected by the user. Or, a plurality of interactive task options can be displayed at a certain side position of the interface of online man-machine audio-video interaction for the user to select, so as to determine the interactive task type of this time according to the interactive task option selected by the user. For example, if the user selects the “start hair detection” option, the interactive task type of this time can be determined as “hair condition detection task”.
[0113] For example, in some scenarios, the user can not explicitly state the task, in which case the type of interaction task required can be automatically inferred from the user's relevant information. For example, according to the user's historical online human-computer interaction information (including historical online human-computer text and / or graphic-text conversation interaction, historical online human-computer audio and video interaction), it is determined that the user previously performed "skin detection" or "face detection", in which case it can be inferred that the user may need to perform hair detection next. For this inferred hair detection task, the user can also be confirmed, such as outputting the following voice "Hello, I am **health manager, do you want to consult hair problems", when receiving the user's confirmation voice (such as yes), the type of interaction task for this time is finally determined as "hair detection task".
[0114] Based on the above given example content, the S12 "determine the type of interaction task" can include any of the following ways:
[0115] Method one, in response to the user triggered task selection operation, the type of interaction task is determined according to the selected task item; wherein the task selection operation can be but not limited to realized through voice, input text, or point selection and the like. Or,
[0116] Method two, according to the user's relevant information, the type of interaction task recommended for the user is determined; wherein the relevant information includes the user's historical online human-computer interaction information.
[0117] In the above S14~S16, in order to be able to collect the user side video that meets the needs, so as to provide better interaction service (such as problem consultation service) for the user, the appropriate shooting guide prompt information can be determined according to the type of interaction task, so as to guide the user to shoot through the shooting guide prompt information, so as to collect the appropriate image frame.
[0118] Among them, the form of shooting guide prompt information can be but not limited to visual prompt and / or voice prompt, visual prompt such as animation demonstration, text prompt, schematic diagram prompt and the like at least one.
[0119] Exemplarily, taking the hair state detection task as an example, the determined shooting guide prompt information can be a simple tutorial animation, which explains from which angles the user needs to shoot the photo of his / her head (such as the top of the head, the forehead, the side and the back of the head), and can play the tutorial animation at a position (such as the lower right corner position) of the online man-machine audio-video interaction interface, for example, the guide shooting content shown by the tutorial animation includes: “Please align the top of your head with the center of the screen”, “Now, slowly rotate your head to make sure that the hairline on both sides can be shot”, etc. Under the guidance of this tutorial animation, the user can adjust the pose of his / her head or adjust the pose of the camera to adjust the shooting angle, so as to shoot his / her hair from multiple angles. In the process of the user shooting his / her hair from multiple angles, the camera will collect a continuous image data, which is the video clip shot under the guidance of the user shooting. In the case of guiding the user to shoot, the user will also be provided with corresponding timely voice feedback, such as image quality feedback (all clear, whether blocked, etc.), whether the current shooting angle meets the requirements, whether the shooting is completed, etc.
[0120] Therefore, the aforementioned video stream collected on the user side contains the video clip shot by guiding the user to adjust the shooting pose.
[0121] Here, in the online man-machine audio-video interaction, by guiding the user to shoot different angle images and providing instant voice feedback, the comprehensiveness and accuracy of image information collection can be ensured, thereby facilitating the accuracy of subsequent conversation text generation.
[0122] Further, for the video stream continuously collected by the image collection device, the image processing device shown in Figure 1 may be called to process the video stream to generate corresponding image description text information after a set time interval is reached.
[0123] Therefore, the method provided by the embodiment can further include the following steps:
[0124] S22, performing image recognition analysis on the acquired video stream according to a set time interval to generate image description text information corresponding to the video stream.
[0125] S24, storing the image description text information for subsequent generation of the prompt word.
[0126] Specifically, the image processing device shown in Figure 1 is called to perform image recognition analysis on the video stream continuously collected by the image collection device according to a set time interval, so as to generate image description text information corresponding to the video stream, and store the image description text information into the corresponding memory.
[0127] The time interval mentioned above can be flexibly set according to the actual situation, and this embodiment does not limit it. For example, the image processor can be called once every 1 second, that is, the image processor can be called once per second. In other words, it can be understood that the image acquisition device will call the image processor once every time a 1-second video stream is collected to perform image recognition and analysis on the 1-second video stream.
[0128] The image processor includes a first preset model and a second preset model. The second preset model is used to perform overall recognition analysis on the image, and the second preset model is used to perform local recognition analysis on the image.
[0129] Based on the above, in a specific feasible technical solution, step S22, "performing image recognition analysis on the acquired video stream and generating descriptive text information corresponding to the video stream," includes:
[0130] S222: Call the first preset model to perform global recognition on multiple consecutive frames of images in the video stream and generate global description text;
[0131] S224. Call the second preset model to perform local object recognition in the continuous multi-frame images and generate local description text.
[0132] Furthermore, the "storing the image description text information" in step S24 above includes:
[0133] S242. The global description text and the local description text corresponding to the consecutive multi-frame images are associated and stored for later use in the generation of the prompt words.
[0134] In S222 and S224 above, the first preset model and the second preset model can be respectively Figure 1 The multimodal model and task model shown herein, where a task model is used to perform a specific task. The generated global descriptive text refers to the textual information content that provides a general description of the image as a whole across multiple consecutive frames. The generated local descriptive text refers to the detailed local object description information generated by extracting the object features of local objects through local object recognition in multiple consecutive frames. The local object can be, but is not limited to, a specific local region or element (such as hair, erythema, etc.) in the image, and the corresponding local descriptive text includes at least one of the following: location, size, density, color, category, and morphology of the local object.
[0135] For example, the first preset model can be called once every 1 second (i.e., the first preset model is called at a frequency of 1 second 1 time) to perform global recognition on a plurality of continuous frames of images collected in the 1 second in the video stream collected by the image collection device, thereby generating a global description text of the plurality of continuous frames of images; further, a second preset model can be called to perform feature extraction on the local object recognized from the plurality of continuous frames of images, thereby generating a local description text of the local object, and the like.
[0136] It should be noted here that the above step S224 is not a necessary step, but an optional step. The step S224 can be triggered and executed when it is determined that local object recognition is needed according to the interaction task type of this interaction. That is, the method provided in the embodiment can further include the following step:
[0137] S223, when it is determined that local object recognition is needed according to the interaction task type, triggering and executing the above step S224.
[0138] For the determination of the interaction task type here, please refer to the related content described in the foregoing other embodiments, which will not be described in detail here.
[0139] For example, if the interaction task type is a "hair state detection task", it is determined that local object recognition is needed, and the local object to be recognized is the hair region in the image.
[0140] For example, if the interaction task type is a "red spot detection task on the skin", it is determined that local object recognition is needed, and the local object to be recognized is the red spot element in the image.
[0141] In the above S242, the global description text and the local description text of the plurality of continuous frames of images can be stored in association in the corresponding memory for subsequent support of prompt word generation. The memory can be a memory module as shown in the memory module. Figure 1 In the association storage, the local description text is saved in the form of a structured parameter. For example, the saved form of the local description text is: identification (such as name) of the local object: **; position: **; size: **; density: **.
[0142] For the user-side voice dialogue stream collected through a pickup device (such as a microphone), an automatic speech recognition (ASR) function is used to convert the voice signal in the voice dialogue stream into corresponding dialogue text, and after detecting the end of the voice dialogue stream (such as when a period of silence is detected, it can be determined that the current voice dialogue stream has ended, i.e., the current voice input of the user side has ended), the automatic speech recognition is stopped, and a prompt word generation stage is entered. In the prompt word generation stage, the image description text information that is more relevant to the semantic of the dialogue text is searched from the stored information, so as to splice the dialogue text with the searched image description text information to obtain the corresponding prompt word. The prompt word is input as an input parameter of the response generator (specifically, a third preset model in the response generator) shown in the middle of the figure, so that the response generator generates a dialogue text based on the input prompt word, for responding to the voice stream of the user side. Figure 1 The response generator (specifically, a third preset model in the response generator) shown in the middle of the figure, so that the response generator generates a dialogue text based on the input prompt word, for responding to the voice stream of the user side.
[0143] Based on the above content, in a specific implementable technical solution, the above-mentioned 104 "generating a prompt word according to the video stream and the voice dialogue stream" includes:
[0144] 1042, text conversion is performed on the voice dialogue stream to generate corresponding dialogue text;
[0145] 1044, the image description text information related to the dialogue text is searched from the stored information;
[0146] 1046, the prompt word is generated based on the dialogue text and the searched image description text information.
[0147] The prompt word generated here can include role setting of the third preset model, detection analysis task instruction, output constraint, etc., for guiding the third preset model to take the image description text information of the video stream as a context, to perform a corresponding detection analysis task on the target object and generate a matched detection analysis text, so as to respond to the user consultation question expressed in the dialogue text according to the obtained detection analysis text.
[0148] That is, the prompt word is used to guide the generation of the detection analysis text of the target object based on the image description text information of the video stream. The detection analysis text includes the body part state analysis result of the user, such as the detection analysis result of the state of hair / skin, etc.
[0149] And the above-mentioned 106 "generating a dialogue text for responding to the voice dialogue stream based on the prompt word" can include:
[0150] 1062, the prompt word is input into the third preset model, and the third preset model outputs the dialogue text.
[0151] In the above, after converting the voice dialogue flow into corresponding dialogue text, a retrieval augmentation (RAG) module can be used to retrieve image description text information related to the semantic of the dialogue text from memory, and then the dialogue text and the retrieved image description text information are spliced to generate a prompt word for inputting into a third preset model.
[0152] The third preset model is a text model, specifically, such as a language model (LLM). The dialogue text generated based on the prompt word by using the third preset model can be used to respond to the user's consultation question (such as hair loss condition) in the voice dialogue flow.
[0153] For example, in the field of medical health services, assuming that the generated prompt word is "You are an intelligent health assistant, please generate an objective hair health initial detection analysis result based on the image description text information of the video stream and the dialogue text of the user, and the output requirement is to use Chinese, and to divide into three parts of summary, observation, and suggestion, with a professional and gentle tone; wherein the image description text information of the video stream is "***", and the dialogue text of the user is "Please detect my hair loss condition". After inputting the prompt word into the third preset model, executing the third preset model, the output dialogue text can be, for example: "From the image, your forehead hairline has a mild backward trend; multi-angle view shows that the hair volume on the top of your head is sparse, and the scalp part is visible, which may be related to genetics or stress; it is recommended to maintain a good rest, and if necessary, consult a dermatologist for professional detection." In the above, the contents such as "the forehead hairline has a mild backward trend, the hair volume on the top of the head is sparse, and the scalp part is visible" contained in the dialogue text are the hair state detection analysis result of the user.
[0154] Based on the above description of the related processing content for the video stream and the voice dialogue flow, and in combination with the above description of the third preset model, Figure 1 It can be seen that the embodiment actually processes the video stream collected by the camera and the user voice input collected by the microphone in an asynchronous manner, which can improve the response speed of the dialogue and the user experience. In this asynchronous manner, the corresponding, Figure 1The image processor (specifically the multi-modal model therein) and the response generator (specifically the text model therein) in the foregoing are operated asynchronously. Here, the reason for not using a synchronous operation scheme such as the multi-modal model and the text model is that the synchronous operation direction has a longer delay, which makes the response time long, thereby causing the user to wait for a long time for the response. Moreover, the synchronous operation mode requires high synchronization of the image and the audio, and the effect is difficult to guarantee. Of course, there is also a scheme in which the multi-modal model is used to process the video stream, and the multi-modal model is also used to generate the response text, that is, only the multi-modal model is called to realize the processing of the video stream and the generation of the response text. However, this scheme has a slightly longer running delay than the asynchronous operation scheme, and the multi-modal model used for the response needs more alignment training data, which is relatively high in cost.
[0155] Further, the response text can be converted into corresponding response speech by the TTS function, and then sent to the user-side client for playing. Of course, the response text can also be sent to the client, so as to display the response text while playing the response speech, thereby facilitating the user to understand.
[0156] Therefore, a specific implementation scheme of the above-mentioned 108 “displaying the response text and / or playing the response speech corresponding to the response text” can include:
[0157] 1082, converting the response text into response speech;
[0158] 1084, sending the response speech and the response text to the user-side client, so as to play the response speech and / or display the response text by the client.
[0159] In the above-mentioned, considering that in some fields, such as the medical and health service field, there are often some special terms, if the response speech contains these special terms, simply playing the response speech may cause the user to not understand. Therefore, the corresponding response text is displayed synchronously while playing the response speech, which can facilitate the user to understand the response speech through text vision.
[0160] Overall, the embodiment upgrades the traditional multi-modal interaction of the intelligent dialogue assistant to an audio-video interaction mode, can support the user to perform natural and smooth multi-modal interaction, optimizes the user interaction experience, has strong practicality and interaction flexibility in actual application, can promote the leap development of the intelligent dialogue assistant, and improves the business value and market competitive advantage thereof.
[0161] The present specification also provides another online human-computer interaction method, and the execution subject of the method is the client in the foregoing system. Specifically, as shown inFigure 4 As shown in the figure, the online human-computer interaction method includes the following steps:
[0162] 202, display a service interface for human-computer interaction;
[0163] 204, in response to an audio-video interaction starting operation triggered through the service interface, start online human-computer audio-video interaction;
[0164] 206, in the process of online human-computer audio-video interaction, collect a video stream and a voice conversation stream on the user side; wherein the video stream contains multi-angle images of a target object, and the target object includes a body part of the user;
[0165] 208, send the video stream and the voice conversation stream to a server;
[0166] 210, receive the reply text returned by the server and the reply voice corresponding to the reply text; wherein the reply text is generated based on a prompt word, and is used to respond to the voice conversation stream; the prompt word is generated according to the video stream and the voice conversation stream, and is used to guide the generation of detection analysis text of the target object based on image description text information of the video stream, the detection analysis text containing a body part state analysis result of the user; and the reply text contains the detection analysis text;
[0167] 212, display the reply text while playing the reply voice.
[0168] The specific implementation of each step of the above embodiment can be referred to the related content in other embodiments, and will not be repeated here. In addition, the method provided in the embodiment can also include some steps disclosed in other embodiments, and can also be referred to the related content in other embodiments, and will not be repeated here.
[0169] The present specification also provides another online human-computer interaction method, part of the steps (such as 302, 304, 306) in the method are executed by the client, and the rest of the steps (such as 308, 3010, 3012) are executed by the server. In addition, the application scenario of the method is an intelligent dialogue assistant for health services (such as a health manager). Specifically, see Figure 5 As shown in the figure, the online human-computer interaction method includes the following steps:
[0170] 302, display a health service interface, the health service interface having an audio-video interaction control;
[0171] 304, in response to an operation on the audio-video interaction control, start online human-computer audio-video interaction;
[0172] 306、In the online human-computer audio-video interaction process, a biological video stream and a voice dialogue stream of the user are collected;
[0173] 308、Based on the biological video stream and the voice dialogue stream, a prompt word is generated;
[0174] 310、The prompt word is input into a third preset model, and a reply text is output by the third preset model; the reply text is used to respond to a health problem consulted by the user, and the health problem is extracted from the voice dialogue stream;
[0175] 312、The reply text is played to the user.
[0176] In the above, the health service interface is a service interface for human-computer interaction as shown in Figure 2A and Figure 2B , and the audio-video interaction control is an "audio-video dialogue" control 11 as shown in Figure 2A .
[0177] In addition, the biological video stream refers to a continuous image sequence of a user's body part (such as a face, a head, a hand, a neck, etc.) collected by an image collection device, and contains image information related to the user's biological features, such as hair state, skin condition, etc., to support health examination analysis.
[0178] The specific implementation of each step in the above embodiment can be referred to the related content in other embodiments, which will not be repeated here. In addition, the method provided in the embodiment can also include some steps disclosed in other embodiments, which can also be referred to the related content in other embodiments, which will not be repeated here.
[0179] The above is described in combination with Figure 3~5 the corresponding specific embodiments of the present specification. It should be noted that the described specific embodiments require that: other embodiments are within the scope of the claims thereof; and, in some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.
[0180] The device embodiments corresponding to the method embodiments provided in the present specification will be introduced below.
[0181] Figure 6 A structural schematic diagram of an online human-computer interaction device provided by an exemplary embodiment of the present specification is shown. As shown in Figure 6As shown, the device comprises: an acquisition module 42, a generation module 44, and a playing module 46. Among them,
[0182] The acquisition module 42 is configured to acquire a video stream and a voice dialogue stream on a user side in an online man-machine audio-video interaction. The video stream contains multi-angle images of a target object, and the target object includes a body part of a user.
[0183] The generation module 44 is configured to generate a prompt word according to the video stream and the voice dialogue stream. The prompt word is used to guide generation of a detection analysis text of the target object based on image description text information of the video stream. The detection analysis text contains a body part state analysis result of the user.
[0184] The generation module 44 is further configured to generate a reply text for replying to the voice dialogue stream based on the prompt word. The reply text contains the detection analysis text.
[0185] The playing and displaying module 44 is configured to display the reply text and / or play a reply voice corresponding to the reply text.
[0186] In an embodiment, the device further comprises an identification and analysis module and a storage module. The identification and analysis module is configured to perform image identification and analysis on the acquired video stream at a set time interval to generate image description text information corresponding to the video stream. The storage module is configured to store the image description text information for subsequent generation of the prompt word.
[0187] In an embodiment, when the identification and analysis module is used to perform image identification and analysis on the acquired video stream to generate image description text information corresponding to the video stream, it is specifically configured to: call a first preset model to perform global identification on continuous multiple frames of images in the video stream to generate global description text; and call a second preset model to perform local identification on local objects in the continuous multiple frames of images to generate local description text. When the storage module is used to store the image description text information, it can be specifically configured to: store the global description text and the local description text corresponding to the multiple frames of images in association.
[0188] In an embodiment, the device further comprises a determination module and a triggering module. The determination module is configured to determine a current interaction task type. The triggering module is configured to trigger the calling of the second preset model to perform local identification on local objects in the multiple frames of images to generate local description text when it is determined that local object identification is required according to the interaction task type. The local description text includes at least one of a position, a size, a density, a color, and a category of a local object.
[0189] In an implementation, the determining module is configured to determine the interactive task type in response to a task selection operation triggered by the user, and determine the interactive task type according to a selected task item; or determine the interactive task type recommended for the user according to relevant information of the user, wherein the relevant information includes historical online human-computer interaction information of the user.
[0190] In an implementation, the determining module is further configured to determine the shooting guidance prompt information adapted to the current interactive task type, and the output module is further configured to output the shooting guidance prompt information to the user to guide the user to adjust a shooting pose, wherein the video stream includes a video clip captured by guiding the user to adjust the shooting pose, and the shooting guidance prompt information includes visual guidance and / or voice guidance.
[0191] In an implementation, the generating module is configured to, when generating the prompt word according to the video stream and the voice dialogue stream, perform text conversion on the voice dialogue stream to generate corresponding dialogue text, retrieve the image description text information related to the dialogue text from the storage information, and generate the prompt word based on the retrieved image description text information and the dialogue text. The generating module is further configured to, when generating the reply text for responding to the voice dialogue stream based on the prompt word, input the prompt word into a third preset model to output the reply text.
[0192] In an implementation, the playing module is configured to, when displaying the reply text and / or playing the reply text to the user, convert the reply text into a reply voice, and send the reply voice and the reply text to a user-side client to play the reply voice and / or display the reply text by the client.
[0193] Figure 7 A structural schematic diagram of an online human-computer interaction device provided by another exemplary embodiment of the present specification is shown. As shown in FIG. 6, the online human-computer interaction device includes a server and a client. Figure 7As shown, the device comprises a display module 52, a starting module 54, a collection and sending module 56, a receiving module 58, a broadcast display module 510. Among them, the display module 52 is used to display a service interface for human-computer interaction. The starting module 54 is used to start the online human-computer audio-video interaction in response to the audio-video interaction starting operation triggered through the service interface. The collection and sending module 56 is used to collect the video stream and voice dialogue stream of the user side in the online human-computer audio-video interaction process; wherein the video stream contains multi-angle images of the target object, and the target object includes the body parts of the user; and further used to send the video stream and voice dialogue stream to the server. The receiving module 58 is used to receive the reply text returned by the server and the reply voice corresponding to the reply text; wherein the reply text is generated based on the prompt word, and is used to respond to the voice dialogue stream; the prompt word is generated according to the video stream and the voice dialogue stream, and is used to guide the generation of the detection analysis text of the target object based on the image description text information of the video stream, and the detection analysis text contains the body part state analysis result of the user; the reply text contains the detection analysis text. The broadcast display module 510 is used to play the reply voice while displaying the reply text.
[0194] Figure 8 The structure diagram of an online human-computer interaction device provided by another example embodiment of the present specification is shown. As shown in the figure, Figure 8 The device comprises a display module 62, a starting module 64, a collection module 66, a generation module 68, an execution module 610, and a playing module 612. Among them, the display module 62 is used to display a health service interface, and the health service interface has an audio-video interaction control. The starting module 64 is used to start the online human-computer audio-video interaction in response to the operation of the audio-video interaction control. The collection module 66 is used to collect the biological video stream and voice dialogue stream of the user in the online human-computer audio-video interaction process; wherein the biological video stream is a continuous image sequence containing the body parts of the user. The generation module 68 is used to generate a prompt word based on the biological video stream and the voice dialogue stream. The execution module 610 is used to input the prompt word into a third preset model to execute the third preset model to output a reply text; the reply text is used to respond to the health problem consulted by the user, and the health problem is extracted from the voice dialogue stream. The playing module 612 is used to play the reply voice corresponding to the reply text to the user.
[0195] It should be noted that the above-mentioned devices can implement the technical solutions described in the corresponding method embodiments. The specific implementation principles of each module or unit can be found in the relevant content of the corresponding method embodiments, and will not be elaborated further here. Furthermore, for ease of description, the above devices are described by function as various modules or units. Of course, when implementing one or more of this specification, the functions of each module or unit can be implemented in one or more software and / or hardware, or a module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0196] Furthermore, embodiments of this specification also provide an electronic device. For example... Figure 9 As shown, the electronic device 700 includes a memory 71 and a processor 72.
[0197] The aforementioned memory 71 can be implemented by at least one volatile or non-volatile storage device of any type, or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Furthermore, the memory, wholly or partially, can be integrated with the processor. The memory can contain both removable and non-removable components.
[0198] The processor 72 described above may include one or more general-purpose processors and / or special-purpose processors.
[0199] Furthermore, memory 71 may contain a non-transitory computer-readable medium storing executable program instructions 712 (e.g., compiled or uncompiled program logic and / or machine code). Processor 72 is capable of executing the program instructions 712 stored in memory to implement any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Additionally, execution of program instructions 712 by processor 72 may result in processor using corresponding data 711.
[0200] For example, the program instructions 712 described above may include an operating system 7122 (e.g., an operating system kernel, device drivers, and / or other modules) and one or more applications 7121 (e.g., a browser, social media application, or game application) installed on the electronic device 700. Similarly, the data 711 described above may include operating system data 7112 and application data 7111. The operating system data 7112 is primarily accessible to the operating system 7122, while the application data 7111 is primarily accessible to one or more applications 7121. The application data 7111 may reside in a file system that is visible or hidden from the user of the electronic device 400.
[0201] Application 7121 can communicate with operating system 7122 through one or more application programming interfaces (APIs). These APIs facilitate application 7122 in reading and / or writing application data, transmitting or receiving information via communication components, and receiving or displaying information on the user interface. In some terms, application 7121 may be simply referred to as "app". Furthermore, application 7121 can be downloaded to the electronic device through one or more online application stores or app markets. However, application 7121 can also be installed on electronic device 400 in other ways, such as through a web browser or a physical interface on electronic device 700 (e.g., a USB port).
[0202] Furthermore, such as Figure 9 As shown, the electronic device also includes: a communication component 73, a display 74, a power supply component 75, an audio component 76, a user interface 77, and other components. Figure 9 The diagram only shows some components and does not imply that the electronic device 700 includes only these components. Figure 9 The components shown. Additionally... Figure 9 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the electronic device 700. The electronic device 700 in this embodiment can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT device; it can also be a server-side device such as a conventional server, cloud server, or server array; or it can be an integrated device combining terminal and server-side devices. If the electronic device 700 in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 9 The components within the dashed box; if the electronic device 700 in this embodiment is implemented as a conventional server, cloud server, or server array, etc., then it may not include... Figure 9 The component within the dashed box.
[0203] The communication component 73 is configured to facilitate wired or wireless communication between the device on which the communication component 73 is located and other devices. The device on which the communication component 73 is located can access a wireless network based on a communication standard, such as a 2G, 3G, 4G / LTE, 5G, or the like, or a combination thereof. In an example embodiment, the communication component 73 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an implementation, the communication component 73 includes a communication interface that enables the electronic device 700 to communicate with other electronic devices, access networks, and transmission networks through analog or digital modulation. For example, the communication interface can include a chipset and antenna for wirelessly communicating with a radio access network or access point. In addition, the communication interface can be a wired interface, such as an Ethernet, token ring, or USB port, or a wireless interface, such as a Wifi, Bluetooth, Global Positioning System (GPS), or wide area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface can support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface can also include multiple physical communication interfaces, such as a Wifi interface, a Bluetooth interface, and a wide area wireless interface.
[0204] The display 74 includes a screen, which can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touch or a slide action, but also detect duration and pressure related to the touch or slide operation.
[0205] The power supply component 75 provides power to the various components of the device on which the power supply component 75 is located. The power supply component 75 can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device on which the power supply component 75 is located.
[0206] The audio component 76 can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) that is configured to receive an external audio signal when the device on which the audio component is located is in an operational mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in memory or transmitted via the communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0207] The user interface 77 described above includes receiving user input and providing output to the user. Thus, the user interface 77 can include input components such as a keypad, keyboard, touch- sensitive or presence-sensitive panel, a computer mouse, a trackball, a joystick, a microphone, a still camera, and a video camera, among others, and output components such as a display screen (which can be combined with a touch-sensitive panel), a CRT, an LCD, an LED, a display using DLP technology, a printer, other known or future developed equivalent devices, among others. The user interface 77 can also generate audible output through a loudspeaker, a loudspeaker jack, an audio output port, an audio output device, a headphone, and other known or future developed equivalent devices. In some embodiments, the user interface 77 can include software, circuitry, or other forms of logic that enables the transmission of data to and from external user input / output devices. Additionally or alternatively, the electronic device 700 can support remote access from other devices through a communication interface or another physical interface (not shown). The user interface 77 can be configured to receive user input, the location and movement of which can be indicated by a pointer or cursor described herein. The user interface 77 can also be configured as a display device for rendering or displaying a text segment.
[0208] Accordingly, the embodiments of the present specification also provide a computer readable storage medium storing a computer program, when the computer program is executed by a processor, the processor is enabled to implement each step in the above-mentioned method embodiments. Wherein, the computer readable storage medium includes volatile or non-volatile or their combination, and can be removable or non-removable. Examples of the computer readable storage medium include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic cassette, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store data
[0209] In addition, the embodiments of the present specification also provide a computer readable storage medium, which stores a computer program, wherein when the computer program is executed in a computer, the computer is enabled to execute the method described above. Figure 3 to Figure 5 The method described above.
[0210] The embodiments of the present disclosure further provide a computer program product comprising computer programs / instructions which, when executed by a processor, implement the method described above. Figure 3 to Figure 5 The method described above.
[0211] Those skilled in the art should be aware that, in one or more examples described above, the functions described in the embodiments disclosed in the present specification can be implemented in hardware, software, firmware or any combination thereof. When implemented in software, the functions can be stored in a computer readable medium or transmitted as one or more instructions or codes on a computer readable medium.
[0212] The above detailed description of the specific implementation is further detailed for the purpose of the embodiments disclosed in the present specification, technical solutions and beneficial effects, and it should be understood that the above detailed description is only for the specific implementation of the embodiments disclosed in the present specification, and is not used to limit the protection scope of the embodiments disclosed in the present specification. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the embodiments disclosed in the present specification shall be included in the protection scope of the embodiments disclosed in the present specification.
Claims
1. An online human-computer interaction method, characterized in that, include: Acquire video streams and voice dialogue streams from the user side during online human-computer audio and video interaction; wherein, the video stream contains multi-angle images of a target object, and the target object includes parts of the user's body; Based on the video stream and the voice dialogue stream, prompt words are generated; the prompt words are used to guide the generation of detection and analysis text for the target object based on the image description text information of the video stream, and the detection and analysis text includes the analysis results of the user's body part status. Based on the prompt words, a response text is generated to respond to the voice dialogue stream, and the response text includes the detection and analysis text; Display the conversation text and / or play the conversation audio corresponding to the conversation text; The online human-computer voice and video interaction is initiated through an audio and video interaction entry provided by the intelligent dialogue assistant. The audio and video interaction entry includes audio and video interaction controls displayed on the service interface for human-computer interaction provided by the intelligent dialogue assistant. After the online human-computer audio and video interaction is initiated, the user interacts with the intelligent dialogue assistant in the form of a video call. The image description text information of the video stream includes: local description text; the local description text is generated when it is determined that local object recognition is required based on the current interaction task type, triggering local object recognition in multiple consecutive images in the video stream; the current interaction task type is determined by: responding to the user's selection operation of multiple interaction task options displayed on the online human-computer audio-visual interaction interface, and determining the current interaction task type based on the interaction task option selected by the user.
2. The method according to claim 1, characterized in that, Also includes: According to a set time interval, the acquired video stream is subjected to image recognition analysis to generate image description text information corresponding to the video stream; The image description text information is stored for later use in generating the prompt words.
3. The method according to claim 2, characterized in that, The acquired video stream is subjected to image recognition analysis to generate image description text information corresponding to the video stream, including: The first preset model is invoked to perform global recognition on multiple consecutive frames of images in the video stream, and global descriptive text is generated. The second preset model is invoked to perform local object recognition in the continuous multi-frame images and generate local descriptive text. And, storing the image description text information, including: The global description text and the local description text corresponding to the multi-frame images are associated and stored.
4. The method according to claim 3, characterized in that, Also includes: Determine the current interaction task type; When it is determined that local object recognition is required based on the interaction task type, the second preset model is triggered to perform local object recognition in the multi-frame images and generate local description text. The local description text includes at least one of the following: the location, size, density, color, and category of the local object.
5. The method according to claim 4, characterized in that, Determine the type of interactive task, including: In response to a user-triggered task selection action, determine the type of the interactive task based on the selected task item; or... Based on the user's relevant information, the type of interactive task recommended to the user is determined; wherein, the relevant information includes the user's historical online human-computer interaction information.
6. The method according to any one of claims 1 to 5, characterized in that, Also includes: Determine the appropriate shooting guidance prompts based on the current interaction task type; The shooting guidance prompts are displayed to the user to guide them in adjusting their shooting posture. The video stream includes video clips captured by guiding the user to adjust their shooting posture; The shooting guidance information includes visual prompts and / or voice prompts.
7. The method according to any one of claims 2 to 5, characterized in that, Generate prompt words based on the video stream and the voice dialogue stream, including: The voice dialogue stream is converted into text to generate corresponding dialogue text; Retrieve the image description text information related to the dialogue text from the stored information; Based on the retrieved image description text information and the dialogue text, the prompt words are generated; And, based on the prompt words, generate response text for responding to the voice dialogue stream, including: The prompt word is input into the third preset model, and the third preset model is executed to output the response text.
8. The method according to any one of claims 1 to 5, characterized in that, Displaying the conversation text and / or playing the corresponding conversation audio, including: Convert the echo text into echo speech; The echo voice and the echo text are sent to the user-side client so that the client can play the echo voice and / or display the echo text.
9. An online human-computer interaction method, characterized in that, include: Displays the service interface provided by the intelligent dialogue assistant for human-computer interaction; The service interface displays an audio and video interaction entry point, which includes audio and video interaction controls. In response to the launch operation initiated through the audio and video interaction portal, online human-computer audio and video interaction is started; wherein, after the online human-computer audio and video interaction is started, the user interacts with the intelligent dialogue assistant in the form of a video call; During the online human-computer audio and video interaction process, video streams and voice dialogue streams are collected from the user side; wherein, the video stream contains multi-angle images of the target object, and the target object includes parts of the user's body; The video stream and voice dialogue stream are sent to the server. The system receives a response text and a corresponding response voice from the server; wherein the response text is generated based on prompt words to respond to the voice dialogue stream; the prompt words are generated based on the video stream and the voice dialogue stream to guide the generation of detection and analysis text for the target object based on the image description text information of the video stream, and the detection and analysis text includes the user's body part state analysis results; the response text includes the detection and analysis text. While playing the recorded audio, the recorded text is displayed. The image description text information of the video stream includes: local description text; the local description text is generated when it is determined that local object recognition is required based on the current interaction task type, triggering local object recognition in multiple consecutive images in the video stream; the current interaction task type is determined by: responding to the user's selection operation of multiple interaction tasks displayed on the online human-computer audio-visual interaction interface, and determining the current interaction task type based on the interaction task selected by the user.
10. An online human-computer interaction method, characterized in that, include: The health service interface provided by the intelligent dialogue assistant is displayed, and the health service interface has audio and video interaction controls. In response to the operation of the audio and video interaction control, online human-computer audio and video interaction is initiated; wherein, after the online human-computer audio and video interaction is initiated, the user interacts with the intelligent dialogue assistant in the form of a video call; During the online human-computer audio and video interaction process, the user's biological video stream and voice dialogue stream are collected; wherein, the biological video stream is a continuous image sequence containing the user's body parts; Based on the biological video stream and the voice dialogue stream, prompt words are generated; the prompt words are used to guide the generation of detection and analysis text for the user's body parts based on the image description text information of the biological video stream. The prompt words are input into a third preset model, and the third preset model is executed to output a response text containing the detection and analysis text; the response text is used to respond to the user's inquiry about health issues, and the health issues are extracted from the voice dialogue stream; Play the corresponding audio message to the user; The image description text information of the biological video stream includes: local description text; the local description text is generated when it is determined that local object recognition is required based on the current interaction task type, triggering local object recognition in multiple consecutive images in the video stream; the current interaction task type is determined by: responding to the user's selection operation of multiple interaction tasks displayed on the online human-computer audio-visual interaction interface, and determining the current interaction task type based on the interaction task selected by the user.
11. A service system, characterized in that, include: The client is used to display the service interface provided by the intelligent dialogue assistant for human-computer interaction. The service interface displays an audio-visual interaction entry point, which includes audio-visual interaction controls. In response to a startup operation initiated through the audio-visual interaction entry point, online human-computer audio-visual interaction is launched. After the online human-computer audio-visual interaction is launched, the user interacts with the intelligent dialogue assistant via video call. During the online human-computer audio-visual interaction, video streams and voice dialogue streams are collected from the user's side. The video streams and voice dialogue streams are sent to the server. The video streams contain multi-angle images of a target object, including parts of the user's body. The server is configured to generate prompts based on the video stream and the voice dialogue stream; the prompts guide the generation of detection and analysis text for the target object based on the image description text information of the video stream, the detection and analysis text including the user's body part state analysis results; based on the prompts, generate response text for responding to the voice dialogue stream, the response text including the detection and analysis text; convert the response text into response speech, and send the response text and response speech to the client; wherein, the image description text information of the video stream includes: local description text; the local description text is generated when it is determined that local object recognition is required based on the current interaction task type, triggering local object recognition in multiple consecutive images in the video stream; the current interaction task type is determined by responding to the user's selection of multiple interaction tasks displayed on the online human-computer audio-visual interaction interface, and determining the current interaction task type based on the interaction task selected by the user; The client is also used to play the conversation audio and display the conversation text.
12. An online human-computer interaction device, characterized in that, include: The acquisition module is used to acquire video streams and voice dialogue streams from the user side during online human-computer audio and video interaction; wherein, the video stream contains multi-angle images of a target object, and the target object includes parts of the user's body; The generation module is used to generate prompt words based on the video stream and the voice dialogue stream; generate response text for responding to the voice dialogue stream based on the prompt words; the prompt words are used to guide the generation of detection and analysis text for the target object based on the image description text information of the video stream, and the detection and analysis text includes the user's body part state analysis results; A playback display module is used to display the echo text and / or play the echo voice corresponding to the echo text; The online human-computer voice and video interaction is initiated through an audio and video interaction entry provided by the intelligent dialogue assistant. The audio and video interaction entry includes audio and video interaction controls displayed on the service interface for human-computer interaction provided by the intelligent dialogue assistant. After the online human-computer audio and video interaction is initiated, the user interacts with the intelligent dialogue assistant in the form of a video call. The image description text information of the video stream includes: local description text; the local description text is generated when it is determined that local object recognition is required based on the current interaction task type, triggering local object recognition in multiple consecutive images in the video stream; the current interaction task type is determined by: responding to the user's selection operation of multiple interaction tasks displayed on the online human-computer audio-visual interaction interface, and determining the current interaction task type based on the interaction task selected by the user.
13. An online human-computer interaction device, characterized in that, include: The display module is used to display the service interface provided by the intelligent dialogue assistant for human-computer interaction. The service interface displays an audio and video interaction entry point, which includes audio and video interaction controls. The startup module is used to respond to the startup operation initiated through the audio and video interaction portal and start the online human-computer audio and video interaction; wherein, after the online human-computer audio and video interaction is started, the user interacts with the intelligent dialogue assistant in the form of a video call; The acquisition module is used to acquire video streams and voice dialogue streams from the user side during the online human-computer audio and video interaction process; wherein, the video stream contains multi-angle images of a target object, and the target object includes parts of the user's body; The sending module is used to send the video stream and the voice dialogue stream to the server. A receiving module is configured to receive a response text and a corresponding response voice returned by the server; wherein the response text is generated based on prompt words to respond to the voice dialogue stream; the prompt words are generated based on the video stream and the voice dialogue stream to guide the generation of detection and analysis text of the target object based on the image description text information of the video stream, and the detection and analysis text includes the user's body part state analysis results; the response text includes the detection and analysis text. The playback display module is used to display the playback text while playing the voice message; The image description text information of the video stream includes: local description text; the local description text is generated when it is determined that local object recognition is required based on the current interaction task type, triggering local object recognition in multiple consecutive images in the video stream; the current interaction task type is determined by: responding to the user's selection operation of multiple interaction tasks displayed on the online human-computer audio-visual interaction interface, and determining the current interaction task type based on the interaction task selected by the user.
14. An online human-computer interaction device, characterized in that, include: The display module is used to display the health service interface provided by the intelligent dialogue assistant, and the health service interface has audio and video interaction controls. The startup module is used to respond to the operation of the audio and video interaction control and start the online human-computer audio and video interaction; wherein, after the online human-computer audio and video interaction is started, the user interacts with the intelligent dialogue assistant in the form of a video call; The acquisition module is used to acquire the user's biological video stream and voice dialogue stream during the online human-computer audio and video interaction process; wherein, the biological video stream is a continuous image sequence containing the user's body parts; The generation module is used to generate prompt words based on the biological video stream and the voice dialogue stream; the prompt words are used to guide the generation of detection and analysis text for the user's body parts based on the image description text information of the biological video stream. An execution module is used to input the prompt words into a third preset model, execute the third preset model to output a response text containing the detection and analysis text; the response text is used to respond to the user's inquiry about health issues, which are extracted from the voice dialogue stream; The playback module is used to play the corresponding voice message to the user; The image description text information of the biological video stream includes: local description text; the local description text is generated when it is determined that local object recognition is required based on the current interaction task type, triggering local object recognition in multiple consecutive images in the video stream; the current interaction task type is determined by: responding to the user's selection operation of multiple interaction tasks displayed on the online human-computer audio-visual interaction interface, and determining the current interaction task type based on the interaction task selected by the user.
15. An electronic device, characterized in that, The method includes a memory and a processor, wherein the memory stores executable program instructions, and the processor executes the program instructions to implement the method of any one of claims 1 to 10.
16. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed in a computer, causes the computer to perform the method described in any one of claims 1 to 10.
17. A computer program product, characterized in that, The computer program product includes a computer program or instructions that, when executed by a processor, implement the method described in any one of claims 1 to 10.
Citation Information
Patent Citations
Knowledge question and answer method, device and equipment and storage medium
CN116561276A
Man-machine conversation method, device and equipment and computer readable storage medium
CN119312817A
Smart health care apparatus, system and method using artificial intelligence
KR102066225B1