Video interactive processing method and device
By capturing user images and detecting interaction mode switching, recognizing user voice data and performing multi-dimensional mapping to generate diagnostic data, this method solves the problems of convenience and accuracy in the interaction between users and medical service providers and the generation of diagnostic data in online services, thus meeting users' diverse online consultation needs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
- Filing Date
- 2026-01-16
- Publication Date
- 2026-05-26
AI Technical Summary
Online service providers face challenges such as diversified user needs and stringent service requirements, leading to increased pressure. Existing technologies struggle to effectively handle video interactions and the generation of diagnostic and treatment data between users and healthcare providers.
By capturing user images, detecting interaction mode switching, recognizing user voice data, and performing medical entity recognition, diagnostic data is generated using multi-dimensional mapping parameters. Combined with user interaction videos, multi-dimensional interactive mapping is performed to realize the display and return of diagnostic data.
It has improved the convenience of interaction between users and medical service providers and the comprehensiveness of diagnosis and treatment data, enhanced the accuracy of diagnosis and treatment data generation and the comprehensiveness of multi-dimensional mapping parameters, and met the diverse online consultation needs of users.
Smart Images

Figure CN122093371A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of data processing technology, and in particular to a video interactive processing method and apparatus. Background Technology
[0002] With the continuous promotion of internet technology and artificial intelligence, the development of online services based on the internet is also becoming increasingly rapid. Online services provide users with convenient means of service, enabling users to interact through online services to meet diverse service needs. This has led to more and more users joining the ranks of online service users. In this process, as the number of online service providers continues to increase and users' requirements for online services become more stringent, it has also brought certain pressures and challenges to online service providers. Summary of the Invention
[0003] This specification provides one or more embodiments of a video interaction processing method, comprising: capturing images of a user interacting with an intelligent agent instance corresponding to a medical service provider via voice, and detecting interaction mode switching based on the captured action image sequence. If the switching detection passes, uploading the user's voice data to a server, and obtaining candidate medical entities obtained by medical entity recognition based on the user's text returned by the server. Performing multi-dimensional interaction mapping on the candidate medical entities based on the user interaction video to obtain multi-dimensional mapping parameters of the candidate medical entities. Uploading the multi-dimensional mapping parameters to the server, and displaying the diagnosis and treatment data generated by the intelligent agent instance based on the multi-dimensional mapping parameters and the user's text returned by the server.
[0004] This specification provides one or more embodiments of another video interaction processing method, including: obtaining candidate medical entities by performing medical entity recognition on user text based on user voice data uploaded by a user terminal. The user voice data is obtained by capturing images of a user interacting with an intelligent agent instance corresponding to a medical service provider, and uploading the data after passing interaction mode switching detection based on the captured action image sequence. The candidate medical entities are returned to the user terminal to obtain multi-dimensional mapping parameters of the candidate medical entities based on the user interaction video through multi-dimensional interaction mapping. The intelligent agent instance is invoked to generate diagnostic data based on the multi-dimensional mapping parameters uploaded by the user terminal and the user text, and the diagnostic data is returned to the user terminal.
[0005] This specification provides one or more embodiments of a video interaction processing apparatus, comprising: a switching detection module configured to acquire images of a user interacting with an intelligent agent instance corresponding to a medical service provider via voice, and to detect a switching of interaction modes based on the acquired action image sequence; an entity acquisition module configured to, if the switching detection passes, upload user voice data to a server, and acquire candidate medical entities obtained by medical entity recognition based on user text returned by the server based on the user voice data; an interaction mapping module configured to perform multi-dimensional interaction mapping on the candidate medical entities based on the user interaction video, and obtain multi-dimensional mapping parameters of the candidate medical entities; and a parameter uploading module configured to upload the multi-dimensional mapping parameters to the server, and to display the diagnostic data generated by the intelligent agent instance based on the multi-dimensional mapping parameters and the user text returned by the server.
[0006] This specification provides one or more embodiments of another video interaction processing apparatus, including: an entity recognition module configured to perform medical entity recognition on user text based on user voice data uploaded by a user terminal to obtain candidate medical entities. The user voice data is obtained by capturing images of a user interacting with an intelligent agent instance corresponding to a medical service provider, and uploading the data after passing interaction mode switching detection based on the captured action image sequence. An entity return module is configured to return the candidate medical entities to the user terminal, obtaining multi-dimensional mapping parameters of the candidate medical entities by performing multi-dimensional interaction mapping on the candidate medical entities based on the user interaction video. A data return module is configured to call the intelligent agent instance to generate diagnostic data based on the multi-dimensional mapping parameters uploaded by the user terminal and the user text, and return the diagnostic data to the user terminal.
[0007] This specification provides one or more embodiments of a video interaction processing device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: acquire images of a user interacting with an agent instance corresponding to a medical service provider via voice, and perform interaction mode switching detection based on the acquired action image sequence. If the switching detection passes, the processor uploads the user's voice data to a server and obtains candidate medical entities obtained by medical entity recognition based on the user's text returned by the server. The processor performs multi-dimensional interaction mapping on the candidate medical entities based on the user interaction video to obtain multi-dimensional mapping parameters for the candidate medical entities. The processor uploads the multi-dimensional mapping parameters to the server and displays the diagnostic data generated by the agent instance based on the multi-dimensional mapping parameters and the user's text, returned by the server.
[0008] This specification provides one or more embodiments of another video interaction processing device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: perform medical entity recognition based on user text from user voice data uploaded by a user terminal to obtain candidate medical entities. The user voice data is obtained by capturing images of a user interacting with an intelligent agent instance corresponding to a medical service provider, and uploading the data after passing interaction mode switching detection based on the captured action image sequence. The candidate medical entities are returned to the user terminal to perform multi-dimensional interaction mapping on the candidate medical entities based on the user interaction video to obtain multi-dimensional mapping parameters of the candidate medical entities. The intelligent agent instance is invoked to generate diagnostic data based on the multi-dimensional mapping parameters uploaded by the user terminal and the user text, and the diagnostic data is returned to the user terminal.
[0009] This specification provides one or more embodiments of a computer-readable storage medium for storing computer-executable instructions, which, when executed, perform the following steps: image acquisition of a user interacting with an intelligent agent instance corresponding to a medical service provider via voice; detection of interaction mode switching based on the acquired action image sequence; if the switching detection passes, uploading the user's voice data to a server; and obtaining candidate medical entities obtained by medical entity recognition based on the user's text returned by the server, using the user's voice data. Multi-dimensional interaction mapping of the candidate medical entities based on the user's interaction video to obtain multi-dimensional mapping parameters for the candidate medical entities. Uploading the multi-dimensional mapping parameters to the server; and displaying the diagnostic data generated by the intelligent agent instance based on the multi-dimensional mapping parameters and the user's text, returned by the server.
[0010] This specification provides one or more embodiments of another computer-readable storage medium for storing computer-executable instructions, which, when executed, perform the following steps: Recognizing candidate medical entities based on user text derived from user voice data uploaded by a user terminal. The user voice data is obtained by capturing images of a user interacting with an intelligent agent instance corresponding to a medical service provider, and uploading the data after passing interaction mode switching detection based on the captured action image sequence. The candidate medical entities are returned to the user terminal to obtain multi-dimensional mapping parameters of the candidate medical entities based on multi-dimensional interaction mapping of the candidate medical entities using user interaction video. The intelligent agent instance is invoked to generate diagnostic data based on the multi-dimensional mapping parameters uploaded by the user terminal and the user text, and the diagnostic data is returned to the user terminal. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 A schematic diagram illustrating the implementation environment of a video interaction processing method provided in one or more embodiments of this specification; Figure 2 A flowchart illustrating a video interaction processing method provided in one or more embodiments of this specification; Figure 3 A schematic diagram of a voice interaction page provided for one or more embodiments of this specification; Figure 4 A schematic diagram of a video interactive page provided for one or more embodiments of this specification; Figure 5 This is a schematic diagram of an interactive video frame after a location point connection is provided for one or more embodiments of this specification; Figure 6 A schematic diagram of a medical label page provided for one or more embodiments of this specification; Figure 7 A schematic diagram of a consultation report page provided for one or more embodiments of this specification; Figure 8 A timing diagram of a video interaction processing method applied to a video interaction scenario, provided by one or more embodiments of this specification; Figure 9 A flowchart illustrating another video interaction processing method provided in one or more embodiments of this specification; Figure 10 A schematic diagram of an embodiment of a video interaction processing device provided in one or more embodiments of this specification; Figure 11 A schematic diagram of another embodiment of a video interaction processing apparatus provided in one or more embodiments of this specification; Figure 12 A schematic diagram of the structure of a video interactive processing device provided for one or more embodiments of this specification; Figure 13 This is a schematic diagram of the structure of another video interactive processing device provided in one or more embodiments of this specification. Detailed Implementation
[0012] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.
[0013] The video interaction processing method provided in one or more embodiments of this specification is applicable to the implementation environment of video interaction. (Refer to...) Figure 1 The implementation environment includes at least: User terminal 101, server 102; in addition, the implementation environment may also include agent instances of the interactive agent corresponding to the medical service provider; The user terminal 101 is used to collect images of users interacting with the intelligent agent instance corresponding to the medical service provider via voice, and to detect the switching of interaction modes based on the collected action image sequence. If the switching detection is successful, it performs multi-dimensional interaction mapping on the candidate medical entities obtained by medical entity recognition based on the user's voice data returned by the server based on the user's interaction video, obtains the multi-dimensional mapping parameters of the candidate medical entities, uploads the multi-dimensional mapping parameters to the server, and displays the diagnosis and treatment data generated based on the multi-dimensional mapping parameters and user text returned by the server. The user terminal 101 can be a mobile phone, personal computer, tablet computer, e-book reader, wearable device, device for information interaction based on AR (Augmented Reality) / VR (Virtual Reality), and laptop computer, etc. Server 102 is used to perform medical entity recognition based on user text of user voice data uploaded by user terminal to obtain candidate medical entities, return the candidate medical entities to user terminal, call intelligent agent instance to generate diagnosis and treatment data based on multi-dimensional mapping parameters and user text uploaded by user terminal, and return the diagnosis and treatment data to user terminal; server 102 can be one or more servers, a server cluster composed of several servers, or a cloud server of a cloud computing platform. The interactive intelligent agent corresponding to the medical service provider can be deployed on server 102 or other servers. There can be multiple intelligent agent instances, including intelligent agent instance 103-1... intelligent agent instance 103-N; In this implementation environment, user terminal 101 can capture images of users interacting with the intelligent agent instance corresponding to the medical service provider via voice, and perform interaction mode switching detection based on the captured action image sequence. If the switching detection passes, the user's voice data is uploaded to server 102, and candidate medical entities are obtained by medical entity recognition based on user text returned by server 102 based on user voice data. Multi-dimensional interaction mapping is performed on the candidate medical entities based on user interaction video, and the obtained multi-dimensional mapping parameters of the candidate medical entities are uploaded to server 102. The diagnosis and treatment data generated by the calling intelligent agent instance based on the multi-dimensional mapping parameters and user text returned by the server are displayed. In this way, parameter mapping of candidate medical entities is performed from multiple interaction dimensions, and diagnosis and treatment data is generated based on the multi-dimensional mapping parameters.
[0014] One or more embodiments of a video interactive processing method provided in this specification are as follows: Reference Figure 2 The video interaction processing method provided in this embodiment can be applied to user terminals, and specifically includes steps S202 to S208.
[0015] Step S202: Image acquisition is performed on the user who interacts with the intelligent agent instance corresponding to the medical service provider via voice, and interaction mode switching detection is performed based on the acquired action image sequence.
[0016] The medical service provider mentioned in this embodiment refers to a service provider related to medical care. The medical service provider can be a medical entity, specifically a medical person, a medical institution, a medical department of a medical institution, and / or a medical and health institution. Medical persons may include doctors and / or nurses, and medical institutions may be hospitals, medical clinics, and / or vaccination institutions.
[0017] The interactive intelligent agent corresponding to the medical service provider refers to an interactive intelligent agent set up specifically for the medical service provider. This interactive intelligent agent can intelligently interact with users on behalf of the medical service provider. The interactive intelligent agent can be an intelligent avatar of the medical service provider. An intelligent agent is an executor that autonomously performs tasks, makes decisions, and learns and adjusts according to environmental changes. Specifically, an intelligent agent can be an executor that integrates the ability to generate diagnostic and treatment data. The intelligent agent can receive natural language commands or interactive operations input by the user, call one or more models, and drive the models to generate corresponding diagnostic and treatment data. Here, the intelligent agent can specifically be an intelligent agent application, an intelligent agent system, or a Large Language Model (LLM), such as a large language model using the Transformer architecture. There can be at least one intelligent agent instance. An intelligent agent instance refers to a specific functional entity that can run independently, formed by instantiating and deploying a program template of an intelligent agent with autonomous perception, decision-making, and execution capabilities in a computer system. The action image sequence refers to an image action image frame sequence composed of at least one action image frame collected from the user.
[0018] In specific implementation, image acquisition is performed on users who interact with the intelligent agent instance corresponding to the medical service provider through voice to obtain action image sequences, and interaction mode switching detection is performed based on the action image sequences; optionally, the medical service provider includes medical personnel, the interactive intelligent agent is created after the medical personnel authorize medical interaction, the interactive intelligent agent is obtained based on a large medical model, and the large medical model is obtained after training a pre-trained medical model based on the medical personnel's diagnosis and treatment data; Specifically, medical interaction authorization can be executed by obtaining medical interaction authorization instructions from medical personnel; the large medical model can be a large model in the medical field, and the pre-trained medical model can be obtained by training a benchmark large model in the medical field based on global medical knowledge and / or medical knowledge of the medical field to which the medical personnel belong. Global medical knowledge can be medical knowledge of all medical fields, and the medical field to which the medical personnel belong can be the field in which the medical personnel perform medical treatment, such as the field of cardiovascular medicine to which the medical personnel belong; the pre-trained medical model, the benchmark large model, and the large medical model can be a large language model based on the Transformer architecture; the medical personnel's diagnosis and treatment data can be the diagnosis and treatment data obtained by the medical personnel in diagnosing and treating patients online and / or offline, and the diagnosis and treatment data can include the patient's medical records and / or prescriptions.
[0019] In addition to the above-mentioned implementation method of creating interactive intelligent agents after medical personnel authorize medical interaction, in order to improve the compliance of interactive intelligent agents and enable them to provide interactive functions to the outside world in a safe and compliant manner, optionally, the interactive intelligent agent is created after the medical service provider performs medical identity authentication in the interaction subroutine, and the interactive intelligent agent is obtained by configuring the parameters of the initial interactive intelligent agent based on the medical service provider's diagnosis and treatment records.
[0020] The initial interactive agent can be an interactive agent that has learned global medical knowledge; the interactive agent can be obtained by the medical service provider and / or the administrator of the interactive subroutine by configuring the parameters of the initial interactive agent based on the medical service provider's diagnosis and treatment records; this embodiment can be applied to user terminals, interactive subroutines or interactive applications, where the interactive subroutine can be a medical subroutine and the interactive application can be a medical application.
[0021] In practical applications, users may have a need to conduct voice interactions online regarding medical issues. To address this, and to improve the convenience of online consultations and meet the diverse needs of users, in one optional implementation of this embodiment, before capturing images of the user interacting with the intelligent agent instance corresponding to the medical service provider and performing interaction mode switching detection based on the captured action image sequence, the following operations can also be performed: Based on the target agent tag selected by the user from multiple agent tags corresponding to medical personnel displayed within the application, the user is redirected from the application to the interactive subroutine. Based on the voice interaction commands of the interactive agent corresponding to the target agent tag submitted by the user in the interaction subroutine, the user performs voice interaction with the agent instance.
[0022] Among them, the intelligent agent tags corresponding to multiple medical personnel can be the intelligent agent tags corresponding to each medical personnel among multiple medical personnel. The intelligent agent tags refer to the tags used to identify the interactive intelligent agents corresponding to each medical personnel. The intelligent agent tags corresponding to each medical personnel can record the medical personnel's name, the medical personnel's avatar, the medical field to which the medical personnel belong, the medical institution to which the medical personnel belong, and / or the diseases that the medical personnel are good at treating. Optionally, the intelligent agent tag is obtained by matching the intelligent agent based on the user's medical interaction intent. The medical interaction intent is obtained by recognizing the intent of the medical interaction voice entered by the user through the intelligent interaction control in the application. The intelligent interaction control is displayed after detecting that the user performs a page switching operation in the application. The application can be a payment application, a medical application, a shopping application, a goods delivery application, an office application, an attendance application, or a ticketing application. The page switching operation here can be a page pull-down operation, a page pull-up operation, a page pull-left operation, a page pull-right operation, or a page pull-diagonal operation.
[0023] Specifically, users can submit page switching commands on the target page within the application. The user terminal can then jump from the target page to the intelligent interactive page within the application based on the page switching command. Optionally, intelligent interactive controls are displayed on the intelligent interactive page within the application. The target page can be any page within the application, such as the application's homepage or the page to jump to after triggering the application's homepage. The page switching command can be submitted based on a pull-down, pull-up, pull-left, pull-right, or diagonal pull operation on the target page. Users can input medical interaction voice by triggering the intelligent interaction controls on the intelligent interaction page within the application. The user terminal can perform intent recognition based on the medical interaction voice to obtain the medical interaction intent. The user terminal can upload the medical interaction intent to the server. The server can perform intelligent agent matching based on the medical interaction intent to obtain intelligent agent tags corresponding to multiple medical personnel and return them to the user terminal. The intelligent agent tags corresponding to multiple medical personnel can specifically be a tag list composed of the intelligent agent tags corresponding to each of the multiple medical personnel. The user terminal can jump from the application to the interaction subroutine based on the target intelligent agent tag selected by the user from the multiple intelligent agent tags corresponding to the medical personnel displayed in the application. The user can submit voice interaction commands based on the voice interaction controls of the interactive intelligent agent corresponding to the target intelligent agent tag displayed in the interaction subroutine. The user terminal can then perform voice interaction between the user and the intelligent agent instance based on the voice interaction commands.
[0024] For example Figure 3 The voice interaction page shown shows that the user selects the xx agent, which belongs to the Department of Cardiovascular Medicine of xx Medical Institution. Area 301 displays a video interaction reminder to remind the user to interact with the agent. The voice interaction page also includes a mute control, a hang-up control, a subtitle control, and a voice input reminder that says "I am listening, please speak".
[0025] In practical application scenarios, users interact with the intelligent agent instance corresponding to the medical service provider through voice. During this process, users may need to interact with the intelligent agent instance through video to facilitate visual analysis of the user and improve the comprehensiveness and completeness of subsequent medical data. To address this, an image acquisition component can be called to capture images of users interacting with the intelligent agent instance corresponding to the medical service provider through voice. Subsequently, interaction mode switching detection can be performed based on the acquired action image sequence. Specifically, user action recognition can be performed based on the acquired action image sequence. If the recognized user action is the target action, the action amplitude can be calculated based on the action image sequence. If the action amplitude is greater than the amplitude threshold, the switching detection is considered successful. In the process of calculating the action amplitude based on the action image sequence, the action image sequence can be input into an image point localization model to obtain facial localization points. The action amplitude is then calculated based on the facial localization points. Here, the image point localization model can be a keypoint detection model, such as a face model (MediaPipe Face Mesh). Alternatively, the number of actions can also be calculated based on the action image sequence. If the number of actions is greater than the number of actions threshold, the interaction mode switching detection is considered successful. Here, the target action can be a nodding action and / or a blinking action. The action amplitude can be the nodding action amplitude, and the number of actions can be the blinking action number.
[0026] In addition to the implementation methods provided above, to personalize the interaction mode switching detection and improve the user experience, this embodiment provides an optional implementation method in which, during the interaction mode switching detection process based on the acquired motion image sequence, the interaction mode switching detection is performed on the motion image sequence according to the switching detection method corresponding to the user posture obtained based on device motion data recognition. Specifically, the interaction mode switching detection can be performed in the following manner: The user's posture is obtained by recognizing the user's posture based on device movement data; Interaction mode switching detection is performed on the action image sequence according to the switching detection method corresponding to the user's posture.
[0027] The user posture may include a lying posture, a standing posture, and / or a sitting posture, or the user posture may include a lying posture or a non-lying posture; the device motion data may be the device acceleration of the user terminal, which may be obtained based on the acceleration sensor.
[0028] Specifically, in the process of obtaining the user's posture by recognizing the user's posture based on the device's movement data, the gravity component can be obtained by filtering high-frequency noise from the device's acceleration using a low-pass filter, and the magnitude of the gravity component can be calculated. The magnitude is then matched with a preset threshold for the magnitude of the gravity component of the user's posture to obtain the user posture corresponding to the matched magnitude threshold.
[0029] Based on this, in the first optional implementation provided in this embodiment, during the interaction mode switching detection of the action image sequence according to the switching detection method corresponding to the user posture, if the user posture is the first user posture, the interaction mode switching detection is performed based on the key point coordinates extracted from each action image in the action image sequence. Specifically, the interaction mode switching detection of the action image sequence can be performed in the following way: If the user pose is the first user pose, construct at least two motion vectors based on the keypoint coordinates extracted from each motion image in the motion image sequence; The motion amplitude is detected based on at least two motion vectors to obtain the user motion amplitude corresponding to each motion image. If the number of motion images with user motion amplitudes exceeding the preset motion amplitude is greater than the number threshold, the interaction mode switching detection is determined to be successful.
[0030] The first user posture can be a standing posture or a sitting posture.
[0031] Specifically, in the process of detecting the amplitude of motion based on at least two motion vectors and obtaining the user's motion amplitude corresponding to each motion image, the angle between the two motion vectors of each motion image can be calculated. Specifically, the cosine value of the two motion vectors can be calculated and converted into an angle as the motion amplitude of the motion image.
[0032] For example, key point coordinates of three key points—the center of the eyebrows, the bridge of the nose, and the chin—are extracted from the action image sequence. An action vector v1 is constructed based on the key point coordinates of the center of the eyebrows and the bridge of the nose, and an action vector v2 is constructed based on the key point coordinates of the chin and the bridge of the nose. The angle between the two action vectors of each action image is calculated as the action amplitude of the action image.
[0033] Furthermore, in the second optional implementation provided in this embodiment, during the process of performing interaction mode switching detection on the motion image sequence according to the switching detection method corresponding to the user posture, if the user posture is the second user posture, the interaction mode switching detection can be performed based on the positioning distance between the eye positioning points in each motion image included in the motion image sequence, or the orbicularis oculi muscle movement can be detected based on the motion image sequence, thereby determining the switching detection result. Specifically, the interaction mode switching detection of the motion image sequence can be performed in the following manner: If the user pose is the second user pose, calculate the positioning distance between the eye positioning points based on the positioning point coordinates of the eye positioning points in each action image contained in the action image sequence. The system determines whether the number of eye movements made by the user exceeds a threshold based on the positioning distance. If so, the interaction mode switching detection is confirmed to have passed.
[0034] The second user posture can be a lying posture, and the positioning distance can be the Euclidean distance between the key points of the eyelids of both eyes.
[0035] Specifically, in the process of calculating the positioning distance between eye positioning points based on the coordinates of the eye positioning points in each action image contained in the action image sequence, the coordinates of the eyelid positioning points in each action image contained in the action image sequence can be extracted, and the positioning distance between the eyelid positioning points can be calculated based on the positioning point coordinates. If the positioning distance is less than the eye action reference distance, the duration of the eye action can be determined. If the duration of the action is within the preset duration range, the number of eye actions can be accumulated. If the accumulated number of eye actions exceeds the threshold, the interaction mode switching detection can be determined to have passed; otherwise, the interaction mode switching detection can be determined to have failed. In this way, the action image sequence is used to determine whether to switch the voice interaction mode, improving the convenience of mode switching and enabling quick switching of interaction modes for special user groups such as elderly users.
[0036] Step S204: If the switching detection passes, upload the user's voice data to the server and obtain the candidate medical entities obtained by medical entity recognition based on the user's text returned by the server.
[0037] The above-mentioned process involves image acquisition of users interacting with the intelligent agent instance corresponding to the medical service provider via voice, and detection of interaction mode switching based on the acquired action image sequence. In this step, if the switching detection passes, the user's voice data is uploaded to the server, and candidate medical entities are obtained through medical entity recognition based on the user's voice data returned by the server. Correspondingly, the server can obtain candidate medical entities through medical entity recognition based on the user's voice data uploaded by the user terminal, and return the candidate medical entities to the user terminal. In this embodiment, the user text can be the user text obtained through voice recognition of the user's voice data; the candidate medical entities can be candidate diseases, candidate drugs, and / or candidate auxiliary treatment items. Auxiliary treatment items can be surgical treatment, nebulization treatment, and / or massage treatment, and can also be other types of auxiliary treatment items.
[0038] In practice, if the interaction mode switching detection passes, the service program can switch from voice interaction mode to video interaction mode, upload the collected user voice data to the server, and obtain the candidate medical entities obtained by medical entity recognition based on the user text returned by the server based on the user voice data. In the process of medical entity recognition based on user text of user voice data, candidate medical entities can be obtained by using user text of user voice data collected in video interaction mode, user text of user voice data collected in voice interaction mode, and / or user text of medical interaction voice recorded by the user through intelligent interactive controls.
[0039] In the specific execution process, after the user's voice data is uploaded to the server, the server can also call the agent instance of the corresponding interactive agent of the medical service provider to generate questions based on the user's text of the voice data to obtain response questions, construct response voice based on the response questions, and return the response voice and / or response questions to the user terminal. The user terminal can display the response questions and hear the voice broadcast of the response voice on the video interaction page of the agent instance. Subsequently, the candidate medical entities can be updated based on the user's feedback answers to the response questions, and multi-dimensional interactive mapping can be performed on the updated candidate medical entities.
[0040] It should be noted that the above-mentioned operation of uploading user voice data to the server and obtaining candidate medical entities obtained by medical entity recognition based on user text returned by the server if the handover detection passes can be replaced by uploading user voice data to the server based on the handover detection result and obtaining candidate medical entities obtained by medical entity recognition based on user text returned by the server; or it can be replaced by obtaining candidate medical entities obtained by medical entity recognition based on user text returned by the server if the handover detection passes, and combining it with other processing steps provided in this embodiment to form a new implementation method.
[0041] Step S206: Perform multi-dimensional interactive mapping on the candidate medical entities based on the user interaction video to obtain the multi-dimensional mapping parameters of the candidate medical entities.
[0042] The above-mentioned candidate medical entities are obtained by performing medical entity recognition on user text based on user voice data returned by the server. In this step, multi-dimensional interaction mapping is performed on the candidate medical entities based on user interaction video to obtain multi-dimensional mapping parameters of the candidate medical entities. This achieves interaction mapping of candidate medical entities from multiple dimensions, thereby improving the accuracy and comprehensiveness of multi-dimensional mapping parameters.
[0043] The user interaction video mentioned in this embodiment refers to the user video captured by the video capture component. The user interaction video can be a full-body video of the user or a facial video of the user, for example. Figure 4 As shown, the user terminal collects user interaction video in video interaction mode; the multidimensional mapping parameters may include the user's visual prediction index of facial visual attributes corresponding to the candidate medical entity, the user's physiological index of the candidate medical entity, and / or the user's physiological parameters of the physiological signals corresponding to the candidate medical entity. The visual prediction index may be the predicted probability or prediction score of the user belonging to the facial visual attribute. The facial visual attribute may be facial micro-expression, complexion, facial skin condition, and / or facial vascular condition. For example, if the candidate medical entity is the candidate disease - angina pectoris, the facial micro-expression corresponding to angina pectoris is the pain micro-expression, and the visual prediction index is the predicted probability that the user's facial micro-expression belongs to the pain micro-expression; the physiological index can be calculated based on the physiological parameters of the physiological signals corresponding to the candidate medical entity.
[0044] In practical implementation, to improve the comprehensiveness of multidimensional mapping parameters, multidimensional interactive mapping is performed from a fine-grained perspective to obtain the multidimensional mapping parameters of candidate medical entities. In one optional implementation provided in this embodiment, during the process of performing multidimensional interactive mapping on candidate medical entities based on user interaction videos to obtain the multidimensional mapping parameters of candidate medical entities, the visual prediction index of the user's facial visual attributes corresponding to the candidate medical entities is determined according to the user interaction videos, and the physiological index of the user for the candidate medical entities is also determined. Specifically, the following operations can be performed: The user interaction video and the facial visual attributes corresponding to the candidate medical entities are input into the visual model to predict visual indicators and obtain the user's visual prediction indicators for facial visual attributes. Based on the user interaction video, determine the physiological parameters of the physiological signals of the user in relation to the candidate medical entity, and determine the physiological indicators of the user in relation to the candidate medical entity based on the physiological parameters.
[0045] Among them, facial visual attributes can be facial expressions, facial micro-expressions, complexion, facial skin condition and / or facial vascular condition. Each candidate disease can correspond to its own facial visual attributes. For example, the facial micro-expression corresponding to angina pectoris is the pain micro-expression. Each candidate disease can also correspond to its own physiological signals. For example, the physiological signals of angina pectoris are heart rate and / or anxiety level. The visual model can be a lightweight visual model. For example, the visual model is a lightweight CNN (Convolutional Neural Network) model MobileNetV3 (Lightweight Deep Convolutional Neural Network).
[0046] Specifically, in the process of determining the physiological parameters of the user's physiological signals corresponding to candidate medical entities based on user interaction videos, the physiological parameters of the user's physiological signals corresponding to candidate medical entities can be extracted from the interactive video frames contained in the user interaction video using a remote photoplethysmography device; for example, the heart rate value of the user's heart rate signal can be extracted from the interactive video frames, and the physiological indicators of the user for candidate medical entities can be further determined based on the difference between the heart rate value and the baseline heart rate value; the physiological indicators can measure the degree of user adaptation to candidate medical entities from a physiological perspective, and the visual prediction indicators can measure the degree of user adaptation to candidate medical entities from a visual perspective; the visual prediction indicators, physiological parameters, and / or physiological indicators can be used as multidimensional mapping parameters; In the process of inputting user interaction videos and facial visual attributes corresponding to candidate medical entities into a visual model for visual index prediction, and obtaining the user's visual prediction index for facial visual attributes, the visual positioning module in the visual model can perform visual positioning of interactive video frames in the user interaction video to obtain visual positioning points. The index prediction module in the visual model can then perform visual index prediction based on the coordinates of these positioning points to obtain the user's visual prediction index for facial visual attributes. Additionally, the user terminal can connect the positioning points of the interactive video frames based on the visual positioning points, and then render and display the connected interactive video frames. For example... Figure 5 As shown, localization points are connected in the facial area of the user's interactive video frame, and the connected interactive video frame is rendered and displayed.
[0047] Step S208: Upload the multidimensional mapping parameters to the server and display the diagnosis and treatment data generated by the calling agent instance based on the multidimensional mapping parameters and user text returned by the server.
[0048] The above-mentioned method performs multi-dimensional interactive mapping on candidate medical entities based on user-interactive video to obtain multi-dimensional mapping parameters for the candidate medical entities. In this step, the multi-dimensional mapping parameters are uploaded to the server, and the diagnostic data generated by the calling intelligent agent instance based on the multi-dimensional mapping parameters and user text returned by the server is displayed. Correspondingly, the server can call the intelligent agent instance to generate diagnostic data based on the multi-dimensional mapping parameters uploaded by the user terminal and user text, and return the diagnostic data to the user terminal. The diagnostic data mentioned in this embodiment can be a diagnostic report or a diagnostic tag; for example... Figure 6 As shown, area 601 displays the initial diagnosis label, which includes preliminary symptom diagnosis, chief complaint: persistent headache, and inference from the medical history: migraine, tension headache. Area 602 displays the visual diagnosis label, which includes micro-expression diagnosis, with the following results: pain index xx1, anxiety level xx2, and abnormal data: heart rate xx3 (↑).
[0049] In practice, the user terminal can upload multidimensional mapping parameters to the server, and the server can input the multidimensional mapping parameters into the agent instance. The specific microservice framework of the server can receive the multidimensional mapping parameters. Here, the microservice framework can be Python Flask (a lightweight web microservice framework based on the Python language). The multidimensional mapping parameters are input into the agent instance, and the agent instance calls the fit calculation model to calculate the user fit of the candidate medical entities. The candidate medical entities are sorted according to the user fit to obtain the pushed medical entities, and the medical big model is called to generate the diagnosis and treatment report of the pushed medical entities to obtain the diagnosis and treatment report. Alternatively, the agent instance can call the report engine to generate the diagnosis and treatment report of the pushed medical entities to obtain the diagnosis and treatment report. For example, the report engine can be a template engine. The diagnosis and treatment report can include emotion fluctuation curves and heart rate trend graphs.
[0050] In the specific execution process, based on the above-mentioned determination of the user's physiological indicators for candidate medical entities based on physiological parameters, on the one hand, a diagnosis and treatment report for the candidate medical entities can be generated according to visual prediction indicators, physiological indicators, and text indicators calculated based on user text; in the first optional implementation method provided in this embodiment, the diagnosis and treatment report is generated in the following way, that is, the diagnosis and treatment data can be generated in the following way: User fit of candidate medical entities is calculated based on visual prediction metrics, physiological metrics, and text metrics calculated from user text. Candidate medical entities are sorted according to user suitability to obtain the medical entities to be pushed, and the intelligent agent instance is called to generate the diagnosis and treatment report of the pushed medical entity to obtain the diagnosis and treatment report as diagnosis and treatment data.
[0051] Among them, text metrics can measure the degree to which a user is adapted to a candidate medical entity from a textual perspective; user fit can be the degree of fit between a user and a candidate medical entity; and diagnosis and treatment reports can be intelligent consultation reports.
[0052] Specifically, in the process of calculating the user fit of candidate medical entities based on visual prediction metrics, physiological metrics, and text metrics calculated based on user text, the fit can be calculated based on the visual prediction metrics, physiological metrics, and text metrics calculated based on user text, as well as the corresponding weights of the three metrics, to obtain the user fit of candidate medical entities. Here, the process of calculating the user fit of candidate medical entities based on visual prediction metrics, physiological metrics, and text metrics calculated based on user text can be executed by the intelligent agent instance calling the fit model.
[0053] Based on this, in an optional implementation of this embodiment, the following operations are performed during the process of generating a diagnosis and treatment report for a medical entity: Based on the medical service provider's medical style, text-based diagnosis and treatment data of medical entities are generated and pushed to the user. Based on the facial visual parameters of the user in relation to the pushed medical entity identified from the user interaction video, visual diagnostic data of the pushed medical entity is generated, and a diagnostic report is generated based on the text diagnostic data and the visual diagnostic data.
[0054] Physiological parameters may include facial visual parameters, which may include facial micro-expressions, complexion, facial skin condition and / or facial vascular condition; visual diagnostic data may include facial visual parameters and / or diagnostic and treatment recommendations based on facial visual parameters.
[0055] Specifically, in the process of generating and pushing text-based medical data for medical entities based on user text according to the medical style of the medical service provider, medical data can be obtained by matching medical information in a medical database based on user text, and the medical data associated with the pushed medical entity can be used as text-based medical data; in the process of generating visual medical data for the pushed medical entity based on the user's facial visual parameters identified from the user's interactive video, treatment suggestions can be generated based on the facial visual parameters, and the facial visual parameters and / or treatment suggestions can be used as visual medical data; the treatment in this embodiment can be replaced by interaction, consultation and / or assisted treatment, and the treatment report in this embodiment is only for reference.
[0056] For example Figure 7 The intelligent consultation report shown contains text-based diagnostic data, i.e., detailed preliminary symptom diagnosis, and visual diagnostic data, i.e., detailed micro-expression diagnosis. The specific diagnostic content is shown in the figure and will not be elaborated here.
[0057] Based on the above-mentioned determination of the user's physiological indicators for candidate medical entities based on physiological parameters, on the other hand, diagnostic labels can be constructed based on visual prediction indicators, physiological parameters, and user text. These diagnostic labels may include initial diagnostic labels and / or visual diagnostic labels. In the second optional implementation provided in this embodiment, the diagnostic data is generated in the following manner: Initial diagnosis and treatment labels are generated based on user text and candidate diseases; Visual diagnostic labels are constructed based on visual prediction indicators and physiological parameters to obtain visual diagnostic labels, and the initial diagnostic labels and visual diagnostic labels are used as diagnostic data.
[0058] Based on this, in an optional implementation of this embodiment, the following operations are performed during the process of constructing visual diagnosis labels based on visual prediction indicators and physiological parameters to obtain visual diagnosis labels: The fundamental frequency value sequence is obtained by calculating the fundamental frequency based on the user's voice data, and the psychological index is obtained by calculating the psychological index based on the fundamental frequency value sequence. Visual diagnostic labels are constructed based on visual predictive indicators, psychological indicators, and physiological parameters.
[0059] Among them, psychological indicators can be psychological tension or psychological anxiety.
[0060] Specifically, in the process of calculating psychological indicators based on the fundamental frequency value sequence, the fluctuation range of the fundamental frequency value can be determined based on each fundamental frequency value in the sequence. The user's pitch and / or voice tremor can be determined based on the fluctuation range distribution. The psychological indicator can then be calculated based on the user's pitch and / or voice tremor. Here, the process of calculating the fundamental frequency value sequence based on the user's voice data, determining the fluctuation range distribution of the fundamental frequency value based on each fundamental frequency value in the sequence, and determining the user's pitch and / or voice tremor based on the fluctuation range distribution can be executed by a speech engine, such as OpenSMILE (Open-Source Speech and Music Interpretation by Large SpaceExtraction, an open-source speech parsing toolkit based on large-scale feature space extraction).
[0061] Furthermore, after uploading the multidimensional mapping parameters to the server and displaying and executing the diagnostic data generated by the intelligent agent instance based on the multidimensional mapping parameters and user text returned by the server, in an optional implementation of this embodiment, the following operations are also performed: In response to a user's access command for either the initial diagnosis label or the visual diagnosis label, the system generates and displays diagnosis reports corresponding to the initial diagnosis label and the visual diagnosis label, respectively, based on visual prediction indicators and physiological indicators.
[0062] Alternatively, in response to an access command, a treatment report corresponding to any treatment label can be generated and displayed, or the treatment report corresponding to the initial treatment label and visual treatment label generated by the calling agent instance based on multidimensional mapping parameters and user text can be obtained from the server; or the treatment report corresponding to any treatment label generated by the calling agent instance based on multidimensional mapping parameters and user text can be obtained from the server. In this embodiment, the generated treatment report can be sent to the user terminal after the medical service provider reviews and approves it, or the interactive agent can be fine-tuned based on the correction data of the generated treatment report from the medical service provider.
[0063] In this embodiment, step S202 can be replaced by performing voice recognition on the user's voice during voice interaction with the intelligent agent instance corresponding to the medical service provider. If the recognized voice text triggers the interaction mode switching condition, the voice interaction mode of the interaction subroutine can be switched to video interaction mode. Further, step S204 can be executed, which involves uploading the user's voice data to the server and obtaining candidate medical entities obtained from medical entity recognition based on the user's voice text returned by the server. Steps S206 to S208 can then be executed. The voice text triggering the interaction mode switching condition can be that the recognition result of medical entity recognition based on voice text is empty; or, after step S202 is replaced, it can also be based on the user's touch on the user's terminal screen. The operation involves detecting body parts within the user-interactive video frame, obtaining the body parts, and sending the body parts to the server. The server can then perform medical entity recognition based on the body parts and / or user text from the user's voice data to obtain candidate medical entities and return them to the user terminal. Subsequently, the user terminal can execute steps S206 to S208. Alternatively, after step S202 in this embodiment is executed, the user terminal can perform body part detection within the user-interactive video frame based on the user's triggered operation on the user's terminal screen, obtain the body parts, and upload the user's voice data and / or body parts to the server. The server can then perform medical entity recognition based on the user text from the user's voice data and / or body parts to obtain candidate medical entities. This, combined with other processing steps provided in this embodiment, forms a new implementation method.
[0064] It should be noted that the user data obtained in this specification, such as motion image sequences, user voice data, user interaction videos, device movement data, etc., has been authorized by the user and does not involve user privacy. Specifically, authorization can be granted during the user's first access to the application, its interactive subroutines, or the interactive application itself, or each subsequent access to the application, its interactive subroutines, or the interactive application. The aforementioned optional implementation methods for candidate diseases can be replaced with candidate drugs, candidate adjunctive therapies, and / or candidate diseases.
[0065] It should be added that each optional implementation method and each feasible execution method in steps S202 to S208 provided in this embodiment can be executed independently as needed, or they can be combined and referenced with each other. At the same time, each specific execution step in each optional implementation method or each feasible execution method can also be executed independently or combined as needed. The execution conditions of "if" or "under what circumstances" involved in each step or operation can be directly deleted, and subsequent operations can be executed. This embodiment does not make specific limitations on this.
[0066] It should also be added that, depending on the actual application scenario, step S202 and any of the subsequent steps S204 to S208 can be deleted, or any feature in any step can be deleted. For example, the agent instance of the interactive agent corresponding to the medical service provider in step S202 can be deleted. The execution order of steps S202 to S208 can also be arbitrary.
[0067] The above-described video interaction processing method can be implemented by a user terminal. The following method embodiment provides another video interaction processing method that can be implemented by a server. The two can cooperate with each other during execution. Therefore, when reading the above implementation process, you can refer to the corresponding content of the following other video interaction processing method embodiment. Similarly, when reading the following other video interaction processing method embodiment, you can also refer to the corresponding content of the above method embodiment.
[0068] The following description uses the application of a video interaction processing method provided in this embodiment in a video interaction scenario as an example to further illustrate the video interaction processing method provided in this embodiment. (See also...) Figure 8 The video interaction processing method, which is applied to video interaction scenarios, can be used on user terminals and includes the following steps.
[0069] Step S802: Based on the target intelligent agent tag selected by the user from the intelligent agent tags corresponding to multiple medical service providers displayed in the application, the user jumps from the application to the interactive subroutine.
[0070] Step S804: Based on the voice interaction command of the interactive intelligent agent corresponding to the target intelligent agent tag submitted by the user in the interaction subroutine, conduct voice interaction between the user and the intelligent agent instance of the interactive intelligent agent corresponding to the medical service provider.
[0071] Step S806: Image acquisition is performed on the user who interacts with the intelligent agent instance corresponding to the medical service provider via voice, and interaction mode switching detection is performed based on the acquired action image sequence.
[0072] Step S808: If the switch detection passes, upload the user voice data collected by the interaction subroutine in video interaction mode to the server.
[0073] Step S812: Input the user interaction video and the facial visual attributes corresponding to the candidate diseases returned by the server into the visual model to predict visual indicators and obtain the user's visual prediction indicators for facial visual attributes.
[0074] Step S814: Determine the physiological parameters of the physiological signals corresponding to the candidate disease based on the user's interactive video, and determine the physiological indicators of the candidate disease based on the physiological parameters.
[0075] Step S816: Upload the visual prediction index, physiological parameters, and physiological indicators to the server.
[0076] Step S822: Render and display the initial diagnostic labels and visual diagnostic labels.
[0077] Step S824: Based on the user's access instruction to any one of the initial diagnosis label and visual diagnosis label, obtain and display the diagnosis report generated by the calling agent instance based on the user's text, visual prediction indicators and physiological indicators returned by the server.
[0078] Steps S802 to S808, S812 to S816, and S822 to S824 provided in this embodiment are executed by the user terminal. It should be noted that the steps S802 to S808, S812 to S816, and S822 to S824 executed by the user terminal can cooperate with steps S810, S818 to S820 executed by the server in the following embodiment during execution. Therefore, when reading this embodiment, please refer to the corresponding content of steps S810, S818 to S820 provided in the following method embodiment, and when reading the following method embodiment, please refer to the corresponding content of steps S802 to S808, S812 to S816, and S822 to S824 provided in this embodiment.
[0079] It should be noted that any one or more of steps S802 to S808, S812 to S816, and S822 to S824 can be replaced with the corresponding technical means provided by steps S202 to S208 as needed for implementation and deployment. Furthermore, any one or more of steps S802 to S808, S812 to S816, and S822 to S824 can be combined to form a new implementation method as needed for implementation and deployment. In addition, any one or more of steps S802 to S808, S812 to S816, and S822 to S824 can also be combined with one or more of the steps provided by steps S202 to S208 to form a new implementation method, or combined with one or more of the optional implementation methods provided by steps S202 to S208 to form a new implementation method, as needed for actual deployment. These will not be elaborated on here.
[0080] One or more embodiments of another video interaction processing method provided in this specification are as follows: Reference Figure 9 The video interaction processing method provided in this embodiment can be applied to a server, specifically including steps S902 to S906.
[0081] Step S902: Based on the user text of the user voice data uploaded by the user terminal, medical entity recognition is performed to obtain candidate medical entities.
[0082] Optionally, the user voice data is uploaded after the user's voice interaction with the intelligent agent instance corresponding to the medical service provider is captured by image acquisition, and the interaction mode switching is detected and passed based on the acquired action image sequence.
[0083] In this embodiment, the medical service provider refers to a service provider related to medical care. The medical service provider can be a medical entity, specifically a medical person, a medical institution, a medical department of a medical institution, and / or a medical and health institution. Medical persons can include doctors and / or nurses, and medical institutions can be hospitals, medical clinics, and / or vaccination institutions. In this embodiment, the server can be deployed with multiple interactive intelligent agents corresponding to each medical service provider. Each medical service provider can have one or more intelligent agent instances.
[0084] The interactive intelligent agent corresponding to the medical service provider refers to an interactive intelligent agent set up specifically for the medical service provider. This interactive intelligent agent can intelligently interact with users on behalf of the medical service provider. The interactive intelligent agent can be an intelligent avatar of the medical service provider. An intelligent agent is an executor that autonomously performs tasks, makes decisions, and learns and adjusts according to environmental changes. Specifically, an intelligent agent can be an executor that integrates the ability to generate diagnostic and treatment data. The intelligent agent can receive natural language commands or interactive operations input by the user, call one or more models, and drive the models to generate corresponding diagnostic and treatment data. Here, the intelligent agent can specifically be an intelligent agent application, an intelligent agent system, or a Large Language Model (LLM), such as a large language model using the Transformer architecture. There can be at least one intelligent agent instance. An intelligent agent instance refers to a specific functional entity that can run independently, formed by instantiating and deploying a program template of an intelligent agent with autonomous perception, decision-making, and execution capabilities in a computer system. The action image sequence refers to an image action image frame sequence composed of at least one action image frame collected from the user.
[0085] In practice, the user terminal can capture images of the user interacting with the intelligent agent instance corresponding to the medical service provider through voice to obtain action image sequences, and perform interaction mode switching detection based on the action image sequences. If the switching detection passes, the user's voice data is uploaded to the server. Correspondingly, the server can perform medical entity recognition based on the user text of the user's voice data uploaded by the user terminal to obtain candidate medical entities.
[0086] In the process of medical entity recognition based on user text of user voice data, candidate medical entities can be obtained by using user text of user voice data collected in video interaction mode, user text of user voice data collected in voice interaction mode, and / or user text of medical interaction voice recorded by the user through intelligent interactive controls.
[0087] The user text can be obtained by speech recognition of user voice data; the candidate medical entity can be a candidate disease, a candidate drug and / or a candidate auxiliary treatment item, and the auxiliary treatment item can be surgical treatment, nebulization treatment and / or massage treatment, or other types of auxiliary treatment items.
[0088] In the specific execution process, after the user terminal uploads the user's voice data to the server, the server can also call the intelligent agent instance of the corresponding interactive intelligent agent of the medical service provider to generate questions based on the user's text of the voice data to obtain response questions, construct response voice based on the response questions, and return the response voice and / or response questions to the user terminal. The user terminal can display the response questions and listen to the voice broadcast of the response voice on the video interaction page of the intelligent agent instance. Subsequently, the candidate medical entities can be updated based on the user's feedback answers to the response questions, and multi-dimensional interactive mapping can be performed on the updated candidate medical entities.
[0089] Step S904: Return the candidate medical entity to the user terminal to obtain the multidimensional mapping parameters of the candidate medical entity by performing multidimensional interactive mapping on the candidate medical entity based on the user interactive video.
[0090] Step S906: Invoke the intelligent agent instance to generate diagnostic data based on the multidimensional mapping parameters uploaded by the user terminal and the user text, and return the diagnostic data to the user terminal.
[0091] The user text mentioned in this embodiment can be user text obtained by speech recognition of user voice data; the candidate medical entity can be a candidate disease and / or a candidate drug.
[0092] The user terminal can perform multi-dimensional interactive mapping on candidate medical entities based on user interactive video to obtain multi-dimensional mapping parameters of the candidate medical entities. This enables interactive mapping of candidate medical entities from multiple dimensions, improving the accuracy and comprehensiveness of the multi-dimensional mapping parameters. The multi-dimensional mapping parameters are then uploaded to the server. The server can call the intelligent agent instance to generate diagnosis and treatment data based on the multi-dimensional mapping parameters uploaded by the user terminal and the user text, and return the diagnosis and treatment data to the user terminal.
[0093] Among them, user interaction video refers to user video captured by calling the video capture component. User interaction video can be a full-body video of the user or a facial video of the user. Multidimensional mapping parameters can include visual prediction indicators of the user's facial visual attributes corresponding to the candidate medical entity, physiological indicators of the user for the candidate medical entity, and / or physiological parameters of the user's physiological signals corresponding to the candidate medical entity. The visual prediction indicator can be the predicted probability or prediction score of the user belonging to the facial visual attribute. For example, if the candidate medical entity is the candidate disease - angina pectoris, the facial visual attribute corresponding to angina pectoris is the pain attribute, and the visual prediction indicator is the predicted probability of the user's face belonging to the pain attribute. The physiological indicators can be calculated based on the physiological parameters of the user's physiological signals corresponding to the candidate medical entity.
[0094] The following description uses the application of a video interaction processing method provided in this embodiment in a video interaction scenario as an example to further illustrate the video interaction processing method provided in this embodiment. (See also...) Figure 8 This is a video interaction processing method applied to video interaction scenarios. It can be applied to servers and specifically includes the following steps.
[0095] Step S810: Based on the user's voice data uploaded by the user terminal, perform disease identification on the user's text, obtain candidate diseases, and return them to the user terminal.
[0096] Step S818: Construct initial diagnosis and treatment labels based on physiological parameters, user text, and candidate diseases uploaded by the user terminal, and construct visual diagnosis and treatment labels based on visual prediction indicators and physiological parameters uploaded by the user terminal to obtain visual diagnosis and treatment labels.
[0097] Step S820: Return the initial diagnostic label and visual diagnostic label to the user terminal.
[0098] This specification provides an embodiment of a video interactive processing device as follows: In the above embodiments, a video interaction processing method is provided, and correspondingly, a video interaction processing device is also provided, which will be described below with reference to the accompanying drawings.
[0099] Reference Figure 10 The diagram illustrates an embodiment of a video interactive processing device provided in this embodiment.
[0100] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.
[0101] This embodiment provides a video interactive processing device, including: The switching detection module 1002 is configured to collect images of users who interact with the intelligent agent instance corresponding to the medical service provider via voice, and to detect the switching of interaction modes based on the collected action image sequence. The entity acquisition module 1004 is configured to upload user voice data to the server if the switch detection passes, and obtain candidate medical entities obtained by medical entity recognition based on user text returned by the server based on the user voice data. The interactive mapping module 1006 is configured to perform multi-dimensional interactive mapping on the candidate medical entity based on the user interactive video, and obtain the multi-dimensional mapping parameters of the candidate medical entity. The parameter upload module 1008 is configured to upload the multidimensional mapping parameters to the server and display the diagnosis and treatment data generated by the intelligent agent instance based on the multidimensional mapping parameters and the user text returned by the server.
[0102] Another embodiment of the video interactive processing device provided in this specification is as follows: In the above embodiments, another video interaction processing method is provided, and correspondingly, another video interaction processing device is also provided, which will be described below with reference to the accompanying drawings.
[0103] Reference Figure 11 The diagram illustrates an embodiment of a video interactive processing device provided in this embodiment.
[0104] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.
[0105] This embodiment provides a video interactive processing device, including: The entity recognition module 1102 is configured to perform medical entity recognition on user text based on user voice data uploaded by user terminal to obtain candidate medical entities; the user voice data is obtained by capturing images of users interacting with intelligent agent instances corresponding to medical service providers through voice, and uploading the data after passing the interaction mode switching detection based on the captured action image sequence. The entity return module 1104 is configured to return the candidate medical entity to the user terminal to obtain the multidimensional mapping parameters of the candidate medical entity by performing multidimensional interactive mapping on the candidate medical entity based on the user interactive video. The data return module 1106 is configured to call the intelligent agent instance to generate diagnosis and treatment data based on the multidimensional mapping parameters uploaded by the user terminal and the user text, and return the diagnosis and treatment data to the user terminal.
[0106] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0107] This specification provides an embodiment of a video interactive processing device as follows: Corresponding to the video interaction processing method described above, based on the same technical concept, one or more embodiments of this specification also provide a video interaction processing device for executing the video interaction processing method provided above. Figure 12 This is a schematic diagram of the structure of a video interactive processing device provided for one or more embodiments of this specification.
[0108] This embodiment provides a video interactive processing device, including: like Figure 12As shown, device 1200 mainly consists of a communication interface 1202, a user interface 1204, a processor 1206, and a data storage 1208. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 1210. The communication interface 1202 enables device 1200 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 1202 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 1202 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 1202 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 1202 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces. The user interface 1204 includes receiving user input and providing output to the user. Therefore, user interface 1204 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 1204 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 1204 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, device 1200 may support remote access from other devices via communication interface 1202 or another physical interface (not shown). User interface 1204 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 1204 may also be configured as a display device for rendering or displaying text fragments.
[0109] Processor 1206 may include one or more general-purpose processors and / or dedicated processors. Data storage 1208 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 1206. Data storage 1208 may include removable and non-removable components.
[0110] Processor 1206 is capable of executing program instructions 1218 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 1208 to perform the various functions described herein. Data storage 1208 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 1200, enable device 1200 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 1218 by processor 1206 may result in processor 1206 using data 1212. For example, program instructions 1218 may include an operating system 1222 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 1200 and one or more application programs 1220 (e.g., a browser, social application, or game application). Similarly, data 1212 may include operating system data 1216 and application data 1214. Operating system data 1216 is primarily accessible to operating system 1222, while application data 1214 is primarily accessible to one or more application programs 1220. Application data 1214 may reside in a file system visible or hidden from the user of device 1200. Application 1220 may communicate with operating system 1222 via one or more application programming interfaces (APIs). These APIs facilitate application 1220 reading and / or writing application data 1214, transmitting or receiving information via communication interface 1202, receiving or displaying information on user interface 1204, etc. In some terms, application 1220 may be simply referred to as an "app". Furthermore, application 1220 may be downloaded to device 1200 through one or more online app stores or app markets. However, applications may also be installed on device 1200 in other ways, such as through a web browser or a physical interface on device 1200 (e.g., a USB port).
[0111] In one specific embodiment, the video interaction processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the video interaction processing device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Image acquisition is performed on users who interact with the intelligent agent instance corresponding to the medical service provider via voice, and interaction mode switching is detected based on the acquired action image sequence. If the switching detection passes, the user's voice data is uploaded to the server, and the candidate medical entities obtained by medical entity recognition based on the user's text returned by the server based on the user's voice data are obtained. Based on user interaction video, multi-dimensional interaction mapping is performed on the candidate medical entities to obtain the multi-dimensional mapping parameters of the candidate medical entities; The multidimensional mapping parameters are uploaded to the server, and the diagnostic data generated by the intelligent agent instance based on the multidimensional mapping parameters and the user text, returned by the server, is displayed.
[0112] Another embodiment of the video interactive processing device provided in this specification is as follows: Corresponding to the other video interaction processing method described above, based on the same technical concept, one or more embodiments of this specification also provide another video interaction processing device, which is used to execute the other video interaction processing method provided above. Figure 13 This is a schematic diagram of the structure of a video interactive processing device provided for one or more embodiments of this specification.
[0113] This embodiment provides a video interactive processing device, including: like Figure 13As shown, device 1300 mainly consists of a communication interface 1302, a user interface 1304, a processor 1306, and a data storage 1308. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 1310. The communication interface 1302 enables device 1300 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 1302 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 1302 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 1302 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 1302 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces. The user interface 1304 includes receiving user input and providing output to the user. Therefore, user interface 1304 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 1304 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 1304 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, device 1300 may support remote access from other devices via communication interface 1302 or another physical interface (not shown). User interface 1304 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 1304 may also be configured as a display device for rendering or displaying text fragments.
[0114] Processor 1306 may include one or more general-purpose processors and / or special-purpose processors. Data storage 1308 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 1306. Data storage 1308 may include removable and non-removable components.
[0115] Processor 1306 is capable of executing program instructions 1318 (e.g., compiled or uncompiled program logic and / or machine code) stored in data store 1308 to perform the various functions described herein. Data store 1308 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 1300, enable device 1300 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 1318 by processor 1306 may result in processor 1306 using data 1312. For example, program instructions 1318 may include an operating system 1322 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 1300 and one or more application programs 1320 (e.g., a browser, social application, or game application). Similarly, data 1312 may include operating system data 1316 and application data 1314. Operating system data 1316 is primarily accessible to operating system 1322, while application data 1314 is primarily accessible to one or more application programs 1320. Application data 1314 may reside in a file system visible or hidden to the user of device 1300. Application 1320 may communicate with operating system 1322 via one or more application programming interfaces (APIs). These APIs facilitate application 1320 reading and / or writing application data 1314, transmitting or receiving information via communication interface 1302, receiving or displaying information on user interface 1304, etc. In some terms, application 1320 may be simply referred to as an "app". Furthermore, application 1320 may be downloaded to device 1300 through one or more online app stores or app markets. However, applications may also be installed on device 1300 in other ways, such as through a web browser or a physical interface on device 1300 (e.g., a USB port).
[0116] In one specific embodiment, the video interaction processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the video interaction processing device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Candidate medical entities are obtained by performing medical entity recognition on user text based on user voice data uploaded by user terminals; the user voice data is obtained by capturing images of users interacting with intelligent agent instances corresponding to medical service providers through voice, and uploading the data after passing interaction mode switching detection based on the captured action image sequence. The candidate medical entity is returned to the user terminal to obtain the multidimensional mapping parameters of the candidate medical entity by performing multidimensional interactive mapping on the candidate medical entity based on the user interactive video. The intelligent agent instance is invoked to generate diagnostic data based on the multidimensional mapping parameters uploaded by the user terminal and the user text, and the diagnostic data is returned to the user terminal.
[0117] This specification provides an embodiment of a computer-readable storage medium as follows: Corresponding to the video interaction processing method described above, and based on the same technical concept, one or more embodiments of this specification also provide a computer-readable storage medium.
[0118] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, perform the following steps: Image acquisition is performed on users who interact with the intelligent agent instance corresponding to the medical service provider via voice, and interaction mode switching is detected based on the acquired action image sequence. If the switching detection passes, the user's voice data is uploaded to the server, and the candidate medical entities obtained by medical entity recognition based on the user's text returned by the server based on the user's voice data are obtained. Based on user interaction video, multi-dimensional interaction mapping is performed on the candidate medical entities to obtain the multi-dimensional mapping parameters of the candidate medical entities; The multidimensional mapping parameters are uploaded to the server, and the diagnostic data generated by the intelligent agent instance based on the multidimensional mapping parameters and the user text, returned by the server, is displayed.
[0119] It should be noted that the embodiments of a computer-readable storage medium described in this specification and the embodiments of a video interactive processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0120] Another embodiment of a computer-readable storage medium provided in this specification is as follows: In response to another video interaction processing method described above, and based on the same technical concept, one or more embodiments of this specification also provide another computer-readable storage medium.
[0121] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, perform the following steps: Candidate medical entities are obtained by performing medical entity recognition on user text based on user voice data uploaded by user terminals; the user voice data is obtained by capturing images of users interacting with intelligent agent instances corresponding to medical service providers through voice, and uploading the data after passing interaction mode switching detection based on the captured action image sequence. The candidate medical entity is returned to the user terminal to obtain the multidimensional mapping parameters of the candidate medical entity by performing multidimensional interactive mapping on the candidate medical entity based on the user interactive video. The intelligent agent instance is invoked to generate diagnostic data based on the multidimensional mapping parameters uploaded by the user terminal and the user text, and the diagnostic data is returned to the user terminal.
[0122] It should be noted that the embodiments of another computer-readable storage medium described in this specification and the embodiments of another video interactive processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0123] This specification provides an example of a computer program product as follows: Corresponding to the video interaction processing method described above, based on the same technical concept, one or more embodiments of this specification also provide a computer program product.
[0124] A computer program product includes a computer program / instructions that, when executed by a processor, perform the following steps: Image acquisition is performed on users who interact with the intelligent agent instance corresponding to the medical service provider via voice, and interaction mode switching is detected based on the acquired action image sequence. If the switching detection passes, the user's voice data is uploaded to the server, and the candidate medical entities obtained by medical entity recognition based on the user's text returned by the server based on the user's voice data are obtained. Based on user interaction video, multi-dimensional interaction mapping is performed on the candidate medical entities to obtain the multi-dimensional mapping parameters of the candidate medical entities; The multidimensional mapping parameters are uploaded to the server, and the diagnostic data generated by the intelligent agent instance based on the multidimensional mapping parameters and the user text, returned by the server, is displayed.
[0125] It should be noted that the embodiments of a computer program product described in this specification and the embodiments of a video interaction processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0126] Another example of a computer program product provided in this specification is as follows: Corresponding to the other video interaction processing method described above, based on the same technical concept, one or more embodiments of this specification also provide another computer program product.
[0127] A computer program product includes a computer program / instructions that, when executed by a processor, perform the following steps: Candidate medical entities are obtained by performing medical entity recognition on user text based on user voice data uploaded by user terminals; the user voice data is obtained by capturing images of users interacting with intelligent agent instances corresponding to medical service providers through voice, and uploading the data after passing interaction mode switching detection based on the captured action image sequence. The candidate medical entity is returned to the user terminal to obtain the multidimensional mapping parameters of the candidate medical entity by performing multidimensional interactive mapping on the candidate medical entity based on the user interactive video. The intelligent agent instance is invoked to generate diagnostic data based on the multidimensional mapping parameters uploaded by the user terminal and the user text, and the diagnostic data is returned to the user terminal.
[0128] It should be noted that the embodiments of another computer program product described in this specification and the embodiments of another video interaction processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0129] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments. For example, the device embodiment, equipment embodiment and computer-readable storage medium embodiment are all similar to the method embodiment, so the description is relatively simple. When reading the relevant content of the device embodiment, equipment embodiment and computer-readable storage medium embodiment, please refer to the description of the method embodiment.
[0130] Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is only one of many possible execution orders and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims.
[0131] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.
[0132] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0133] In the 1930s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many improvements to the methodology today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that an improvement to the methodology cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0134] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0135] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0136] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0137] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0138] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable test processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable test processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0139] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable test processing equipment to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0140] These computer program instructions can also be loaded onto a computer or other programmable test processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0141] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0142] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0143] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0144] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of features includes not only those features but also other features not expressly listed, or features inherent to such process, method, article, or apparatus. Without further limitations, a feature defined by the phrase "comprising one..." does not exclude the presence of other identical features in the process, method, article, or apparatus that includes said feature.
[0145] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0146] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0147] The above description is merely an embodiment of this document and is not intended to limit the scope of this document. Various modifications and variations can be made to this document by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this document should be included within the scope of the claims of this document.
Claims
1. A video interactive processing method, comprising: Image acquisition is performed on users who interact with the intelligent agent instance corresponding to the medical service provider via voice, and interaction mode switching is detected based on the acquired action image sequence. If the switching detection passes, the user's voice data is uploaded to the server, and the candidate medical entities obtained by medical entity recognition based on the user's text returned by the server based on the user's voice data are obtained. Based on user interaction video, multi-dimensional interaction mapping is performed on the candidate medical entities to obtain the multi-dimensional mapping parameters of the candidate medical entities; The multidimensional mapping parameters are uploaded to the server, and the diagnostic data generated by the intelligent agent instance based on the multidimensional mapping parameters and the user text, returned by the server, is displayed.
2. The video interaction processing method according to claim 1, wherein the medical service provider includes medical personnel, and the interactive intelligent agent is created after the medical personnel authorize medical interaction; The interactive intelligent agent is obtained by constructing a large medical model, which is obtained by training a pre-trained medical model based on the medical personnel's diagnosis and treatment data.
3. The video interaction processing method according to claim 2, before executing the step of image acquisition of the user who performs voice interaction with the intelligent agent instance corresponding to the medical service provider, and the interaction mode switching detection step based on the acquired action image sequence, further includes: Based on the target agent tag selected by the user from multiple agent tags corresponding to medical personnel displayed in the application, the user is redirected from the application to the interactive subroutine. Based on the voice interaction command of the interactive intelligent agent corresponding to the target intelligent agent tag submitted by the user in the interaction subroutine, the user performs voice interaction with the intelligent agent instance.
4. The video interaction processing method according to claim 3, wherein the agent tag is obtained by agent matching based on the user's medical interaction intent; The medical interaction intent is obtained by recognizing the medical interaction voice entered by the user through the intelligent interaction control in the application; the intelligent interaction control is displayed after detecting that the user performs a page switching operation in the application.
5. The video interaction processing method according to claim 1, wherein the step of performing multi-dimensional interaction mapping on the candidate medical entity based on user interaction video to obtain the multi-dimensional mapping parameters of the candidate medical entity includes: The user interaction video and the facial visual attributes corresponding to the candidate medical entities are input into a visual model to predict visual indicators, thereby obtaining the user's visual prediction indicators for the facial visual attributes. Based on the user interaction video, determine the physiological parameters of the physiological signals of the user in relation to the candidate medical entity, and determine the physiological indicators of the user in relation to the candidate medical entity based on the physiological parameters.
6. The video interaction processing method according to claim 5, wherein the diagnostic data is generated in the following manner: The user fit of the candidate medical entity is calculated based on the visual prediction index, the physiological index, and the text index calculated based on the user text. The candidate medical entities are sorted according to the user suitability to obtain the pushed medical entities, and the intelligent agent instance is called to generate the diagnosis and treatment report of the pushed medical entities to obtain the diagnosis and treatment report as the diagnosis and treatment data.
7. The video interaction processing method according to claim 6, wherein generating the diagnosis and treatment report of the pushed medical entity includes: Based on the user's text, the text diagnosis data of the pushed medical entity is generated according to the medical style of the medical service provider. Based on the facial visual parameters of the user in relation to the pushed medical entity identified from the user interaction video, visual diagnostic data of the pushed medical entity is generated, and a diagnostic report is generated based on the text diagnostic data and the visual diagnostic data.
8. The video interaction processing method according to claim 5, wherein the diagnostic data is generated in the following manner: Initial diagnostic labels are constructed based on the user text and candidate diseases; Visual diagnostic labels are constructed based on the visual prediction indicators and the physiological parameters to obtain visual diagnostic labels, and the initial diagnostic labels and the visual diagnostic labels are used as the diagnostic data.
9. The video interaction processing method according to claim 8, wherein the step of constructing visual diagnosis and treatment labels based on the visual prediction index and the physiological parameters to obtain visual diagnosis and treatment labels includes: Based on the user's voice data, a fundamental frequency value sequence is obtained by calculating the fundamental frequency, and a psychological index is obtained by calculating the psychological index based on the fundamental frequency value sequence. The visual diagnostic labels are constructed based on the visual prediction indicators, the psychological indicators, and the physiological parameters.
10. The video interaction processing method according to claim 8, after executing the steps of uploading the multidimensional mapping parameters to the server and displaying the diagnostic data generated by calling the intelligent agent instance based on the multidimensional mapping parameters and the user text returned by the server, the method further includes: In response to the user's access command for either the initial treatment label or the visual treatment label, a treatment report corresponding to the initial treatment label and the visual treatment label is generated and displayed based on the visual prediction index and the physiological index.
11. The video interaction processing method according to claim 1, wherein the interactive intelligent agent is created after the medical service provider performs medical identity authentication in the interaction subroutine; The interactive agent is obtained by configuring parameters of the initial interactive agent based on the medical records of the medical service provider.
12. The video interaction processing method according to claim 1, wherein the step of detecting the interaction mode switch based on the acquired motion image sequence includes: The user's posture is obtained by performing user posture recognition based on device movement data; The interaction mode switching of the motion image sequence is detected according to the switching detection method corresponding to the user posture.
13. The video interaction processing method according to claim 12, wherein the step of performing interaction mode switching detection on the action image sequence according to the switching detection method corresponding to the user posture includes: If the user pose is the first user pose, at least two motion vectors are constructed based on the key point coordinates extracted from each motion image in the motion image sequence; Based on the at least two action vectors, the action amplitude is detected to obtain the user action amplitude corresponding to each action image. If the number of action images with user action amplitudes exceeding the preset action amplitude is greater than the number threshold, the interaction mode switching detection is determined to be successful.
14. The video interaction processing method according to claim 12, wherein the step of performing interaction mode switching detection on the action image sequence according to the switching detection method corresponding to the user posture includes: If the user pose is the second user pose, the positioning distance between the eye positioning points is calculated based on the positioning point coordinates of the eye positioning points in each action image contained in the action image sequence. Based on the positioning distance, it is determined whether the number of the user's eye movements exceeds the threshold. If so, the interaction mode switching detection is confirmed to be successful.
15. A video interactive processing method, comprising: Candidate medical entities are obtained by recognizing user text based on user voice data uploaded by user terminals. The user voice data is obtained by capturing images of users interacting with the intelligent agent instance corresponding to the medical service provider through voice interaction, and uploading the data after passing the interaction mode switching detection based on the captured action image sequence. The candidate medical entity is returned to the user terminal to obtain the multidimensional mapping parameters of the candidate medical entity by performing multidimensional interactive mapping on the candidate medical entity based on the user interactive video. The intelligent agent instance is invoked to generate diagnostic data based on the multidimensional mapping parameters uploaded by the user terminal and the user text, and the diagnostic data is returned to the user terminal.
16. A video interactive processing apparatus, comprising: The switching detection module is configured to collect images of users who interact with the intelligent agent instance corresponding to the medical service provider via voice, and to detect the switching of interaction modes based on the collected action image sequence. The entity acquisition module is configured to upload user voice data to the server if the switch detection passes, and obtain candidate medical entities obtained by medical entity recognition based on user text returned by the server based on the user voice data. The interactive mapping module is configured to perform multi-dimensional interactive mapping on the candidate medical entities based on user interactive videos, and obtain multi-dimensional mapping parameters of the candidate medical entities. The parameter upload module is configured to upload the multidimensional mapping parameters to the server and display the diagnostic data generated by the intelligent agent instance based on the multidimensional mapping parameters and the user text returned by the server.
17. A video interactive processing apparatus, comprising: The entity recognition module is configured to perform medical entity recognition based on user text from user voice data uploaded by the user terminal to obtain candidate medical entities. The user voice data is obtained by capturing images of users interacting with the intelligent agent instance corresponding to the medical service provider through voice interaction, and uploading the data after passing the interaction mode switching detection based on the captured action image sequence. The entity return module is configured to return the candidate medical entity to the user terminal in order to obtain the multidimensional mapping parameters of the candidate medical entity by performing multidimensional interactive mapping on the candidate medical entity based on the user interactive video. The data return module is configured to call the intelligent agent instance to generate diagnostic data based on the multidimensional mapping parameters uploaded by the user terminal and the user text, and return the diagnostic data to the user terminal.
18. A video interactive processing device, comprising: processor; And, a memory configured to store computer-executable instructions, which, when executed, cause the processor to: Image acquisition is performed on users who interact with the intelligent agent instance corresponding to the medical service provider via voice, and interaction mode switching is detected based on the acquired action image sequence. If the switching detection passes, the user's voice data is uploaded to the server, and the candidate medical entities obtained by medical entity recognition based on the user's text returned by the server based on the user's voice data are obtained. Based on user interaction video, multi-dimensional interaction mapping is performed on the candidate medical entities to obtain the multi-dimensional mapping parameters of the candidate medical entities; The multidimensional mapping parameters are uploaded to the server, and the diagnostic data generated by the intelligent agent instance based on the multidimensional mapping parameters and the user text, returned by the server, is displayed.
19. A video interactive processing device, comprising: processor; And, a memory configured to store computer-executable instructions, which, when executed, cause the processor to: Candidate medical entities are obtained by performing medical entity recognition on user text based on user voice data uploaded by user terminals; the user voice data is obtained by capturing images of users interacting with intelligent agent instances corresponding to medical service providers through voice, and uploading the data after passing interaction mode switching detection based on the captured action image sequence. The candidate medical entity is returned to the user terminal to obtain the multidimensional mapping parameters of the candidate medical entity by performing multidimensional interactive mapping on the candidate medical entity based on the user interactive video. The intelligent agent instance is invoked to generate diagnostic data based on the multidimensional mapping parameters uploaded by the user terminal and the user text, and the diagnostic data is returned to the user terminal.
20. A computer-readable storage medium for storing computer-executable instructions that, when executed, implement the steps of the method of claim 1 or 15.