Method and device for standardizing video inquiry diagnosis and treatment behaviors and electronic equipment

Through the client and cloud collaborative architecture, and by utilizing multimodal models and compliance evaluation models, we have achieved automated compliance verification and real-time monitoring of the video consultation process, solving the problems of traditional methods such as strong dependence on human resources, high operating costs, and poor timeliness, and achieving efficient diagnostic and treatment behavior norms.

CN120727239APending Publication Date: 2025-09-30WEDOC CLOUD (HANGZHOU) HLDG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510827107.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

The traditional standardized video consultation and treatment method relies on large human resource investment, high operating costs and poor timeliness, and cannot achieve real-time standardization and effective management of doctors.

Method used

Through the client and cloud collaborative architecture, a multimodal model is used to verify the compliance of doctor images before consultation, and multimodal real-time analysis of video streams, audio streams and text content is performed in real time to generate adjustment suggestions and early warning reminders, and generate standardized reports in combination with the compliance evaluation model.

Benefits of technology

It has achieved full automation of the consultation process, reduced operating costs, improved timeliness, and can standardize doctors' diagnosis and treatment behaviors in real time, reducing dependence on human resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120727239A_ABST
    Figure CN120727239A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for standardizing video inquiry diagnosis and treatment behaviors and electronic equipment. The system can automatically perform compliance verification on dressing, environment and sitting posture of a doctor before interrogation, perform multi-modal real-time analysis on real-time video stream, real-time audio stream and text content in the interrogation process, and perform compliance evaluation on logs, early warning records and multi-modal features before and in the interrogation process, namely, full-process automation is realized, and the efficiency is improved. Compared with the prior art, the method has the advantages that manpower resource investment is not needed, operation cost is greatly reduced, in addition, diagnosis and treatment behaviors of doctors in online video inquiry can be standardized in real time by adjusting suggestions and early warning reminding, and timeliness is good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical standards, and in particular to a method, device and electronic equipment for standardizing video consultation and treatment behaviors. Background Art

[0002] With the rapid development of internet technology, online medical treatment has become a relatively common option. Video consultations are the only online medical treatment method that allows both the doctor and the patient to see each other's real-time status in real time. However, while convenient, video consultations also introduce another challenge to the online medical experience: the difficulty in controlling the doctor's appearance and behavior during video consultations.

[0003] In order to improve patients' video consultation and medical experience, Internet diagnosis and treatment platforms often adopt methods such as concentrated pre-job training and publicity, manual sampling after diagnosis, and the establishment of a reward and punishment mechanism for doctors.

[0004] Existing solutions suffer from two obvious flaws. First, they rely heavily on human resources. Centralized training and promotion require specific time and venues, spot checks require dedicated personnel, and reward and punishment mechanisms rely heavily on manual review of each diagnosis and treatment. Second, they all suffer from issues of premature or delayed implementation. Premature promotion and promotion can lead to a long, uncertain period before doctors officially start seeing patients, resulting in limited effectiveness. Post-diagnosis spot checks and reward and punishment mechanisms can only constrain subsequent diagnosis and treatment.

[0005] In summary, the traditional method of standardizing video consultation and treatment behaviors has technical problems such as strong dependence on human resources, high operating costs, and poor timeliness. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a method, device and electronic equipment for standardizing video consultation and treatment behaviors, so as to alleviate the technical problems of traditional methods for standardizing video consultation and treatment behaviors, such as strong dependence on human resources, high operating costs and poor timeliness.

[0007] In a first aspect, an embodiment of the present invention provides a method for standardizing video consultation and treatment behaviors, including:

[0008] The client's multimodal model is used to perform compliance verification on the doctor's clothing, environment, and sitting posture in the pre-consultation doctor image. If the compliance verification fails, adjustment suggestions are generated and displayed based on the compliance verification results.

[0009] Through the collaborative architecture of the client and the cloud, a multimodal real-time analysis of the real-time video stream, real-time audio stream, and text content of the consultation process is performed to obtain a real-time compliance score. When the real-time compliance score is less than a preset score threshold, an early warning reminder is issued through the client;

[0010] After the consultation is completed, the logs, warning records and multimodal features before and during the consultation are input into the compliance evaluation model to obtain the total compliance score, pre-consultation score, consultation process score, communication score and improvement suggestions, and then generate a consultation standard report.

[0011] Furthermore, the client's multimodal model verifies the compliance of the doctor's attire, environment, and sitting posture in the pre-consultation images, including:

[0012] The client obtains the doctor's image before the consultation;

[0013] Processing the doctor's image before the consultation using the multimodal model to obtain the doctor's clothing information, environment information, and sitting posture information;

[0014] The doctor's clothing information, environment information and sitting posture information are respectively checked for compliance according to preset clothing standards, preset environment standards and preset sitting posture standards.

[0015] Furthermore, when the compliance check fails, after generating and displaying adjustment suggestions based on the compliance check result, the method further includes:

[0016] Return to the step of the client obtaining the doctor's image before the consultation until the compliance check is passed.

[0017] Furthermore, a multimodal real-time analysis of the real-time video stream, real-time audio stream, and text content of the consultation process is performed through the collaborative architecture of the client and the cloud, including:

[0018] The client obtains the real-time video stream, real-time audio stream and text content of the consultation process;

[0019] Preprocessing the real-time video stream and the real-time audio stream to obtain key frames in the real-time video stream and audio segments corresponding to the key frames in the real-time audio stream;

[0020] Sending the key frame, the audio segment, and the text content to the cloud;

[0021] The cloud-based visual model detects the key frames to obtain posture features, expression features, and behavior features;

[0022] The cloud performs speech recognition and semantic understanding on the audio segment to obtain speech features;

[0023] The large language model in the cloud processes the text content to obtain the diagnosis and treatment speech features;

[0024] The posture features, the expression features, the behavior features, the voice features and the diagnosis and treatment speech features are fused, and a compliance score is calculated for the fused features to obtain the real-time compliance score.

[0025] Furthermore, when the real-time compliance score is less than a preset score threshold, an early warning reminder is issued through the client, including:

[0026] When the real-time compliance score is less than the preset score threshold, the cloud generates warning information and sends the warning information to the client, so that the client issues a warning reminder based on the warning information, wherein the warning information includes: the real-time compliance score, warning level and violation behavior.

[0027] Furthermore, the logs, warning records and multimodal features before and during the consultation are input into the compliance evaluation model, including:

[0028] generating a time series feature according to the log, the warning record, and the multimodal feature;

[0029] Performing Transformer encoding on the temporal features to obtain a global context representation;

[0030] Constructing a medical inquiry knowledge graph based on the global context representation, wherein the nodes in the medical inquiry knowledge graph include: doctors, patients, and environments, and the edges in the medical inquiry knowledge graph have weights, and the weights are interaction activity;

[0031] Using a graph neural network to process the medical consultation knowledge graph, the total compliance score, the pre-consultation score, the medical consultation process score, the communication score, and the improvement suggestions are obtained;

[0032] The consultation standard report is generated based on the time series characteristics, the total compliance score, the pre-consultation score, the consultation process score, the communication score and the improvement suggestions.

[0033] Furthermore, the multimodal model includes: a distilled and pruned version of ViT-Base, a micro-branch of HRNet-W32, and a multi-task output head, and the multimodal model is a model quantized to INT8;

[0034] Data transmission between the client and the cloud is encrypted;

[0035] Access to medical consultation data is controlled and sensitive information is masked.

[0036] In a second aspect, an embodiment of the present invention further provides a device for standardizing video consultation and treatment behaviors, including:

[0037] A compliance verification unit, configured to verify the compliance of the doctor's attire, environment, and sitting posture on the pre-consultation doctor's image using the client's multimodal model, and to generate and display adjustment suggestions based on the compliance verification results if the compliance verification fails;

[0038] A multimodal real-time analysis unit is used to perform multimodal real-time analysis of the real-time video stream, real-time audio stream, and text content of the consultation process through a collaborative architecture between the client and the cloud to obtain a real-time compliance score, and to issue an early warning reminder through the client when the real-time compliance score is less than a preset score threshold;

[0039] The compliance evaluation unit is used to input the logs, warning records and multimodal features before and during the consultation into the compliance evaluation model after the consultation is completed, obtain the total compliance score, pre-consultation score, consultation process score, communication score and improvement suggestions, and then generate a consultation standard report.

[0040] In a third aspect, an embodiment of the present invention further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods described in the first aspect when executing the computer program.

[0041] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to execute any method described in the first aspect above.

[0042] In an embodiment of the present invention, a method for standardizing video consultation and treatment behavior is provided, including: using a multimodal model of a client to perform compliance verification on the doctor's dress, environment and sitting posture of the doctor's image before the consultation, and when the compliance verification fails, generating and displaying adjustment suggestions based on the compliance verification results; using a collaborative architecture of the client and the cloud to perform multimodal real-time analysis on the real-time video stream, real-time audio stream and text content of the consultation process to obtain a real-time compliance score, and when the real-time compliance score is less than a preset score threshold, issuing an early warning reminder through the client; after the consultation is completed, the logs, early warning records and multimodal features before and during the consultation are input into the compliance evaluation model to obtain the total compliance score, pre-consultation score, consultation process score, communication score and improvement suggestions, and then generate a consultation standardization report. From the above description, it can be seen that the method of standardizing video consultation and treatment behaviors of the present invention can automatically perform compliance verification on the doctor's attire, environment and sitting posture before the consultation, perform multimodal real-time analysis on the real-time video stream, real-time audio stream and text content of the consultation process, and conduct compliance evaluation on the logs, warning records and multimodal features before and during the consultation. That is, the entire process is automated, without relying on human resource investment, which greatly reduces operating costs. In addition, the doctor's diagnosis and treatment behaviors in online video consultations can be standardized in real time through adjustment suggestions and warning reminders, with good timeliness, alleviating the technical problems of traditional methods of standardizing video consultation and treatment behaviors, such as strong dependence on human resources, high operating costs and poor timeliness. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0044] Figure 1 A flowchart of a method for standardizing video consultation and treatment behavior provided by an embodiment of the present invention;

[0045] Figure 2 A flowchart of another method for standardizing video consultation and treatment behavior provided by an embodiment of the present invention;

[0046] Figure 3 A schematic diagram of a client pushing a video data packet to the cloud according to an embodiment of the present invention;

[0047] Figure 4 A schematic diagram of the metadata field in a data packet used by a client to push a video to the cloud according to an embodiment of the present invention;

[0048] Figure 5A schematic diagram of a client pushing a data packet of an audio segment to the cloud according to an embodiment of the present invention;

[0049] Figure 6 A schematic diagram of the request body field in a data packet used by a client to push an audio segment to the cloud according to an embodiment of the present invention;

[0050] Figure 7 A schematic diagram illustrating fields of a warning message provided in an embodiment of the present invention;

[0051] Figure 8 A schematic diagram of the overall system architecture and functional modules provided by an embodiment of the present invention;

[0052] Figure 9 A schematic diagram of a device for standardizing video consultation and treatment behavior provided by an embodiment of the present invention;

[0053] Figure 10 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0054] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0055] The traditional method of standardizing video consultation and treatment behaviors requires centralized pre-job training, manual spot checks and manual review after diagnosis, which is highly dependent on human resources and has high operating costs. In addition, the training effect lags, and post-diagnosis spot checks cannot correct diagnosis and treatment behaviors in real time, resulting in poor timeliness.

[0056] Based on this, the method of standardizing video consultation and treatment behaviors of the present invention can automatically perform compliance verification on the doctor's attire, environment and sitting posture before the consultation, perform multimodal real-time analysis on the real-time video stream, real-time audio stream and text content of the consultation process, and conduct compliance evaluation on the logs, warning records and multimodal features before and during the consultation. That is, the entire process is automated, without relying on human resource investment, which greatly reduces operating costs. In addition, the doctor's diagnosis and treatment behaviors in online video consultations can be standardized in real time through adjustment suggestions and warning reminders, with good timeliness.

[0057] To facilitate understanding of this embodiment, a method for standardizing video consultation and treatment behaviors disclosed in an embodiment of the present invention is first introduced in detail.

[0058] Example 1:

[0059] According to an embodiment of the present invention, an embodiment of a method for standardizing video consultation and treatment behavior is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0060] Figure 1 FIG. 1 is a flow chart of a method for standardizing video consultation and treatment behaviors according to an embodiment of the present invention. Figure 1 As shown, the method includes the following steps:

[0061] Step S102: Using the client's multimodal model, the pre-consultation doctor's image is subjected to a compliance check of the doctor's attire, surroundings, and sitting posture. If the compliance check fails, adjustment suggestions are generated and displayed based on the compliance check results.

[0062] In an embodiment of the present invention, video consultation is an online diagnosis and treatment method in which a doctor uses Internet technology to conduct a video conversation and communication with a patient remotely on a terminal (such as a client).

[0063] Specifically, the above process is actually a pre-consultation posture check. Its goal is to ensure that the doctor's attire, environment, sitting posture, camera angle, and other aspects meet standard requirements before initiating a remote video consultation. Specifically, the camera captures the doctor's pre-consultation image (i.e., the doctor's current environment and appearance) and feeds it into a multimodal large model (specifically, a multimodal model). The model automatically compares the image against a set of rules, including pre-set dress, environment, and sitting posture standards.

[0064] If any of the doctor's attire, environment, and sitting posture fail to meet the standards, the compliance check is determined to have failed. Based on the compliance check results (specific unqualified items), the client prompts and gives adjustment suggestions, such as "Please look up, please turn the collar of your white coat from the inside out to the normal position, adjust the camera angle, and adjust the camera exposure." The above violation types and timestamps are recorded in the log before the consultation.

[0065] Step S104: Perform multimodal real-time analysis of the real-time video stream, real-time audio stream, and text content of the consultation process through a collaborative architecture between the client and the cloud to obtain a real-time compliance score. When the real-time compliance score is less than a preset score threshold, an early warning reminder is issued through the client.

[0066] Specifically, this process is actually the monitoring of real-time diagnosis and treatment behaviors. The goal is to capture and analyze the doctor's audio and video streams in real time during the consultation process to ensure that the diagnosis and treatment language, demeanor and posture continue to comply with medical standards. The specific implementation process is that the cloud continuously calls the multimodal analysis engine to synchronously process the client's video stream and audio stream. Once behaviors such as off-camera, non-professional language, and interruption of patient communication are detected, that is, the real-time compliance score is less than the preset score threshold, an immediate prompt or early warning reminder will be given through the client. The multimodal technology of the present invention refers to integrating data from different modalities (such as images, text, audio, video, etc.) to enhance the model's understanding and reasoning capabilities.

[0067] Step S106: After the consultation is completed, the logs, warning records and multimodal features before and during the consultation are input into the compliance evaluation model to obtain the total compliance score, pre-consultation score, consultation process score, communication score and improvement suggestions, and then generate a consultation standard report.

[0068] Specifically, this process is actually a post-consultation comprehensive evaluation process. The goal is to quantitatively score the overall compliance of the video consultation and treatment process after the video consultation is completed, and provide improvement suggestions. The specific implementation process is to input the logs, warning records, and multimodal feature vectors of the pre-consultation (i.e. before the consultation) and the consultation process (i.e. real-time monitoring of each stage) into the compliance evaluation model, output the total compliance score, pre-consultation score, consultation process score, communication score, and improvement suggestions, and then generate a personalized consultation standard report.

[0069] In an embodiment of the present invention, a method for standardizing video consultation and treatment behavior is provided, including: using a multimodal model of a client to perform compliance verification on the doctor's dress, environment and sitting posture of the doctor's image before the consultation, and when the compliance verification fails, generating and displaying adjustment suggestions based on the compliance verification results; using a collaborative architecture of the client and the cloud to perform multimodal real-time analysis on the real-time video stream, real-time audio stream and text content of the consultation process to obtain a real-time compliance score, and when the real-time compliance score is less than a preset score threshold, issuing an early warning reminder through the client; after the consultation is completed, the logs, early warning records and multimodal features before and during the consultation are input into the compliance evaluation model to obtain the total compliance score, pre-consultation score, consultation process score, communication score and improvement suggestions, and then generate a consultation standardization report. From the above description, it can be seen that the method of standardizing video consultation and treatment behaviors of the present invention can automatically perform compliance verification on the doctor's attire, environment and sitting posture before the consultation, perform multimodal real-time analysis on the real-time video stream, real-time audio stream and text content of the consultation process, and conduct compliance evaluation on the logs, warning records and multimodal features before and during the consultation. That is, the entire process is automated, without relying on human resource investment, which greatly reduces operating costs. In addition, the doctor's diagnosis and treatment behaviors in online video consultations can be standardized in real time through adjustment suggestions and warning reminders, with good timeliness, alleviating the technical problems of traditional methods of standardizing video consultation and treatment behaviors, such as strong dependence on human resources, high operating costs and poor timeliness.

[0070] The above content briefly introduces the method of standardizing video consultation and treatment behaviors of the present invention. The specific contents involved are described in detail below.

[0071] In an optional embodiment of the present invention, a client-side multimodal model is used to perform compliance verification on the doctor's clothing, environment, and sitting posture in the pre-consultation doctor image, specifically including the following steps:

[0072] (1) The client obtains the doctor's image before the consultation;

[0073] Specifically, the above-mentioned doctor images before the consultation include: the image of the doctor in a white coat at the entrance inspection area, the image of the doctor sitting in front of the client, etc. The entrance inspection refers to the abbreviation of the hospital's "entrance and exit control area".

[0074] (2) A multimodal model is used to process the doctor's image before the consultation to obtain the doctor's clothing information, environment information, and sitting posture information;

[0075] Specifically, the multimodal model consists of a distilled and pruned version of ViT-Base, a micro-branch of HRNet-W32, and a multi-task output head. The initial base model of the multimodal model includes a visual backbone (specifically, a pre-trained Vision Transformer (ViT-Base)) and a pose estimation model (specifically, a lightweight HRNet-W32). These base models are then distilled and pruned using a "teacher-student" framework, where the teacher is a complete joint ViT and HRNet network, and the student retains only four Transformer encoder layers and a micro-branch of HRNet. The pruning strategy involves pruning the number of heads in the student's Multi-Head Attention layer by 30%, and sparsifying and pruning the FFN linear layer, retaining 50% of the channels. Finally, the resulting model is quantized to INT8, resulting in a multimodal model with an inference time of ≤80ms for a single frame and a memory footprint of <150MB.

[0076] The above-mentioned ViT-Base and HRNet-W32 are executed in parallel (the original input image is processed independently, and then the global features of ViT are spliced ​​with the local features of HRNet at a specific stage. Feature fusion can be selected in the fourth stage of HRNet (when outputting a 14×14 resolution feature map)). The global features extracted by ViT-Base serve as the supplementary input of HRNet. ViT can provide scene context (such as the white coat area), and HRNet can provide accurate limb position information. The present invention has created a hybrid architecture of Transformer and HRNet. The global attention mechanism of ViT can make up for the local perception limitations of HRNet, and the high-resolution retention capability of HRNet can improve the accuracy of key point detection, breaking through the performance bottleneck of traditional single-model solutions. In addition, targeted layer retention is performed on ViT (4-layer Encoder), and a custom design is also made for the HRNet micro-branch. The total number of parameters of the two models is less than 100M, and the memory usage after INT8 quantization is less than 150MB, which can meet the real-time requirements of medical scenarios. Compared with a single model solution, the HRNet of the present invention improves the key point detection accuracy by 12% and reduces the false alarm rate by 25%. In addition, it is 3 times faster than the original ViT+HRNet combination and reduces memory usage by 60%.

[0077] Take white coat compliance testing as an example:

[0078] The function of ViT-Base is to identify clothing categories in the picture; locate possible white coat areas; and output semantic segmentation masks.

[0079] HRNet-W32 functions: accurately locate key parts such as collars and cuffs; measure the deviation between actual sleeve length and standard value; and calculate clothing coverage.

[0080] The following is a detailed introduction to the distilled and pruned versions of ViT-Base and HRNet-W32's micro-branches:

[0081] The visual backbone retains the first four Transformer encoder layers of the original ViT-Base; the last eight encoder layers and the classification head are removed. Each Transformer encoder layer includes a multi-head attention mechanism (retaining 70% of the heads after pruning), a feed-forward network (FFN; retaining 50% of the channels after pruning), layer normalization (LayerNorm), and residual connections. When pruning attention heads, the 12 attention heads in each Transformer layer are ranked by importance, retaining the eight most relevant to medical posture recognition (a pruning rate of 33%). L1-norm is used as the head importance metric.

[0082] Pose estimation branch: A micro-variant of HRNet-W32 is used; 4 stages are retained (from the original version); the number of channels in each stage is reduced to [32, 64, 128, 256] (from [32, 64, 128, 256, 512] in the original version); the last upsampling layer is removed. The number of residual blocks in each of the four stages is halved (from 4 to 2), the fifth high-resolution stage is removed, and the output feature map resolution remains at 1 / 4 the input size.

[0083] Through the above strategy, the student model (i.e., multimodal model) design significantly improves the efficiency of the model while ensuring the accuracy of medical behavior analysis, enabling it to run in real time (>15FPS) on client devices while meeting the strict requirements of medical scenarios for algorithm reliability.

[0084] In practical applications, such as clothing detection, ViT and HRNet's fused features are used; posture (i.e., sitting position) keypoint detection uses HRNet's independent features; and for environment classification, ViT's independent features (i.e., pose estimation) completely rely on HRNet's high-resolution features. Environment classification leverages ViT's global context understanding, while clothing detection combines the strengths of both features. The feature maps generated by ViT and HRNet are shared by multiple task heads, thereby obtaining different task information (such as clothing information, environment information, and sitting position information).

[0085] The above-mentioned INT8 quantization solution significantly improves the feasibility of client deployment through sophisticated layering strategies and medical scenario adaptation, while ensuring the accuracy of behavioral analysis. It is the key technical support for the real-time operation of the end-cloud collaborative architecture.

[0086] The doctor's clothing information includes: the area of ​​the white coat area occupied by the screen and the length of the sleeves; the environmental information includes: the certificates on the wall, the medical records on the table (used to determine whether there are sensitive signs in the background), the average brightness of the face area, and the distribution of environmental noise; the sitting posture information includes: the angle between the head and the torso, and the angle between the shoulder and the horizontal line of the screen.

[0087] (3) Compliance verification is performed on the doctor's clothing information, environment information, and sitting posture information according to the preset clothing standards, preset environment standards, and preset sitting posture standards.

[0088] Specifically, the preset dress code may be that the white coat area occupies more than 40% of the screen area; the sleeve length is above the elbow;

[0089] The preset environmental standards can be that there are no sensitive signs in the background (such as no certificates on the wall, no medical records on the desk); the average brightness of the face area is within a preset range (such as a brightness value of 50-200); and the ambient noise decibel is less than 45dB;

[0090] The preset sitting posture standard can be that the angle between the head and the torso ranges from 70° to 110°; the shoulders are parallel to the horizontal line of the screen with a deviation of ≤10°.

[0091] It should be noted that the above standards (or rules) support online incremental updates. In addition, the above models also support retraining, so they can adapt to the rapid iteration of regulations and hospital standards.

[0092] In an optional embodiment of the present invention, when the compliance check fails, after generating and displaying an adjustment suggestion based on the compliance check result, the method further includes:

[0093] Return to the steps of obtaining the doctor's image before the consultation until the compliance check is passed. Figure 2 shown.

[0094] In an optional embodiment of the present invention, a multimodal real-time analysis of the real-time video stream, real-time audio stream, and text content of the consultation process is performed through a collaborative architecture between the client and the cloud, specifically including the following steps:

[0095] (1) The client obtains the real-time video stream, real-time audio stream and text content of the consultation process;

[0096] Specifically, a 30fps, 640×480 resolution video stream is captured through an RGB camera, and a 16kHz single-channel audio stream is collected through a directional microphone. The above text content can be the medical records uploaded by the doctor during the consultation. After obtaining the real-time video stream and real-time audio stream, the PTP protocol can be used to ensure audio and video synchronization (error <10ms).

[0097] (2) preprocessing the real-time video stream and the real-time audio stream to obtain key frames in the real-time video stream and audio segments corresponding to the key frames in the real-time audio stream;

[0098] Specifically, after obtaining the real-time video stream and real-time audio stream after audio and video synchronization, the two are preprocessed, including: performing MTCNN face tracking on the real-time video stream to crop the region of interest; performing Conv-TasNet sound source separation on the real-time audio stream (referring to the process of recovering the original independent sound source from the mixed signal. In a real environment, the audio signal received by the microphone often contains sounds from multiple sound sources (such as human voice, background music, environmental noise, etc.). The goal of sound source separation technology is to separate these mixed sounds and extract each individual sound source) to obtain the human voice and enhance the human voice. Then, extract the key frame from the above-mentioned cropped region of interest (2 frames per second) and obtain the audio segment corresponding to the key frame (0.5 seconds).

[0099] (3) Send key frames, audio segments, and text content to the cloud;

[0100] (4) The cloud-based visual model detects key frames and obtains posture features, expression features, and behavior features;

[0101] Specifically, the visual model (such as ViT-Large) outputs posture features, expression features, and behavior features.

[0102] (5) The cloud performs speech recognition and semantic understanding on the audio segment to obtain speech features;

[0103] Specifically, speech recognition can be Speech-to-Text, and semantic understanding can be implemented using BERT to obtain speech features.

[0104] (6) A large language model in the cloud processes the text content to obtain the characteristics of the diagnosis and treatment speech;

[0105] Specifically, LLM (large language model, such as GPT-3.5) extracts the characteristics of diagnostic and treatment language.

[0106] (7) The posture features, expression features, behavior features, voice features and diagnosis and treatment speech features are integrated, and the compliance score of the integrated features is calculated to obtain a real-time compliance score.

[0107] Specifically, the above fusion can be attention fusion (i.e., cross-modal attention fusion mechanism), that is, weighted fusion of different features, and then compliance score calculation of the fused features. Specifically, the fused features can be summarized into a score of 0-100 through a fully connected layer to obtain a real-time compliance score.

[0108] The real-time compliance score here specifically refers to the multimodal real-time analysis of the audio and video of a short period of time during the consultation process (such as 1 second, because 2 key frames and corresponding audio are extracted in 1 second), and the real-time compliance score of the consultation process in this short period of time is obtained.

[0109] In an optional embodiment of the present invention, when the real-time compliance score is less than a preset score threshold, an early warning reminder is issued through the client, including:

[0110] When the real-time compliance score is less than the preset score threshold, the cloud generates an early warning message and sends it to the client so that the client can issue an early warning reminder based on the early warning information. The early warning information includes: real-time compliance score, warning level and violation behavior.

[0111] Specifically, the preset score threshold can be 75 points. The cloud-generated warning information includes the real-time compliance score, warning level, and violation behavior. For example, Warning Message 1 indicates a compliance score of 73 points, indicating a minor violation and requiring users to maintain a professional attitude. Warning Message 2 indicates a compliance score of 65 points, indicating a moderate violation, with three instances of non-professional language and an irregular sitting posture detected. Warning Message 3 indicates a compliance score of 52 points, indicating a severe violation, with off-camera behavior and an interruption of interaction exceeding 10 seconds.

[0112] When the client issues an early warning reminder based on the early warning information, it can remind the doctor through a floating layer and / or TTS voice and / or vibration, so that the doctor can make timely adjustments based on the early warning information.

[0113] It should be noted that the cloud sends warning information based on WebRTC DataChannel, and the round-trip delay is ≤100ms.

[0114] The above warning information is obtained by analyzing the fused features. If the fused feature analysis shows that there are no faces for two consecutive seconds, it is determined to be a violation of off-camera. If the fused feature analysis shows that there are X uses of words from the "non-professional terminology dictionary", such as "oh yo" and "nonsense", it is determined to be a violation of non-professional terminology. If the fused feature analysis shows that the doctor's attitude pauses for more than 5 seconds after the patient asks a question, it is determined to be a violation of interaction interruption. Of course, these rules support online incremental updates.

[0115] In an optional embodiment of the present invention, the logs before and during the consultation, early warning records, and multimodal features are input into the compliance evaluation model (all the following processes are the processing process of the compliance evaluation model), which specifically includes the following steps:

[0116] (1) Generate time series features based on logs, warning records, and multimodal features;

[0117] Specifically, the above-mentioned time series features can be a demeanor score curve, a speech compliance curve, a warning number sequence, etc.

[0118] (2) Transformer encoding of temporal features to obtain global context representation;

[0119] Specifically, the feature vector per second (that is, time series features, a set of feature vectors per second) is input into a 10-layer encoder, and the output is a global context representation (the global context representation is a comprehensive representation vector obtained by deep encoding the time series features through Transformer, and its characteristics include: time independence: compressing the information of the entire time series; multimodal fusion: including cross-modal interaction features such as vision and speech; semantic integrity: can be decoded into medical behavior evaluations of different granularities).

[0120] (3) Construct a medical consultation knowledge graph based on the global context representation, where the nodes in the medical consultation knowledge graph include: doctors, patients, and environments, and the edges in the medical consultation knowledge graph have weights, which are the interaction activity;

[0121] Specifically, the global context representation contains the features of doctor nodes, patient nodes, and environment nodes, and the medical consultation knowledge graph is constructed based on the global context representation. When constructing the medical consultation knowledge graph, weighted edges are constructed based on the interaction activity between nodes.

[0122] The global representation method fundamentally solves the three major challenges in medical behavior analysis: "long time series, multiple modalities, and strict standards." Its value lies not only in its technical implementation, but also in providing a core carrier that is computable, traceable, and improvable for digital medical quality control.

[0123] (4) Using a graph neural network to process the consultation knowledge graph, we can obtain the total compliance score, pre-consultation score, consultation process score, communication score, and improvement suggestions;

[0124] Specifically, the graph neural network can be GraphSAGE. The features of GraphSAGE include: neighbor sampling: that is, constructing a computational subgraph through random walks; feature aggregation: hierarchical aggregation of neighbor node information; parameter sharing: all nodes in the same layer share the aggregation function.

[0125] The above consultation process score refers to the score of the entire consultation process and needs to be distinguished from the above real-time compliance score. The above improvement suggestions are based on the copy automatically pulled from the rule library for low-scoring items.

[0126] (5) Generate a consultation standard report based on time series characteristics, total compliance score, pre-consultation score, consultation process score, communication score and improvement suggestions.

[0127] Specifically, a standardized medical consultation report is generated based on the Jinja2 template engine, which can visually display line charts (posture score curves), bar charts (module scores), and screenshots of violation fragments. The exportable formats include: HTML and PDF, which can be achieved through a one-click download link.

[0128] As shown below, the obtained consultation standard report (ID: CN20231125-028) can be:

[0129] Total compliance score: 84 / 100, ★★★★;

[0130] Pre-diagnosis score: 88 / 100;

[0131] Questioning process score: 82 / 100;

[0132] Communication score: 79 / 100;

[0133] Key findings:

[0134] There are two violations in the communication process:

[0135] 00:03:12 The risks of medication were not fully disclosed;

[0136] 00:07:45 Medical terminology is underused.

[0137] Process monitoring found that:

[0138] Improper hand disinfection procedures (00:01:30);

[0139] Electronic medical record recording is delayed (timeout 3 minutes).

[0140] Improvement Suggestions

[0141] Immediate measures: Participate in special training on communication skills; review the "Informed Consent Guidelines".

[0142] Long-term plan: Participate in simulated consultation drills every month; subscribe to the "Doctor-Patient Communication" journal for study.

[0143] This method transforms deep representations (i.e., global contextual representations) into a medical inquiry knowledge graph and uses graph neural networks to capture complex patterns of interaction between doctors, patients, and the environment. The resulting compliance decisions incorporate both macroeconomic assessments and specific improvement recommendations. All outputs undergo medical ethics review to ensure compliance with the latest clinical guidelines. Deep integration with hospital HIS systems is also supported, enabling a closed-loop quality management system of "evaluation-feedback-improvement."

[0144] In an optional embodiment of the present invention, data transmission between the client and the cloud is encrypted; access to the medical consultation data is subject to permission control, and sensitive information is masked.

[0145] The interface protocol and security and privacy of the present invention are described in detail below:

[0146] Interface protocol: POST / api / v1 / stream / frame

[0147] Function: The client pushes a frame of video and its context information to the cloud, such as Figure 3 As shown, the metadata field description is as follows Figure 4 The resolution field is described as follows:

[0148]

[0149] POST / api / v1 / stream / audio:

[0150] Function: The client pushes an audio segment to the cloud for speech recognition and semantic understanding, such as Figure 5 As shown, the request body fields are described as follows Figure 6 shown.

[0151] WS / api / v1 / alert:

[0152] Function: The cloud pushes warning information to the client.

[0153] Connection description: The client uses the same JWT at wss: / / <host> / api / v1 / alert establishes a WebSocket connection. Once the connection is successful, the cloud can push JSON format warning packages at any time. Warning message field description is as follows Figure 7 shown.

[0154] Data security and privacy protection: Transmission encryption: TLS1.3, full-link encryption; Storage encryption: AES-256-GCM, segmented encryption of logs and multimodal feature vectors; Access control: Based on OAuth2.0 Scope authorization, doctors can only access their own consultation data; Privacy shielding: Sensitive information (medical record number, original face image) in the report is coded or hashed.

[0155] The overall system architecture and functional modules corresponding to the method of the present invention are as follows: Figure 8 As shown in the figure, this method uses multimodal information fusion such as vision, voice, and text to automatically verify, monitor in real time, and conduct comprehensive post-evaluation of doctors' demeanor and treatment behaviors throughout the entire remote video consultation process, reducing manual review costs and improving compliance and user experience.

[0156] The implementation points and advantages of the present invention are:

[0157] 1. Client-side lightweighting and distillation pruning: ensuring real-time pre-inspection (before consultation) (<100ms) with less than 2% loss in model accuracy;

[0158] 2. Hybrid Parallel Multimodal Fusion: The cross-modal attention mechanism is introduced for the first time in remote medical compliance monitoring, improving monitoring accuracy by 15%;

[0159] 3. Unique "lightweight client-side and powerful cloud" collaborative architecture;

[0160] 4. Graph Neural Network Closed-Loop Evaluation: GNN models the three-dimensional relationship between doctors, patients, and the environment to achieve personalized and explainable compliance scoring;

[0161] 5. Dynamic self-learning rule base: supports online incremental updates and model retraining, adapting to the rapid iteration of regulations and in-hospital standards;

[0162] 6. Full-link security and auditability: Complies with HIPAA (or equivalent regulations) requirements to ensure medical privacy and data security.

[0163] Innovations include:

[0164] Innovation in technical fields: For the first time, multimodal large model technology was applied to the field of telemedicine behavioral norms; a new "pre-examination-real-time-post-event" full-process automated monitoring model was created; technical solution innovation: the cross-modal attention fusion mechanism improved monitoring accuracy by 15%; graph neural networks achieved explainable compliance scoring; the full-link security design complies with medical privacy regulations such as HIPAA; technical effect innovation: the cost of manual review was reduced by more than 80%; the real-time identification of violations was improved to 100ms; and it supports the rapid iteration and update of regulations and hospital standards.

[0165] Example 2:

[0166] An embodiment of the present invention also provides a device for standardizing video consultation and treatment behaviors. The device for standardizing video consultation and treatment behaviors is mainly used to execute the method for standardizing video consultation and treatment behaviors provided in Example 1 of the present invention. The following is a detailed introduction to the device for standardizing video consultation and treatment behaviors provided in the embodiment of the present invention.

[0167] Figure 9 is a schematic diagram of a device for standardizing video consultation and treatment behavior according to an embodiment of the present invention, such as Figure 9 As shown, the device mainly includes: a compliance verification unit 10, a multimodal real-time analysis unit 20, and a compliance evaluation unit 30, wherein:

[0168] The compliance verification unit is used to verify the compliance of the doctor's attire, environment, and sitting posture in the pre-consultation doctor's image using the client's multimodal model. If the compliance verification fails, adjustment suggestions are generated and displayed based on the compliance verification results.

[0169] A multimodal real-time analysis unit is used to perform multimodal real-time analysis of the real-time video stream, real-time audio stream, and text content of the consultation process through a collaborative architecture between the client and the cloud to obtain a real-time compliance score. When the real-time compliance score is lower than the preset score threshold, an early warning reminder is issued through the client;

[0170] The compliance evaluation unit is used to input the logs, warning records and multimodal features before and during the consultation into the compliance evaluation model after the consultation, to obtain the total compliance score, pre-consultation score, consultation process score, communication score and improvement suggestions, and then generate a consultation standard report.

[0171] In an embodiment of the present invention, a device for standardizing video consultation and treatment behaviors is provided, including: using a multimodal model of a client to perform compliance verification on the doctor's dress, environment and sitting posture of the doctor's image before the consultation, and when the compliance verification fails, generating and displaying adjustment suggestions based on the compliance verification results; using a collaborative architecture of the client and the cloud to perform multimodal real-time analysis on the real-time video stream, real-time audio stream and text content of the consultation process to obtain a real-time compliance score, and when the real-time compliance score is less than a preset score threshold, issuing an early warning reminder through the client; after the consultation is completed, the logs, early warning records and multimodal features before and during the consultation are input into the compliance evaluation model to obtain the total compliance score, pre-consultation score, consultation process score, communication score and improvement suggestions, and then generate a consultation standardization report. From the above description, it can be seen that the device for standardizing video consultation and treatment behaviors of the present invention can automatically perform compliance verification on the doctor's attire, environment and sitting posture before the consultation, perform multimodal real-time analysis on the real-time video stream, real-time audio stream and text content of the consultation process, and conduct compliance evaluation on the logs, warning records and multimodal features before and during the consultation. That is, the entire process is automated, without relying on human resource investment, which greatly reduces operating costs. In addition, the doctor's diagnosis and treatment behaviors in online video consultations can be standardized in real time through adjustment suggestions and warning reminders, which has good timeliness and alleviates the technical problems of traditional methods of standardizing video consultation and treatment behaviors, such as strong dependence on human resources, high operating costs and poor timeliness.

[0172] Optionally, the compliance verification unit is also used for: the client obtains the doctor's image before the consultation; uses a multimodal model to process the doctor's image before the consultation to obtain the doctor's clothing information, environment information and sitting posture information; and performs compliance verification on the doctor's clothing information, environment information and sitting posture information according to preset clothing standards, preset environment standards and preset sitting posture standards.

[0173] Optionally, the device is further configured to return to the step of executing the client obtaining the doctor's image before the consultation until the compliance check is passed.

[0174] Optionally, the multimodal real-time analysis unit is also used for: the client obtains the real-time video stream, real-time audio stream and text content of the medical consultation process; pre-processes the real-time video stream and real-time audio stream to obtain key frames in the real-time video stream and audio segments corresponding to the key frames in the real-time audio stream; sends the key frames, audio segments and text content to the cloud; the cloud-based visual model detects the key frames to obtain posture features, expression features and behavior features; the cloud-based audio segment performs speech recognition and semantic understanding to obtain speech features; the cloud-based large language model processes the text content to obtain diagnosis and treatment speech features; the posture features, expression features, behavior features, speech features and diagnosis and treatment speech features are fused, and the compliance score of the fused features is calculated to obtain a real-time compliance score.

[0175] Optionally, the multimodal real-time analysis unit is also used to: when the real-time compliance score is less than a preset score threshold, generate warning information in the cloud and send the warning information to the client, so that the client can issue a warning reminder based on the warning information, wherein the warning information includes: real-time compliance score, warning level and violation behavior.

[0176] Optionally, the compliance evaluation unit is also used to: generate time series features based on logs, warning records and multimodal features; perform Transformer encoding on the time series features to obtain a global context representation; construct a medical consultation knowledge graph based on the global context representation, wherein the nodes in the medical consultation knowledge graph include: doctors, patients and environment, and the edges in the medical consultation knowledge graph have weights, and the weights are the interaction activity; use a graph neural network to process the medical consultation knowledge graph to obtain a total compliance score, a score before consultation, a score during consultation, a communication score and improvement suggestions; generate a medical consultation standard report based on the time series features, the total compliance score, the score before consultation, the score during consultation, the communication score and improvement suggestions.

[0177] Optionally, the multimodal model includes: a distilled and pruned version of ViT-Base, a micro-branch of HRNet-W32, and a multi-task output head. The multimodal model is a model quantized to INT8; data transmission between the client and the cloud is encrypted; access to the medical consultation data is controlled, and sensitive information is masked.

[0178] The device provided in the embodiment of the present invention has the same implementation principle and technical effects as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment.

[0179] like Figure 10 As shown, an electronic device 600 provided in an embodiment of the present application includes: a processor 601, a memory 602 and a bus, wherein the memory 602 stores machine-readable instructions executable by the processor 601. When the electronic device is running, the processor 601 communicates with the memory 602 through the bus, and the processor 601 executes the machine-readable instructions to perform the steps of the method for standardizing video consultation and treatment behaviors as described above.

[0180] Specifically, the above-mentioned memory 602 and processor 601 can be general-purpose memory and processor, which are not specifically limited here. When the processor 601 runs the computer program stored in the memory 602, it can execute the above-mentioned method of standardizing video consultation and treatment behaviors.

[0181] The processor 601 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 601 or by instructions in the form of software. The above-mentioned processor 601 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 602, and processor 601 reads the information in memory 602 and performs the steps of the above method in conjunction with its hardware.

[0182] Corresponding to the above-mentioned method for standardizing video consultation and treatment behaviors, an embodiment of the present application also provides a computer-readable storage medium, which stores machine-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to execute the steps of the above-mentioned method for standardizing video consultation and treatment behaviors.

[0183] The device for standardizing video consultation and treatment behaviors provided in the embodiment of the present application can be specific hardware on the device or software or firmware installed on the device. The implementation principle and technical effects of the device provided in the embodiment of the present application are the same as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the aforementioned method embodiment, and will not be repeated here.

[0184] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0185] For another example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0186] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0187] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0188] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method for standardizing video consultation and treatment behaviors described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store program codes.

[0189] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and are not to be understood as indicating or implying relative importance.

[0190] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The scope of protection of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-mentioned embodiments within the technical scope disclosed in the present application, or perform equivalent replacements for some of the technical features thereof. However, these modifications, changes, or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application. They should all be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.< / host>

Claims

1. A method for standardizing video consultation and treatment behavior, characterized in that: include: The client's multimodal model is used to perform compliance verification on the doctor's clothing, environment, and sitting posture in the pre-consultation doctor image. If the compliance verification fails, adjustment suggestions are generated and displayed based on the compliance verification results. Through the collaborative architecture of the client and the cloud, a multimodal real-time analysis of the real-time video stream, real-time audio stream, and text content of the consultation process is performed to obtain a real-time compliance score. When the real-time compliance score is less than a preset score threshold, an early warning reminder is issued through the client; After the consultation is completed, the logs, warning records and multimodal features before and during the consultation are input into the compliance evaluation model to obtain the total compliance score, pre-consultation score, consultation process score, communication score and improvement suggestions, and then generate a consultation standard report.

2. The method according to claim 1, characterized in that The client's multimodal model verifies the compliance of the doctor's attire, environment, and sitting posture in pre-consultation images, including: The client obtains the doctor's image before the consultation; Processing the doctor's image before the consultation using the multimodal model to obtain the doctor's clothing information, environment information, and sitting posture information; The doctor's clothing information, environment information and sitting posture information are respectively checked for compliance according to preset clothing standards, preset environment standards and preset sitting posture standards.

3. The method according to claim 2, characterized in that When the compliance check fails, after generating and displaying an adjustment suggestion based on the compliance check result, the method further includes: Return to the step of the client obtaining the doctor's image before the consultation until the compliance check is passed.

4. The method according to claim 1, wherein Through the collaborative architecture of the client and the cloud, multimodal real-time analysis of the real-time video stream, real-time audio stream, and text content of the consultation process is performed, including: The client obtains the real-time video stream, real-time audio stream and text content of the consultation process; Preprocessing the real-time video stream and the real-time audio stream to obtain key frames in the real-time video stream and audio segments corresponding to the key frames in the real-time audio stream; Sending the key frame, the audio segment, and the text content to the cloud; The cloud-based visual model detects the key frames to obtain posture features, expression features, and behavior features; The cloud performs speech recognition and semantic understanding on the audio segment to obtain speech features; The large language model in the cloud processes the text content to obtain the diagnosis and treatment speech features; The posture features, the expression features, the behavior features, the voice features and the diagnosis and treatment speech features are fused, and a compliance score is calculated for the fused features to obtain the real-time compliance score.

5. The method according to claim 1, wherein When the real-time compliance score is less than a preset score threshold, an early warning reminder is issued through the client, including: When the real-time compliance score is less than the preset score threshold, the cloud generates warning information and sends the warning information to the client, so that the client issues a warning reminder based on the warning information, wherein the warning information includes: the real-time compliance score, warning level and violation behavior.

6. The method according to claim 1, characterized in that Inputting the logs, warning records, and multimodal features before and during the consultation into the compliance evaluation model includes: generating a time series feature according to the log, the warning record, and the multimodal feature; Performing Transformer encoding on the temporal features to obtain a global context representation; Constructing a medical inquiry knowledge graph based on the global context representation, wherein the nodes in the medical inquiry knowledge graph include: doctors, patients, and environments, and the edges in the medical inquiry knowledge graph have weights, and the weights are interaction activity; Using a graph neural network to process the medical consultation knowledge graph, the total compliance score, the pre-consultation score, the medical consultation process score, the communication score, and the improvement suggestions are obtained; The consultation standard report is generated based on the time series characteristics, the total compliance score, the pre-consultation score, the consultation process score, the communication score and the improvement suggestions.

7. The method according to claim 1, characterized in that The multimodal model includes: a distilled and pruned version of ViT-Base, a micro-branch of HRNet-W32, and a multi-task output head. The multimodal model is a model quantized to INT8; Data transmission between the client and the cloud is encrypted; Access to medical consultation data is controlled and sensitive information is masked.

8. A device for standardizing video consultation and treatment behavior, characterized in that: include: A compliance verification unit, configured to verify the compliance of the doctor's attire, environment, and sitting posture on the pre-consultation doctor's image using the client's multimodal model, and to generate and display adjustment suggestions based on the compliance verification results if the compliance verification fails; A multimodal real-time analysis unit is used to perform multimodal real-time analysis of the real-time video stream, real-time audio stream, and text content of the consultation process through a collaborative architecture between the client and the cloud to obtain a real-time compliance score, and to issue an early warning reminder through the client when the real-time compliance score is less than a preset score threshold; The compliance evaluation unit is used to input the logs, warning records and multimodal features before and during the consultation into the compliance evaluation model after the consultation is completed, obtain the total compliance score, pre-consultation score, consultation process score, communication score and improvement suggestions, and then generate a consultation standard report.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to execute the method according to any one of claims 1 to 7.