Medical scene dynamic interactive decision-making system and method based on multi-modal perception
By analyzing the multimodal fusion of medical images and medical record texts, the physician's inquiry statements are reconstructed and the patient's physiological voice and facial expressions are monitored. This solves the problem of dynamic adjustment and inquiry completion of multimodal perception technology in medical scenarios, and achieves more accurate and efficient auxiliary diagnosis and interactive response.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU SUNO BIOTECH
- Filing Date
- 2026-01-06
- Publication Date
- 2026-05-08
AI Technical Summary
Existing multimodal perception technologies lack the ability to link the saliency of image frame content with the semantic differences of medical case text in medical scenarios. This makes it difficult to dynamically adjust and complete the query when faced with missing or insufficient semantic coverage of case image information, affecting the completeness of assisted diagnosis and real-time response capabilities.
By extracting grayscale, edge, and texture features from medical images, candidate lesion areas are screened, inter-frame edge changes and texture differences are calculated, image structural saliency is analyzed, symptom keywords are identified by combining patient medical record text, semantically uncovered identifier data is generated, physician inquiry statements are reconstructed, and patients' voice, physiological signals, and facial expressions are monitored in real time to construct an abnormal co-occurrence factor set, adjust the interactive inquiry order, and generate dynamic interactive decision instructions.
It enhances the system's dynamic attention capability in the visual information dimension, effectively captures semantic omission areas, and realizes a complete closed loop from static image screening to dynamic interactive expression, thereby improving the accuracy and efficiency of multimodal perception in medical scenarios for assisted reasoning and interactive response.
Smart Images

Figure CN121999992A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal perception technology, and in particular to a dynamic interactive decision-making system and method for medical scenarios based on multimodal perception. Background Technology
[0002] The field of multimodal perception technology involves the perception, fusion, and analysis of the same object or scene using multiple sensors or information sources. Core aspects include the synchronous acquisition of multidimensional data such as voice, images, and text, the fusion processing of heterogeneous data, and decision support mechanisms for multi-source information. The technical fields mainly include the collaborative acquisition of multimodal signals, the temporal alignment and feature extraction of multimodal information, cross-modal information fusion mechanisms, and intelligent reasoning and discrimination based on fusion results. It is a key technical path for building perception, understanding, and interaction capabilities in complex intelligent systems.
[0003] Among them, the multimodal perception-based dynamic interactive decision-making system for medical scenarios refers to a decision support approach that uses multi-source information such as medical images, electronic medical records, physiological signals, and voice input, combined with the dynamic changes in the patient's scenario, to perform medical auxiliary analysis and interactive response. Traditional methods typically use medical image processing to extract key lesion features, utilize structured electronic medical records to mine and extract basic patient condition data, obtain the patient's physiological state through physiological signal recognition, receive doctor-patient interaction instructions through voice recognition mechanisms, and then use rule-based knowledge reasoning mechanisms or trained classification decision models to reason and respond to multi-source data, thereby completing the system's perception and response in the medical scenario.
[0004] Existing technologies for multimodal data fusion processing largely rely on rule-based reasoning models or well-trained classification systems. They lack the ability to identify the saliency of image frame content and the semantic differences in case text. When faced with cases where there is a lack of local information or insufficient semantic coverage in the case images, the system struggles to make dynamic adjustments and complete queries. Especially when patients present with complex clinical symptoms or multiple ambiguous areas, the system often makes judgment errors or misses key signs due to image-text matching deviations. Furthermore, the lack of a joint perception mechanism for changes in patients' language, physiology, and facial emotions makes it impossible to fully capture dynamic abnormal signals in doctor-patient interactions, thus affecting the completeness of auxiliary diagnosis and real-time response capabilities. Summary of the Invention
[0005] To address the technical problems existing in the prior art, embodiments of the present invention provide a dynamic interactive decision-making system and method for medical scenarios based on multimodal perception. The technical solution is as follows:
[0006] On the one hand, a dynamic interactive decision-making system for medical scenarios based on multimodal perception is provided, the system comprising:
[0007] The master node identification module extracts grayscale, edge and texture features from medical images, filters candidate lesion areas, calculates inter-frame edge changes and texture differences, analyzes image structural saliency, selects image frames as image master nodes and passes them to the coverage determination module and reconstruction guidance module.
[0008] The coverage judgment module extracts the patient's medical record text based on the dominant node of the image, identifies symptom keywords and compares them with the lesion labels corresponding to the dominant node of the image, generates semantically uncovered identification data, and transmits it to the reconstruction guidance module.
[0009] The reconstruction guidance module extracts uncovered keywords and locates missing semantic directions through the dominant nodes of the image and the semantically uncovered identifier data, reconstructs the physician's inquiry statement, outputs the physician's inquiry statement set and passes it to the factor construction module.
[0010] The factor construction module monitors the patient's voice, physiological signals, and facial expressions in real time through the physician's question set. It extracts the keyword intervals and speech rate changes in the voice, combines physiological indicators to locate signal segments, identifies waveform inversion and amplitude jump features, and combines facial expression muscle tension changes to construct an abnormal co-occurrence factor set and transmit it to the path rearrangement module.
[0011] As a further aspect of the present invention, the image dominant node includes image frame number, dominant region location, and structural saliency score; the semantically uncovered identifier data includes a list of missing labels, keyword matching degree, and semantic coverage rate; the physician inquiry statement set includes target symptom words, reconstructed statement templates, and candidate guiding words; and the abnormal co-occurrence factor set includes voice emotion mutation points, facial expression abnormal patterns, and physiological signal abnormal features.
[0012] As a further aspect of the present invention, the master node identification module includes:
[0013] The candidate screening submodule extracts grayscale, edge and texture features from medical images, combines boundary information to extract region contours, filters regions with abnormal grayscale distribution and constructs a set of edge contours, compares the consistency between contours and gradients, identifies potential lesion regions, and generates candidate region consistency coefficient values.
[0014] The edge texture analysis submodule obtains the edge distribution map of the corresponding region in consecutive image frames based on the candidate region consistency coefficient value, calculates the edge intensity difference and texture offset rate between frames, filters out regions whose edge offset amplitude and texture distribution change rate exceed the edge offset reference range and texture fluctuation reference interval, and obtains the joint offset rate value of structural change.
[0015] The edge offset reference range is defined as the upper and lower boundaries of normal edge changes by extracting the edge intensity difference of non-lesion areas in consecutive image frames, statistically analyzing the average level and fluctuation amplitude.
[0016] The texture fluctuation reference range is defined by extracting the gray-level co-occurrence matrix energy, contrast and local binary mode differences between frames in non-lesion areas, and setting upper and lower thresholds based on common variation ranges.
[0017] The dominant frame recognition submodule extracts the structural contour and internal texture aggregation degree of the corresponding region based on the joint offset rate value of the structural change, calculates the structural density and compares it with the salience density benchmark value, identifies the image frames that meet the deviation standard and uses them as dominant nodes, and establishes the image dominant node.
[0018] The specific formula for extracting the structural contour and internal texture aggregation degree of the corresponding region is as follows:
[0019] ;
[0020] Calculate the structural density value ;
[0021] in, Represents the structural density value. Represents the total number of pixels in the extracted area. Representing the The grayscale gradient value of each pixel in the horizontal structural gradient direction. Representing the The grayscale gradient value of each pixel in the vertical structural gradient direction. Representing the The gradient weight values of the texture clustering degree of each pixel. This represents the structural variance adjustment coefficient, used to dynamically adjust the texture structure response weights. Representing the The absolute deviation of the grayscale value of each pixel within the local window. This represents the average local structural density value calculated from the joint structural texture features of all pixels in the extracted region. The index number of the pixel. To extract the total number of pixels within the region;
[0022] The saliency density benchmark value is a standard set by statistically analyzing the distribution range of high-density areas based on the degree of aggregation of structural and textural features in known lesion areas.
[0023] The deviation standard is the minimum identifiable boundary set by comparing the numerical differences in the structure and texture aggregation scores between the dominant region and the ordinary region.
[0024] As a further aspect of the present invention, the coverage determination module includes:
[0025] The dominant node extraction submodule identifies the structural features of the corresponding region based on the dominant node of the image. By analyzing the grayscale, edge and texture features of the lesion candidate area, it combines, classifies and groups each feature by type. Based on the type grouping results, it outputs the lesion type label of each region, establishes the corresponding set of nodes and lesions, and generates the image lesion label comparison quantity.
[0026] The symptom keyword recognition submodule calls the image lesion label comparison quantity and obtains the patient's medical record text data. It extracts symptom description words from the text through word segmentation, filters keywords corresponding to disease symptoms, constructs a term set based on the frequency of keyword occurrence and part-of-speech structure, and generates symptom keyword extraction quantity.
[0027] The semantic comparison and judgment submodule compares the content of the image lesion label control quantity with the extracted symptom keywords, constructs a term missing index based on the semantic matching degree of the set, calculates the matching gap rate between the number of missing terms and the total number of labels, establishes the number of uncovered terms and the corresponding node identifier, and generates semantic uncovered identifier data.
[0028] As a further aspect of the present invention, the reconstruction guidance module includes:
[0029] The keyword extraction submodule extracts the lesion content of the associated nodes in the missing identifiers through the dominant nodes of the image and the semantically uncovered identifier data, locates the set of missing terms corresponding to the identifier number and extracts keywords, establishes a keyword list after screening non-symptom terms, and generates the number of uncovered keywords.
[0030] The semantic direction determination submodule calls the number of uncovered keywords, obtains the corresponding anatomical structure labels and spatial orientations in the dominant nodes of the image, filters the misaligned content of missing terms in the label set, establishes the directional pairing relationship between missing terms and structural labels, calculates the frequency ranking, and generates missing semantic direction trend values.
[0031] The query reconstruction submodule matches the tag templates in the semantic library based on the missing semantic direction trend value, extracts high-frequency structural keywords and combines them with the corresponding query sentence patterns, constructs a query sentence rearrangement sequence driven by missing keywords and generates sentence pattern groups in sequence, and generates a set of physician query sentences.
[0032] As a further aspect of the present invention, the factor construction module includes:
[0033] The semantic rhythm extraction submodule acquires the patient's voice signal in real time through the physician's question set. By extracting the time interval and pronunciation duration between keywords in the voice, it segments the speech rate change and interval time sequence, calculates the speech rate slope and rhythm fluctuation degree of each segment, normalizes the amplitude of speech rhythm variation on the time axis, and establishes the speech rhythm fluctuation rate value.
[0034] The physiological signal recognition submodule extracts electrocardiogram and electrodermal signals for the corresponding time period based on the time range covered by the speech rhythm fluctuation rate value, performs amplitude difference and polarity direction detection on the signal waveform, calculates the number of jumps and the frequency of reversals per unit time, groups and classifies them according to jump amplitude and reversal frequency, and obtains the fluctuation signal jump ratio value.
[0035] The facial collaborative mapping submodule calls the time point information marked by the fluctuation signal jump ratio, extracts the facial image sequence of the corresponding time period and divides the region, tracks the pixel intensity change in the region, calculates the muscle tension change, identifies the mutation points in facial muscle tension data, speech rhythm and various physiological signals, analyzes the overlapping segments of the three types of mutation points on the time axis and counts the co-occurrence frequency, and establishes an abnormal co-occurrence factor set.
[0036] As a further aspect of the present invention, the specific formula for detecting abrupt changes in speech rhythm is as follows:
[0037] ;
[0038] Calculate the trend value of sudden changes in speech rhythm;
[0039] in, For the first The absolute value of the rate of change of the slope of the segmental speech rhythm. For the first The speech rate slope of a segment of speech, For the first The speech rate slope of a segment of speech, For the first Section and the The start time difference between segments For the first The average syllable interval length of keywords in the paragraph For the first The average duration of pronunciation of keywords in the paragraph For the first The average speech energy density of a segment containing keywords. This is the index number of the current speech segment in the rhythm sequence. This is an index variable used to iterate through the speech segment numbers within the summation interval.
[0040] As a further aspect of the present invention:
[0041] The path reordering module extracts the symptom label order in the auxiliary consultation path through the abnormal co-occurrence factor set, matches the semantic direction of the labels with the dominant nodes of the image and keywords, reorders the interactive inquiry order according to the matching situation, adjusts the execution order of the physician's interactive response, and generates a dynamic interactive decision instruction set for the medical scenario.
[0042] The dynamic interactive decision-making instruction set for medical scenarios includes query priority order, tag semantic mapping path, and interactive response adjustment scheme.
[0043] As a further aspect of the present invention, the path reordering module includes:
[0044] The symptom tag extraction submodule extracts the auxiliary consultation path through the abnormal co-occurrence factor set, identifies the symptom tag corresponding to each node in the path, records the arrangement order of the tags in the original path, constructs a tag-order position correspondence set, and generates tag path order value;
[0045] The assisted consultation path refers to a standardized inquiry sequence pre-set by the system and constructed based on a clinical knowledge base and disease evolution process. It is used to provide physicians with a structured symptom label order and inquiry template during intelligent medical interaction, and to assist in the systematic and comprehensive collection of patient status. The path, as a basic interaction framework, contains multiple symptom nodes arranged according to disease characteristics, and their order is dynamically adjusted according to the dominant nodes of the image and the direction of semantic keywords.
[0046] The semantic matching determination submodule calls the tag path order value and matches the dominant node of the image with the semantic direction content of the keyword. It calculates the number of matches based on the position of the tag in the image region and the directional terms, constructs a tag and node matching table, filters the priority by quantity, and generates the path semantic matching coefficient.
[0047] The interaction instruction generation submodule adjusts the response order of symptom tags in the interaction process according to the path semantic matching coefficient, rearranges the tag call structure in the original path according to the matching priority, establishes the mapping relationship between tag nodes and response content, and generates a dynamic interactive decision instruction set for medical scenarios.
[0048] On the other hand, a dynamic interactive decision-making method for medical scenarios based on multimodal perception is provided. This method is applied to a dynamic interactive decision-making system for medical scenarios based on multimodal perception, and includes:
[0049] S1: By extracting grayscale, edge and texture features from medical images, regions with closed contours and complex textures are selected as lesion candidate areas. The morphological change trend of regions in consecutive frames is calculated to determine the degree of structural mutation. The differences in image structure in the frame sequence are optimized, and image frames with prominent structural changes are selected to generate dominant image nodes.
[0050] S2: Based on the dominant node of the image, obtain the annotation information of the corresponding image frame and the symptom keywords in the case text, analyze the semantic matching relationship between the keywords and lesion labels, filter out keywords that do not appear in the semantic corresponding items in the labels, identify the missing positions in the text semantic chain, generate a set of keywords that cannot be associated, and construct semantically uncovered identification data.
[0051] S3: Using the dominant node of the image and the semantically uncovered identifier data, locate the contextual logical position of the keywords, calculate the topic association direction in the symptom expression sequence, filter sentence structure templates that match the semantic direction, adjust the word order and position combination of keywords in the sentence, analyze the grammatical integrity of the generated sentences, construct a set of questions, and output a set of physician inquiry statements.
[0052] S4: Through the physician's question set, monitor the patient's voice, physiological signals, and facial expressions in real time, analyze the speech rate change trend between voice keywords, compare the positions of frequency abrupt changes and waveform changes in physiological signals, screen for abnormal tension transition areas in image sequences, integrate abnormal synchronization events, and generate a set of abnormal co-occurrence factors.
[0053] S5: Using the set of abnormal co-occurrence factors, mark the time points of abnormal statements and the label sequence in the physician inquiry statement set, analyze the correspondence between the symptoms guided by the statements and the labels of the dominant nodes in the images, analyze the consistency between semantic direction and label logic, adjust the sorting structure of statements in the inquiry process, optimize the interaction path order, and construct a dynamic interactive decision instruction set for medical scenarios.
[0054] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
[0055] By identifying key image nodes through image structure saliency analysis and combining image texture changes for frame filtering, the system gains dynamic attention capabilities in the visual information dimension. Semantic content comparison is used to determine the coverage relationship between image information and medical case text, effectively capturing semantically missing areas. When generating physician inquiry statements, reconstruction and completion are performed in the direction of the missing information. Furthermore, real-time monitoring of speech, physiological signals, and facial expression data is overlaid, and features are extracted from unstructured interactive information to construct a set of abnormal response factors. Under multimodal conditions, a complete closed loop from static image filtering to dynamic interactive expression is achieved, enhancing the system's overall performance in semantic integrity, response timeliness, and interaction depth. This improves the accuracy and efficiency of multimodal perception in assisting reasoning and interactive responses in medical scenarios. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a system flowchart of the present invention;
[0058] Figure 2 This is a schematic diagram of the system framework of the present invention;
[0059] Figure 3 This is a schematic diagram of the method steps of the present invention. Detailed Implementation
[0060] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0061] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0062] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0063] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0064] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0065] This invention provides a dynamic interactive decision-making system for medical scenarios based on multimodal perception. Please refer to [link / reference]. Figures 1 to 2 This invention provides a technical solution: a dynamic interactive decision-making system for medical scenarios based on multimodal perception, comprising:
[0066] The master node identification module extracts grayscale, edge and texture features from medical images, filters candidate lesion areas, calculates inter-frame edge changes and texture differences, analyzes image structural saliency, selects image frames as image master nodes and passes them to the coverage determination module and reconstruction guidance module.
[0067] For example, CT images used to identify suspicious nodules in lung cancer screening, images of abnormally shaped blood cells used to identify abnormal morphology in hematological pathology analysis, or cryo-electron microscopy images of structurally abnormal viral particles used in vaccine quality control.
[0068] The coverage judgment module extracts patient medical record text or related quality control logs based on the dominant node of the image, identifies symptom keywords or process parameter deviations, compares them with the lesion labels or defect classifications corresponding to the dominant node of the image, generates semantically uncovered identification data, and passes it to the reconstruction guidance module.
[0069] The reconstruction guidance module extracts uncovered keywords or parameters and locates the missing semantic direction through image dominant nodes and semantically uncovered identifier data, reconstructs physician inquiry statements or quality control instructions, outputs physician inquiry statement sets or process inquiry instruction sets and passes them to the factor construction module.
[0070] The factor construction module monitors patients' speech, physiological signals, and facial expressions in real time using a set of physician inquiry statements. It extracts keyword intervals and speech rate changes in speech, identifies waveform inversion and amplitude jump features by locating signal segments with physiological indicators, and constructs a set of abnormal co-occurrence factors by combining facial expression muscle tension changes. This set is then passed to the path rearrangement module. In automated quality control scenarios, this module can be adapted as a process factor construction module. By analyzing the correlation between production line sensor data (such as temperature, pressure, and vibration frequency) and defects at dominant nodes, a set of process abnormal factors can be constructed.
[0071] The path reordering module extracts the symptom label order or examination steps in the auxiliary consultation path or standard operating procedure through the abnormal co-occurrence factor set, matches the semantic direction of the dominant node and keywords of the label with the image, reorders the interactive inquiry order or the priority of the operation process according to the matching, adjusts the execution order of physician interactive response or system automated response, and generates a dynamic interactive decision instruction set for medical scenarios.
[0072] The image-dominant node includes the image frame number, dominant region location, and structural salience score; the semantically uncovered identifier data includes the missing label list, keyword matching degree, and semantic coverage rate; the physician inquiry statement set includes target symptom words, reconstructed statement templates, and candidate guiding words; the abnormal co-occurrence factor set includes voice emotion mutation points, facial expression abnormal patterns, and physiological signal abnormal features; and the medical scenario dynamic interaction decision instruction set includes inquiry priority order, label semantic mapping path, and interaction response adjustment scheme.
[0073] Please see Figure 1 and Figure 2 The master node identification module includes:
[0074] The candidate screening submodule extracts grayscale, edge and texture features from medical images, combines boundary information to extract region contours, filters regions with abnormal grayscale distribution and constructs a set of edge contours, compares the consistency between contours and gradients, identifies potential lesion regions, and generates candidate region consistency coefficient values.
[0075] The candidate filtering submodule first acquires the grayscale information of the medical image. It reads the grayscale value of each pixel and stores it in a two-dimensional matrix. For example, for a 512×512 CT image, it scans the pixels row by row to record the grayscale value distribution. During this process, it calculates the average grayscale and variance of each 8×8 small region to identify areas with outlier grayscale distribution. Next, it extracts image edge features and performs edge difference processing on the image using a fixed threshold method. For example, for pixel (i,j), the edge intensity value is calculated based on the difference between its four neighbors. If the value is greater than the edge strength threshold of 20, it is marked as an edge point. Then, texture features are obtained by calculating the contrast and entropy of the gray-level co-occurrence matrix in each small region of the image to reflect texture differences. Subsequently, the region contour is extracted by combining boundary information. The effective region boundary is determined by the contour closure condition based on the edge point aggregation method. Starting from the upper left corner, the edge point group is traversed clockwise, and the coordinates of the closed region edge sequence are recorded. For example, in a certain image segment, the extracted edge point sequence is [(10,10),(10,20),(20,20),(20,10)] to form a closed region. Then, regions with abnormal gray-level distribution are screened, and the aforementioned abnormal gray-level mean square regions are intersected with the edge closed regions. Set matching is performed. If the center pixel value deviates from the average gray level of the entire image by more than a set threshold σ=12 (e.g., if the average gray level of the entire image is 120 and the center gray level of this region is 140, the deviation is 20, and it is marked as abnormal), then the region is marked as a potential lesion candidate. Then, an edge contour set is constructed, and the gradient change sequence corresponding to each contour edge is statistically analyzed. The gradient direction of the contour edge point is compared with the angle consistency with its adjacent points. If the gradient angle between more than 90% of the edge points is less than 15°, then the contour is considered to be consistent with the gradient direction. Finally, the consistency coefficient value of the candidate region is generated according to the proportion of consistent points. For example, if 92 out of 100 edge points in a candidate region have an angle less than 15°, then the consistency coefficient is 0.92.
[0076] Application Scenario 1: In blood quality analysis, this submodule is used to screen for abnormal red blood cells from microscopic blood smear images. First, color features are extracted through color space conversion (e.g., RGB to CIELAB) to identify cells with excessively large or absent central pale areas. Then, the Canny edge detection operator is used to precisely locate cell membrane boundaries, marking non-circular (e.g., sickle-shaped, elliptical) or irregularly shaped (e.g., spiky, serrated) contours. The roundness of the contour is calculated... The deviation from the standard red blood cell model is used to generate a consistency coefficient value for candidate abnormal cells.
[0077] Application Scenario 2: In vaccine-related disease control imaging analysis, this submodule is used to identify invalid vaccine particles in cryo-electron microscopy images. First, the image contrast is normalized, and a speckle detection algorithm (such as the Laplacian of Gaussian operator, LoG) is used to initially locate potential virus-like particles (VLPs). Then, the contour of each candidate particle is elliptical-fitted, and its major and minor axis ratios are calculated to filter out particles with abnormal morphology due to aggregation or breakage. By comparing the projected density distribution of the candidate particles with the template of the standard VLP model, the structural consistency coefficient value of the candidate particles is generated.
[0078] Table 1. Edge Intensity of Non-lesion Areas
[0079] Image frame number Intensity difference at the edge of non-lesion areas Average strength difference Fluctuation range 1 12.5 13.9 1.25 2 14.8 13.9 0.90 3 13.3 13.9 0.60 4 15.0 13.9 1.10
[0080] As shown in Table 1, the edge intensity changes of non-lesion areas in consecutive image frames are recorded, which can be used to extract the upper and lower limits of the edge offset reference range.
[0081] The edge texture analysis submodule obtains the edge distribution map of the corresponding region in consecutive image frames based on the candidate region consistency coefficient value, calculates the edge intensity difference and texture offset rate between frames, filters out regions whose edge offset amplitude and texture distribution change rate exceed the edge offset reference range and texture fluctuation reference interval, and obtains the joint offset rate value of structural change.
[0082] The edge texture analysis submodule receives the consistency coefficient values of candidate regions. Regions with a consistency coefficient greater than 0.85 are selected for analysis. The edge map distribution of this region is extracted frame by frame in consecutive image frames. For example, the edge intensity distribution of a candidate region is extracted in frames 1 and 2 respectively. The intensity value of each edge point is recorded, such as [15,18,20,22] and [14,17,21,25]. The difference is calculated point by point [1,1,1,-3], and the average difference intensity is set to 1.5 units. If this average intensity difference exceeds the upper and lower limits of the edge offset reference range (for example, if the average edge difference calculated from a non-lesion region is 1.2 with a standard deviation of 0.3, then the upper and lower limits are [0.9,1...]), the analysis is performed. If the difference is less than 0.5], it is recorded as an abnormal offset. The texture offset rate is reflected by the contrast and energy of the gray-level co-occurrence matrix and the inter-frame change rate of the LBP histogram. The difference rate is calculated as follows: if the contrast of frame 1 is 0.25 and that of frame 2 is 0.34, the change rate is 36%. If it exceeds the upper limit set in the texture fluctuation reference range (the average change of the reference non-lesion area is 12%, the standard deviation is 8%, and the upper limit is 28%), then the texture is considered abnormal. Then, the edge offset amplitude and the normalized value of the texture fluctuation change are weighted and added together. The weights are set to 0.4 and 0.6, respectively. If the two normalized values are 0.7 and 1.2, respectively, then the joint offset rate of structural change is 0.4×0.7+0.6×1.2=1.02.
[0083] The edge offset reference range is defined by extracting the edge intensity difference of non-lesion areas in consecutive image frames, statistically analyzing the average level and fluctuation amplitude, and setting it as the upper and lower boundaries of normal edge changes.
[0084] The texture fluctuation reference range is determined by extracting the gray-level co-occurrence matrix energy, contrast, and local binary mode differences between frames in non-lesion areas, and setting upper and lower thresholds based on common variation ranges.
[0085] The dominant frame recognition submodule extracts the structural contour and internal texture aggregation degree of the corresponding region based on the joint offset rate value of structural changes, calculates the structural density and compares it with the saliency density benchmark value, identifies image frames that meet the deviation standard and uses them as dominant nodes, and establishes image dominant nodes.
[0086] The specific formula for extracting the structural contour and internal texture clustering of the corresponding region is as follows:
[0087] ;
[0088] Calculate the structural density value ;
[0089] in, Represents the structural density value. Represents the total number of pixels in the extracted area. Representing the The grayscale gradient value of each pixel in the horizontal structural gradient direction. Representing the The grayscale gradient value of each pixel in the vertical structural gradient direction. Representing the The gradient weight values of the texture clustering degree of each pixel. This represents the structural variance adjustment coefficient, used to dynamically adjust the texture structure response weights. Representing the The absolute deviation of the grayscale value of each pixel within the local window. This represents the average local structural density value calculated from the joint structural texture features of all pixels in the extracted region. The index number of the pixel. To extract the total number of pixels within the region;
[0090] formula:
[0091] ;
[0092] Detailed explanation of the formula and its calculation derivation:
[0093] The formula is used to calculate the structural density value of an image region. This value reflects the combined characteristics of structural contours and texture aggregation within an image region, and is used for the identification of dominant nodes in subsequent image frames.
[0094] Parameter meanings and settings:
[0095] Extract the total number of pixels within the region, set to (3×3 pixel area);
[0096] : No. The grayscale gradient value of each pixel in the horizontal structural gradient direction is calculated using the Sobel operator and set as follows: ;
[0097] : No. The grayscale gradient value of each pixel in the vertical structural gradient direction is calculated using the Sobel operator and set as follows: ;
[0098] : No. The gradient weight value of the texture clustering degree of each pixel is calculated using the local standard deviation and set to [value]. ;
[0099] The structural variance adjustment coefficient is set based on the overall grayscale variance of the image, and is set to [value missing]. ;
[0100] : No. The absolute deviation of the grayscale value of a pixel within a local window is calculated as the absolute value of the difference between the pixel's grayscale value and the local mean, and is set as follows: ;
[0101] The average local structure density value calculated from the joint structural texture features of all pixels in the region is set as follows: ;
[0102] Substitute the parameters into the formula to calculate:
[0103] ;
[0104] ;
[0105] ;
[0106] ;
[0107] ;
[0108] ;
[0109] result This indicates a high structural density in the image region, reflecting obvious structural contours and texture clustering features. This result is used for dominant node identification in subsequent image frames, assisting in determining whether an image frame meets the deviation criteria, thereby establishing the dominant node in the image.
[0110] In blood quality analysis scenarios, the dominant frame recognition submodule calculates cell morphology abnormality scores. The formula is: ;in The roundness of the cell outline. For edge roughness, This represents the normalized deviation of the area from the standard value. These are their respective weighting coefficients. Blood smear images with values exceeding a threshold (e.g., 0.75) are identified as the dominant nodes.
[0111] In vaccine-related disease control imaging analysis scenarios, the dominant frame recognition submodule calculates the vaccine particle integrity score. The formula is: ;in The total number of particles, For the first The similarity between the shape of each particle and the standard template. To ensure its internal density uniformity, The global aggregation coefficient is calculated from the reciprocal of the average distance between particles. As weights. Image frames with values below a set baseline (e.g., 0.8) are identified as dominant nodes.
[0112] The saliency density benchmark is a standard set by statistically analyzing the distribution range of high-density areas based on the degree of aggregation of structural and textural features in known lesion regions.
[0113] The deviation standard is the minimum identifiable boundary set by comparing the numerical differences in the structure and texture aggregation scores between the dominant region and the ordinary region.
[0114] Please see Figure 1 and Figure 2 The coverage judgment module includes:
[0115] The dominant node extraction submodule identifies the structural features of the corresponding region based on the dominant node of the image. By analyzing the grayscale, edge and texture features of the lesion candidate area, it combines, classifies and groups each feature. Based on the type grouping results, it outputs the lesion type label of each region, establishes the corresponding set of nodes and lesions, and generates the image lesion label comparison quantity.
[0116] The dominant node extraction submodule, based on the region node number information extracted from the aforementioned dominant frame, sequentially quantizes and records the structural features of the region associated with each node. First, it extracts a rectangular window region corresponding to the node coordinates from the grayscale image and extracts the grayscale value data of this region into a vector set. The grayscale fluctuation of the region is determined by calculating the difference between the maximum and minimum grayscale values. If the difference is greater than 30, it is marked as a high-contrast region. For example, the grayscale value range of region 1 is 120 to 150, with a fluctuation of 30. Edge features are calculated by extracting the gradient magnitude at the edge points of this region and calculating the average gradient. Assuming there are 100 edge points in this region, the gradient values are calculated individually and then averaged to obtain 18.2 units. Texture features are obtained by statistically analyzing the proportion of the dominant mode in LBP encoding and combining it with the contrast of the co-occurrence matrix to obtain the texture clustering degree. For example, if the dominant mode proportion is 0.65 and the contrast is 0.3, the texture clustering degree is calculated as follows: Then, the three features of each region are clustered and grouped. Based on the three categories of inflammation, nodules and fibrosis, the classification is based on whether the median gray value, the mean edge gradient, and the mean texture aggregation fall within the grouping threshold range. The median gray value is set to be higher than 135 for high density group, the mean gradient is set to be higher than 20 for hard boundary group, and the texture aggregation is set to be lower than 0.3 for structure concentration group. For example, the median gray value of node 1 is 135, the gradient is 18.2 and the texture aggregation is 0.28, which meets the characteristics of high density and structure concentration, so it is marked as high density inflammation type. Finally, the mapping relationship between each node number and the corresponding region type is established one by one, and the node lesion type label set is output in sequence.
[0117] Table 2. Characteristics of Lesions at Dominant Nodes
[0118] Dominant Node Number Gray-scale feature range Marginal gradient mean Texture clustering Grouping type Lesion label Node 1 [120,150] 18.2 0.28 High-density inflammation Tag A Node 2 [110,130] 20.5 0.32 Calcified nodules Tag B Node 3 [125,145] 19.1 0.26 fibrosis Tag C
[0119] As shown in Table 2, after classifying the node numbers and combinations of grayscale, edge, and texture features, the corresponding lesion label set is output.
[0120] The symptom keyword recognition submodule calls the image lesion label comparison quantity and obtains the patient's medical record text data. It extracts symptom descriptive words from the text through word segmentation, filters keywords corresponding to disease symptoms, constructs a term set based on the frequency and part-of-speech structure of the keywords, and generates the symptom keyword extraction quantity.
[0121] After acquiring the number of image lesion label comparisons, the symptom keyword recognition submodule loads the patient's text medical record data and performs preliminary cleaning, deleting non-symptom fields such as patient basic information and treatment history, retaining only the chief complaint and present illness paragraphs. It then decomposes the data into basic word sequences using natural language processing. For example, the original paragraph "patient has had cough, sputum, and chest pain for the past two months" is divided into the word sequence ["patient", "past two months", "cough", "sputum", "accompanied", "chest pain"]. In the initial screening stage, stop words and non-disease symptom terms are removed, leaving ["cough"]. [cough, phlegm, chest pain], the frequency of each word in the text is obtained for each word in the set. If it exceeds the threshold (e.g., 2 times), it is considered a high-frequency word. During part-of-speech analysis, its part of speech is checked to see if it is "n" or "vn". The word is retained. If both frequency and part-of-speech requirements are met, it is included in the candidate keyword set. For example, if "cough" appears 3 times in the medical record, its part of speech is "n", and it is retained as a keyword. Finally, all candidate keywords are recorded in the symptom keyword extraction array. The array result is as follows: ["cough", "phlegm", "chest pain"].
[0122] In the blood quality analysis scenario, this submodule extracts keywords such as "anemia," "jaundice," and "splenomegaly" from the "clinical diagnosis" section of the test request form to generate a set of clinical characteristic keywords.
[0123] In the vaccine disease control imaging analysis scenario, this submodule extracts key process parameters from the vaccine batch production log, such as "pH value = 7.2", "storage temperature = 4.1°C", and "oscillation frequency = 0Hz", and generates a set of process parameter keywords.
[0124] The semantic comparison and judgment submodule compares the label content in the image lesion label control quantity with the number of extracted symptom keywords, constructs a term missing index based on the semantic matching degree of the set, calculates the matching gap rate for the number of missing terms and the total number of labels, establishes the number of uncovered terms and the corresponding node identifier, and generates semantic uncovered label data.
[0125] After receiving the extracted symptom keywords, the semantic comparison and judgment submodule compares them item by item with the lesion label content already marked in the image lesion label comparison. First, it parses the set of typical symptom phrases corresponding to each label. For example, the corresponding terms under "Label A" are ["cough", "chest pain", "shortness of breath"]. The keyword extraction quantity is compared with the intersection of this set to extract the number of missing terms. That is, in this example, "shortness of breath" is not covered, so the missing term is 1. The total number of terms for label A is recorded as 3, and the missing term is 1. The matching gap rate is calculated. In this way, semantic comparison is performed on all tags to obtain the number of missing terms, the total number of terms, and the gap rate for each tag. Tags with a gap rate greater than the set threshold of 0.25 are recorded as uncovered and their number mapping relationship with the dominant node is established. Finally, semantic uncovered identification data in the form of {node1:[tag A-0.33]} is output.
[0126] In a blood quality analysis scenario, if the image label of the dominant node is "sickle cell", but the clinical diagnostic keywords do not include "sickle cell anemia" or "hereditary disease history", the system determines that there is a semantic gap and generates the label data {blood smear node 1:{[}sickle cell-1.0{]}}, indicating that there may be a mismatch between the test results and clinical information.
[0127] In the vaccine disease control imaging analysis scenario, if the image label of the dominant node is "particle aggregation", while the process parameter keywords in the production log are all "normal range", the system determines that there is semantic non-coverage and generates the label data {electron microscope image node 1:{[}particle aggregation-1.0{]}}, indicating that there is a contradiction between the imaging results and the production records.
[0128] Please see Figure 1 and Figure 2 The refactoring of the bootloader module includes:
[0129] The keyword extraction submodule extracts the lesion content of the associated nodes in the missing identifiers through the dominant nodes of the image and the semantically uncovered identifier data, locates the set of missing terms corresponding to the identifier number and extracts keywords, establishes a keyword list after screening non-symptom terms, and generates the number of uncovered keywords.
[0130] After receiving the dominant image node and semantically uncovered identifier data, the keyword extraction submodule first calls the node numbers in the uncovered identifiers item by item, and locates the corresponding lesion areas and label sets. It extracts all associated descriptive fields from the node label set. For example, the description under the node 1 label is "increased density in the hilar region, blurred image structure". Then, it compares the node number with the missing term set corresponding to the semantic gap index table. For example, the missing terms for node 1 are ["shortness of breath", "coughing up phlegm"]. It filters all terms in the node label description, removes structural words, conjunctions, and judgment words, and only retains the candidate term set that may be used as symptoms, such as ["shortness of breath", "coughing up phlegm", "image"]. Then, it removes non-"n" and "vn" type terms such as "image" through part-of-speech filtering. Finally, the remaining terms are ["shortness of breath", "coughing up phlegm"]. It constructs a list of missing keywords for the node and counts the number as 2, which is recorded as the number of uncovered keywords. After completing this process in the entire node set, it summarizes and outputs the statistical results of the number of missing keywords under each node.
[0131] The semantic direction determination submodule calls the number of uncovered keywords, obtains the corresponding anatomical structure labels and spatial orientations in the dominant nodes of the image, filters the misaligned content of missing terms in the label set, establishes the directional pairing relationship between missing terms and structural labels, calculates the frequency ranking, and generates missing semantic direction trend values.
[0132] After calling the aforementioned number of uncovered keywords, the semantic direction determination submodule reads the structural annotation information of the dominant node in the image according to the node number, and matches the mapping relationship between the structural label and the anatomical atlas. For example, "Node 1" is labeled as "hilum," "Node 2" as "lower lobe of the right lung," and "Node 3" as "tracheal branch." It performs semantic alignment judgment on each keyword and structural label in the missing term set. The judgment method is to analyze whether the typical corresponding structure of the keyword in the standard symptom-anatomical comparison table contains the current structural label keyword. For example, if the term "coughing up phlegm" is associated with "bronchus, alveoli" in the standard library, but the current node structure is "hilum," it is judged as an unaligned term and recorded as a direction inconsistency. The module counts the number of successful and unsuccessful alignments of all missing keywords in each node, and calculates the semantic direction trend value as the proportion of unaligned terms to the total number of missing terms. For example, if Node 3 has 3 missing terms, 2 of which are unaligned, then the trend value is... Finally, the structural orientation deviation of all nodes is output as a semantic orientation trend data set.
[0133] The inquiry reconstruction submodule matches the tag templates in the semantic library based on the missing semantic direction trend value, extracts high-frequency structural keywords and combines them with the corresponding inquiry sentence patterns, constructs a missing keyword-driven question rearrangement sequence and generates sentence pattern groups in sequence, and generates a set of physician inquiry sentences.
[0134] After receiving the semantic direction trend value of each node, the query reconstruction submodule sorts the nodes by trend value from high to low priority, and performs query reconstruction operation on nodes with trend value greater than 0.5. It reads the structure-question template combination pairs related to the missing keywords in the semantic library, such as the sentence "Do you have persistent cough?" corresponding to "coughing-bronchus". It filters out the sentence groups that match the structure of the current node, and then calls the corresponding template content according to the missing words in sequence. For example, the missing words of node 3 are "cough, chest tightness", and the structure is "tracheal branch". It retrieves sentences such as "Do you have an irritating cough?" and "Have you felt chest tightness recently?" from the sentence library, and combines them into new sentence groups according to the order of missing words. Finally, it constructs the query sequence under the direction of "cough → chest tightness" and outputs it as a set of doctor query statements.
[0135] In a blood quality analysis scenario, if the "sickle cell" image label does not match clinical information, this submodule will reconstruct the review instructions for the laboratory physician, for example, generating: "[Review Instruction]: The dominant node image shows sickle cell characteristics, which is inconsistent with the clinical diagnosis. Please review the patient's electronic medical record and focus on verifying the 'hereditary disease history' and 'hemoglobin electrophoresis' test results."
[0136] In the vaccine disease control imaging analysis scenario, if the "particle aggregation" image label does not match the production log, this submodule will reconstruct the investigation instruction from the quality control department, for example, generating: "[Investigation Instruction]: The electron microscopy image of batch number [XXX] sample shows particle aggregation, which does not match the parameters in the production log. Please retrieve the detailed sensor data for this batch's 'buffer solution pH value,' 'post-filling temperature,' and 'transport vibration record.'"
[0137] Table 3: Trend of Missing Keywords
[0138] Node number Number of missing keywords Corresponding anatomical label Number of unaligned words Semantic directional trend value Node 1 2 hilum 1 0.500 Node 2 1 right lower lobe 1 1.000 Node 3 3 Tracheal branches 2 0.666
[0139] As shown in Table 3, the pairing of missing keywords with anatomical structures at different nodes reflects the degree of semantic incompleteness. The higher the trend value, the greater the disconnect between the structural information of the region and the semantics of the symptoms.
[0140] Please see Figure 1 and Figure 2 The factor construction module includes:
[0141] The semantic rhythm extraction submodule acquires the patient's voice signal in real time through the physician's question set. By extracting the time interval and pronunciation duration between keywords in the voice, it segments the speech rate change and interval time series, calculates the speech rate slope and rhythm fluctuation degree of each segment, normalizes the variation amplitude of the speech rhythm on the time axis, and establishes the speech rhythm fluctuation rate value.
[0142] The specific formula for detecting abrupt changes in speech rhythm is as follows:
[0143] ;
[0144] Calculate the trend value of sudden changes in speech rhythm;
[0145] in, For the first The absolute value of the rate of change of the slope of the segmental speech rhythm. For the first The speech rate slope of a segment of speech, For the first The speech rate slope of a segment of speech, For the first Section and the The start time difference between segments For the first The average syllable interval length of keywords in the paragraph For the first The average duration of pronunciation of keywords in the paragraph For the first The average speech energy density of a segment containing keywords. This is the index number of the current speech segment in the rhythm sequence. This is an index variable used to iterate through the speech segment numbers within the summation interval;
[0146] formula:
[0147] ;
[0148] Detailed explanation of the formula and its calculation derivation:
[0149] The formula is used to calculate the rate of change of speech rate slope between adjacent speech segments. The results are used to identify abrupt changes in speech rhythm;
[0150] Parameter meanings and settings:
[0151] : No. The speech rate slope of a speech segment, measured in syllables per square second, is obtained by linear regression analysis of the syllable pronunciation time within the speech segment, and is set as follows: ;
[0152] : No. The speech rate slope of a segment of speech, measured in syllables per square second, is set to... ;
[0153] : No. Section and the The start time difference between segments is obtained from the start time record and is set as follows: Second;
[0154] : , set as Second;
[0155] : No. The average syllable interval duration of keywords in a segment is set to . Second;
[0156] : No. The average pronunciation duration of keywords in the segment is set to Second;
[0157] : No. The average pronunciation duration of keywords in the segment is set to Second;
[0158] : No. The average speech energy density of keywords in the segment is set to decibel;
[0159] : No. The average speech energy density of keywords in the segment is set to decibel;
[0160] Substitute the parameters into the formula to calculate:
[0161] ;
[0162] result Indicates the first sound Section and the There are significant changes in the speech rate slope between segments. This change value is used to represent the abrupt change trend of speech rhythm and serves as the basis for calculating the speech rhythm fluctuation rate value.
[0163] The physiological signal recognition submodule extracts the electrocardiogram and electrodermal signals for the corresponding time period based on the time range covered by the speech rhythm fluctuation rate value, performs amplitude difference and polarity direction detection on the signal waveform, calculates the number of jumps and the frequency of reversals per unit time, groups and classifies them according to jump amplitude and reversal frequency, and obtains the jump ratio value of the fluctuation signal.
[0164] The physiological signal recognition submodule extracts the raw waveform data of two signal sources, electrocardiogram (ECG) and electrodermal conductance (TEA), within the same time period based on the specific time segment covered by the speech rhythm fluctuation value. It segments each signal waveform at 0.01-second intervals, first calculating the adjacent amplitude difference sequence for each signal sample within each time segment, and then performing operations... The amplitude change is measured. If the difference exceeds a set amplitude threshold of 40μV, it is recorded as a jump event. For example, in an electrodermal signal, adjacent values of 320μV and 380μV have a difference of 60μV, which is greater than the threshold, so a jump is marked and the time point is recorded. Next, the signal polarity is detected. When two consecutive sampling points have opposite signs, it is considered a polarity reversal event. For example, in an electrocardiogram signal, if the potential jumps from +120μV to -100μV within a certain time period, a reversal occurs, and the reversal frequency is recorded. The frequency value is obtained by accumulating the number of reversals per second. Then, the frequency of jump events and reversal events is calculated together with the total number of samples to obtain the value per unit time. The number of transitions and the reversal frequency are used. For example, in a certain period of sampling 1000 points, the number of transitions is 12, and the reversal frequency is 3.2 times / second. Then, the transition amplitude and reversal frequency are used as input parameters to divide the data into three classification intervals. For example, a transition amplitude greater than 100μV is considered a high-amplitude transition, and a reversal frequency greater than 4 times / second is considered a high-frequency reversal. If the current data transitions to 120μV and the reversal frequency is 5.1 times / second, it is classified as "high amplitude-high frequency". Finally, after normalizing each group of data, the transition ratio is calculated by dividing the normalized transition amplitude value by the normalized reversal frequency value. For example, if the normalized values are 0.85 and 4.1, the transition ratio is... Record this value in the signal time period labeling information;
[0165] In vaccine disease control imaging analysis scenarios, this module is adapted as a process anomaly factor construction module. It does not process physiological signals, but instead retrieves multi-dimensional sensor data streams (such as temperature, pH, and vibration) from the corresponding batch production process based on the [investigation instructions] issued by the reconstruction guidance module. By analyzing outliers (such as temperature exceeding the threshold 3σ) or mutation rates (such as a sharp change in pH slope) of each sensor data during the production time period corresponding to the dominant "particle aggregation" node, process anomaly co-occurrence factors are constructed. For example, it identifies a brief 2°C temperature spike in the cold storage at time point T1, accompanied by an abnormal decrease in the vibration frequency of the mixing tank.
[0166] Table 4 Characteristics of Jump Signals
[0167] Time period number Jump amplitude (μV) Reversal frequency (times / second) Number of jumps Fluctuation signal jump ratio Period 1 85 3.2 12 0.267 Period 2 120 5.1 25 0.208 Time period 3 45 1.8 7 0.257
[0168] As shown in Table 4, the characteristic data of the signal in terms of transition amplitude and reversal frequency at different time periods can be used for grouping and ratio judgment analysis.
[0169] The facial collaborative mapping submodule calls the time point information marked by the jump ratio of the fluctuation signal, extracts the facial image sequence of the corresponding time period and divides the region, tracks the pixel intensity change in the region, calculates the muscle tension change, identifies the mutation points in facial muscle tension data, speech rhythm and various physiological signals, analyzes the overlapping segments of the three types of mutation points on the time axis and counts the co-occurrence frequency, and establishes a set of abnormal co-occurrence factors.
[0170] After receiving the jump ratio annotation data for the aforementioned time period, the facial collaborative mapping submodule slices the facial image sequence according to the corresponding time period, capturing 25 frames of video images per second. It then divides the facial region in the image into multiple blocks, such as the forehead, corners of the eyes, corners of the mouth, and cheeks. Within each region, it tracks pixel intensity changes, recording the grayscale changes in consecutive frames and calculating the gradient sequence for each pixel. For example, if the grayscale value of a point in the corner of the mouth region is 120, 125, 130, 122, and 118 within 5 frames, it calculates the difference between adjacent frames and records the direction of change. If the grayscale fluctuation direction of most points in the same region is synchronized, it is considered a muscle tension fluctuation, further analyzed in the local area. Within a domain, a grayscale change threshold is set, such as 10 grayscale levels. If the proportion of pixels exceeding the threshold exceeds 50%, it is considered that there is a significant tension change in that area. Subsequently, the muscle tension change points on the time axis are matched with the rhythmic abrupt change points and the jump time points of physiological signal fluctuations in speech rhythm recognition. The co-occurrence condition is judged based on the time interval between two adjacent abrupt change points being less than 0.2 seconds. After a successful match, the number of co-occurrences is accumulated, and the co-occurrence frequency of the three types of abrupt change points in the same time period is counted. For example, if speech + physiological signal co-occurs 3 times, speech + facial occurs 2 times, and all three types co-occur 1 time within 5 seconds, all co-occurrence relationships are stored in the abnormal co-occurrence factor set, and finally a structured data set containing time points, co-occurrence types, and region names is formed.
[0171] In scenarios such as blood quality analysis where doctor-patient communication of diagnostic results is required, when a doctor informs a patient that "your blood test results suggest possible aplastic anemia," this submodule can analyze the patient's facial micro-expressions (such as frowning and drooping eyelids) and vocal responses (such as slower speech and increased pauses) in real time. If strong co-occurrence factors of negative emotions are detected, the system can prompt the doctor to pause in-depth explanation and instead provide emotional reassurance and supportive information.
[0172] For non-interactive vaccine analysis scenarios, this submodule is skipped, and the results from the process anomaly factor construction module are directly entered into the path reordering module.
[0173] Please see Figure 1 and Figure 2 The path reordering module includes:
[0174] The symptom tag extraction submodule extracts the auxiliary consultation path through the abnormal co-occurrence factor set, identifies the symptom tag corresponding to each node in the path, records the arrangement order of the tags in the original path, constructs a set of tags and order positions, and generates tag path order values.
[0175] After obtaining the set of abnormal co-occurrence factors, the symptom tag extraction submodule extracts an auxiliary consultation path based on the cross-information of time points and signal types recorded in the set. This path is preset as a standard inquiry sequence constructed based on the evolutionary process of respiratory diseases, initially containing nodes in the order of ["cough", "fever", "chest pain", "hemoptysis"]. The system reads the abnormal time periods and dominant symptom regions in the co-occurrence factors, matching the N1 node with the most mutation overlap with "chest pain". Subsequently, the system sequentially parses the time points and symptom tags associated with each node, ranking "chest pain" by recording co-occurrence frequency. The first label is placed first, and the other labels are sorted according to their co-occurrence frequency and original path order. For example, "cough" was originally first, but its co-occurrence frequency is lower than "chest pain" but higher than "fever", so "cough" is adjusted to the second position. Finally, a matching set of labels and node order is constructed, such as ["chest pain"-1, "cough"-2, "fever"-3, "hemoptysis"-4]. The original path position and the new position after adjustment of each label are recorded in the fields "original position of label" and "position after path adjustment" respectively. In the example, "chest pain" was originally the 3rd position and was adjusted to the 1st position. The change in order is recorded and the set of label path order values is output.
[0176] Table 5 Tag Path Sequence Table
[0177] Path node number Symptom labels Original location of the tag Location after path adjustment N1 Chest pain 3 1 N2 cough 1 2 N3 fever 2 3 N4 Hemoptysis 4 4
[0178] As shown in Table 5, the new label order values generated after the abnormal co-occurrence factor triggers path rearrangement record the dynamic displacement process of the labels.
[0179] The assisted consultation path refers to a standardized question sequence pre-set by the system and constructed based on a clinical knowledge base and disease evolution process. It is used to provide physicians with a structured symptom label order and question template during intelligent medical interaction, and to assist in the systematic and comprehensive collection of patient status. As a basic interaction framework, the path contains multiple symptom nodes arranged according to disease characteristics, and their order is dynamically adjusted according to the dominant node of the image and the direction of semantic keywords.
[0180] The semantic matching determination submodule calls the tag path order value and matches the dominant node of the image with the semantic direction content of the keywords. It calculates the number of matches based on the position of the tag in the image region and directional terms, builds a tag and node matching table, filters the priority by quantity, and generates path semantic matching coefficients.
[0181] After obtaining the aforementioned label path order values, the semantic matching determination submodule reads the label content and semantic keyword set of the lesion area in the dominant node of the image, counts the total frequency of each label term in the node area and records the matching position. For example, in the dominant node image, "chest pain" appears in the high-density area of the hilum and corresponds to the speech keyframe time period, "cough" appears in the lung field texture enhancement area, and "fever" and "hemoptysis" do not directly match. According to the semantic distance comparison, "chest pain" matches 2 times, "cough" matches 1 time, and the rest are 0 times. Construct a matching frequency table of labels and dominant nodes, and then sort them in descending order of frequency. The labels with a matching frequency greater than 0 are mapped to priority according to the number of matches, such as "chest pain" as priority 1, "cough" as priority 2, and the rest as 0. Generate a label matching priority table of dominant nodes, and extract the path semantic matching coefficient array accordingly, such as ["chest pain"-2, "cough"-1, "fever"-0, "hemoptysis"-0]. Finally, output the semantic matching strength ranking result for path adjustment reference.
[0182] The interaction instruction generation submodule adjusts the response order of symptom tags in the interaction process according to the path semantic matching coefficient, rearranges the tag call structure in the original path according to the matching priority, establishes the mapping relationship between tag nodes and response content, and generates a dynamic interactive decision instruction set for medical scenarios.
[0183] The interaction instruction generation submodule, based on the aforementioned semantic matching coefficients, uses priority ranking as a reference weight to dynamically rearrange the tag structure in the original path according to priority. It calls the tag priority array to adjust the response order of the initial order. For example, if the original path is ["cough", "fever", "chest pain", "hemoptysis"], the rearranged path is ["chest pain", "cough", "fever", "hemoptysis"]. Then, it establishes a mapping relationship between each tag and the interaction response content. For example, the interaction content corresponding to the tag "chest pain" is "Do you feel chest tenderness or dull pain?". The system updates the interaction order according to the rearrangement result, marks the relationship between node number and content segment combination, and generates an interaction decision instruction set structure, such as {"N1": [tag-"chest pain", content-"Do you feel chest tenderness or dull pain?"], "N2": [tag-"cough", content-"Do you have phlegm or dry cough?"]...}. Finally, this set is output as the dynamically adjusted query strategy instruction.
[0184] In a blood quality analysis scenario, if a patient is detected to have strong negative emotional factors related to the diagnosis of "aplastic anemia," the interaction instruction generation submodule will rearrange the doctor's communication path. The original path is {[}1. Inform the patient of the diagnosis, 2. Explain the cause, 3. Discuss the treatment plan{]}, and the rearranged path is {[}1. Inform the patient of the diagnosis, 2. Provide psychological support channels, 3. Ask the patient about their current feelings, 4. Briefly explain the subsequent examination steps{]}. The generated instruction set is: {"N1": {[} Tag - "Psychological Support", Content - "I know this news is hard to accept, we have professional psychological counselors who can provide help."{]}, "N2": {…}}.
[0185] In the vaccine disease control imaging analysis scenario, based on the process anomaly factor (cold storage temperature surge), this submodule rearranges the standard quality control process. The original process is {[}1. Re-inspect samples, 2. Record results, 3. Archive {]}, and the rearranged process is {[}1. Immediately isolate all products in this batch, 2. Retrieve complete historical temperature data of cold storage [number], 3. Arrange to inspect the refrigeration unit of cold storage [number], 4. Expand the sampling scope to adjacent batches {]}. The generated dynamic interactive decision instruction set is: {“Instruction 1”: {[} operation - “Isolate batch”, object - “[batch number XXX]” {]}, “Instruction 2”: {[} operation - “Retrieve data”, object - “cold storage [number]” {]}…}.
[0186] Please see Figure 3 ,
[0187] S1: By extracting grayscale, edge and texture features from medical images, regions with closed contours and complex textures are selected as lesion candidate areas. The morphological change trend of regions in consecutive frames is calculated to determine the degree of structural mutation. The differences in image structure in the frame sequence are optimized, and image frames with prominent structural changes are selected to generate dominant image nodes.
[0188] S2: Based on the dominant node of the image, obtain the annotation information of the corresponding image frame and the symptom keywords in the case text, analyze the semantic matching relationship between the keywords and lesion labels, filter out keywords that do not appear in the semantic corresponding items in the labels, identify the missing positions in the semantic chain of the text, generate a set of keywords that cannot be associated, and construct semantically uncovered label data.
[0189] S3: By using the dominant nodes in the image and semantically uncovered identifier data, locate the contextual logical position of keywords, calculate the topic association direction in the symptom expression sequence, filter sentence structure templates that match the semantic direction, adjust the word order and position combination of keywords in the sentence, analyze the grammatical integrity of the generated sentences, construct a set of questions, and output a set of physician inquiry statements.
[0190] S4: Through the physician's question set, monitor the patient's voice, physiological signals, and facial expressions in real time, analyze the speech rate change trend between voice keywords, compare the locations of frequency abrupt changes and waveform changes in physiological signals, screen for abnormal tension transition areas in image sequences, integrate abnormal synchronous events, and generate a set of abnormal co-occurrence factors.
[0191] S5: By using the set of abnormal co-occurrence factors, mark the time points of abnormal statements and the label sequence in the set of physician inquiry statements, analyze the correspondence between the symptoms guided by the statements and the labels of the dominant nodes in the images, analyze the consistency between semantic direction and label logic, adjust the sorting structure of statements in the inquiry process, optimize the order of interaction paths, and construct a dynamic interactive decision instruction set for medical scenarios.
[0192] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0193] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0194] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0195] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0196] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0197] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0198] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0199] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0200] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0201] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0202] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A dynamic interactive decision-making system for medical scenarios based on multimodal perception, characterized in that, The system includes: The master node identification module extracts grayscale, edge and texture features from medical images, filters candidate lesion areas, calculates inter-frame edge changes and texture differences, analyzes image structural saliency, selects image frames as image master nodes and passes them to the coverage determination module and reconstruction guidance module. The coverage judgment module extracts the patient's medical record text based on the dominant node of the image, identifies symptom keywords and compares them with the lesion labels corresponding to the dominant node of the image, generates semantically uncovered identification data, and transmits it to the reconstruction guidance module. The reconstruction guidance module extracts uncovered keywords and locates missing semantic directions through the dominant nodes of the image and the semantically uncovered identifier data, reconstructs the physician's inquiry statement, outputs the physician's inquiry statement set and passes it to the factor construction module. The factor construction module monitors the patient's voice, physiological signals, and facial expressions in real time through the physician's question set. It extracts the keyword intervals and speech rate changes in the voice, combines physiological indicators to locate signal segments, identifies waveform inversion and amplitude jump features, and combines facial expression muscle tension changes to construct an abnormal co-occurrence factor set and transmit it to the path rearrangement module.
2. The dynamic interactive decision-making system for medical scenarios based on multimodal perception according to claim 1, characterized in that, The image dominant node includes image frame number, dominant region location, and structural salience score; the semantically uncovered identifier data includes a list of missing labels, keyword matching degree, and semantic coverage rate; the physician inquiry statement set includes target symptom words, reconstructed statement templates, and candidate guiding words; and the abnormal co-occurrence factor set includes voice emotion mutation points, abnormal facial expression patterns, and abnormal physiological signal features.
3. The dynamic interactive decision-making system for medical scenarios based on multimodal perception according to claim 1, characterized in that, The master node identification module includes: The candidate screening submodule extracts grayscale, edge and texture features from medical images, combines boundary information to extract region contours, filters regions with abnormal grayscale distribution and constructs a set of edge contours, compares the consistency between contours and gradients, identifies potential lesion regions, and generates candidate region consistency coefficient values. The edge texture analysis submodule obtains the edge distribution map of the corresponding region in consecutive image frames based on the candidate region consistency coefficient value, calculates the edge intensity difference and texture offset rate between frames, filters out regions whose edge offset amplitude and texture distribution change rate exceed the edge offset reference range and texture fluctuation reference interval, and obtains the joint offset rate value of structural change. The edge offset reference range is defined as the upper and lower boundaries of normal edge changes by extracting the edge intensity difference of non-lesion areas in consecutive image frames, statistically analyzing the average level and fluctuation amplitude. The texture fluctuation reference range is defined by extracting the gray-level co-occurrence matrix energy, contrast and local binary mode differences between frames in non-lesion areas, and setting upper and lower thresholds based on common variation ranges. The dominant frame recognition submodule extracts the structural contour and internal texture aggregation degree of the corresponding region based on the joint offset rate value of the structural change, calculates the structural density and compares it with the salience density benchmark value, identifies the image frames that meet the deviation standard and uses them as dominant nodes, and establishes the image dominant node. The specific formula for extracting the structural contour and internal texture aggregation degree of the corresponding region is as follows: ; Calculate the structural density value ; in, Represents the structural density value. Represents the total number of pixels in the extracted area. Representing the The grayscale gradient value of each pixel in the horizontal structural gradient direction. Representing the The grayscale gradient value of each pixel in the vertical structural gradient direction. Representing the The gradient weight values of the texture clustering degree of each pixel. This represents the structural variance adjustment coefficient, used to dynamically adjust the texture structure response weights. Representing the The absolute deviation of the grayscale value of each pixel within the local window. This represents the average local structural density value calculated from the joint structural texture features of all pixels in the extracted region. The index number of the pixel. To extract the total number of pixels within the region; The saliency density benchmark value is a standard set by statistically analyzing the distribution range of high-density areas based on the degree of aggregation of structural and textural features in known lesion areas. The deviation standard is the minimum identifiable boundary set by comparing the numerical differences in the structure and texture aggregation scores between the dominant region and the ordinary region.
4. The dynamic interactive decision-making system for medical scenarios based on multimodal perception according to claim 3, characterized in that, The coverage determination module includes: The dominant node extraction submodule identifies the structural features of the corresponding region based on the dominant node of the image. By analyzing the grayscale, edge and texture features of the lesion candidate area, it combines, classifies and groups each feature by type. Based on the type grouping results, it outputs the lesion type label of each region, establishes the corresponding set of nodes and lesions, and generates the image lesion label comparison quantity. The symptom keyword recognition submodule calls the image lesion label comparison quantity and obtains the patient's medical record text data. It extracts symptom description words from the text through word segmentation, filters keywords corresponding to disease symptoms, constructs a term set based on the frequency of keyword occurrence and part-of-speech structure, and generates symptom keyword extraction quantity. The semantic comparison and judgment submodule compares the content of the image lesion label control quantity with the extracted symptom keywords, constructs a term missing index based on the semantic matching degree of the set, calculates the matching gap rate between the number of missing terms and the total number of labels, establishes the number of uncovered terms and the corresponding node identifier, and generates semantic uncovered identifier data.
5. The dynamic interactive decision-making system for medical scenarios based on multimodal perception according to claim 4, characterized in that, The reconstruction guidance module includes: The keyword extraction submodule extracts the lesion content of the associated nodes in the missing identifiers through the dominant nodes of the image and the semantically uncovered identifier data, locates the set of missing terms corresponding to the identifier number and extracts keywords, establishes a keyword list after screening non-symptom terms, and generates the number of uncovered keywords. The semantic direction determination submodule calls the number of uncovered keywords, obtains the corresponding anatomical structure labels and spatial orientations in the dominant nodes of the image, filters the misaligned content of missing terms in the label set, establishes the directional pairing relationship between missing terms and structural labels, calculates the frequency ranking, and generates missing semantic direction trend values. The query reconstruction submodule matches the tag templates in the semantic library based on the missing semantic direction trend value, extracts high-frequency structural keywords and combines them with the corresponding query sentence patterns, constructs a query sentence rearrangement sequence driven by missing keywords and generates sentence pattern groups in sequence, and generates a set of physician query sentences.
6. The dynamic interactive decision-making system for medical scenarios based on multimodal perception according to claim 5, characterized in that, The factor construction module includes: The semantic rhythm extraction submodule acquires the patient's voice signal in real time through the physician's question set. By extracting the time interval and pronunciation duration between keywords in the voice, it segments the speech rate change and interval time sequence, calculates the speech rate slope and rhythm fluctuation degree of each segment, normalizes the amplitude of speech rhythm variation on the time axis, and establishes the speech rhythm fluctuation rate value. The physiological signal recognition submodule extracts electrocardiogram and electrodermal signals for the corresponding time period based on the time range covered by the speech rhythm fluctuation rate value, performs amplitude difference and polarity direction detection on the signal waveform, calculates the number of jumps and the frequency of reversals per unit time, groups and classifies them according to jump amplitude and reversal frequency, and obtains the fluctuation signal jump ratio value. The facial collaborative mapping submodule calls the time point information marked by the fluctuation signal jump ratio, extracts the facial image sequence of the corresponding time period and divides the region, tracks the pixel intensity change in the region, calculates the muscle tension change, identifies the mutation points in facial muscle tension data, speech rhythm and various physiological signals, analyzes the overlapping segments of the three types of mutation points on the time axis and counts the co-occurrence frequency, and establishes an abnormal co-occurrence factor set.
7. The dynamic interactive decision-making system for medical scenarios based on multimodal perception according to claim 6, characterized in that, The specific formula for detecting abrupt changes in speech rhythm is as follows: ; Calculate the trend value of sudden changes in speech rhythm; in, For the first The absolute value of the rate of change of the slope of the segmental speech rhythm. For the first The speech rate slope of a segment of speech, For the first The speech rate slope of a segment of speech, For the first Section and the The start time difference between segments For the first The average syllable interval length of keywords in the paragraph For the first The average duration of pronunciation of keywords in the paragraph For the first The average speech energy density of a segment containing keywords. This is the index number of the current speech segment in the rhythm sequence. This is an index variable used to iterate through the speech segment numbers within the summation interval.
8. The dynamic interactive decision-making system for medical scenarios based on multimodal perception according to claim 1, characterized in that: The path reordering module extracts the symptom label order in the auxiliary consultation path through the abnormal co-occurrence factor set, matches the semantic direction of the labels with the dominant nodes of the image and keywords, reorders the interactive inquiry order according to the matching situation, adjusts the execution order of the physician's interactive response, and generates a dynamic interactive decision instruction set for the medical scenario. The dynamic interactive decision-making instruction set for medical scenarios includes query priority order, tag semantic mapping path, and interactive response adjustment scheme.
9. The dynamic interactive decision-making system for medical scenarios based on multimodal perception according to claim 8, characterized in that, The path reordering module includes: The symptom tag extraction submodule extracts the auxiliary consultation path through the abnormal co-occurrence factor set, identifies the symptom tag corresponding to each node in the path, records the arrangement order of the tags in the original path, constructs a tag-order position correspondence set, and generates tag path order value; The assisted consultation path refers to a standardized inquiry sequence pre-set by the system and constructed based on a clinical knowledge base and disease evolution process. It is used to provide physicians with a structured symptom label order and inquiry template during intelligent medical interaction, and to assist in the systematic and comprehensive collection of patient status. The path, as a basic interaction framework, contains multiple symptom nodes arranged according to disease characteristics, and their order is dynamically adjusted according to the dominant nodes of the image and the direction of semantic keywords. The semantic matching determination submodule calls the tag path order value and matches the dominant node of the image with the semantic direction content of the keyword. It calculates the number of matches based on the position of the tag in the image region and the directional terms, constructs a tag and node matching table, filters the priority by quantity, and generates the path semantic matching coefficient. The interaction instruction generation submodule adjusts the response order of symptom tags in the interaction process according to the path semantic matching coefficient, rearranges the tag call structure in the original path according to the matching priority, establishes the mapping relationship between tag nodes and response content, and generates a dynamic interactive decision instruction set for medical scenarios.
10. A dynamic interactive decision-making method for medical scenarios based on multimodal perception, characterized in that, The method is used to implement the multimodal perception-based dynamic interactive decision-making system for medical scenarios as described in any one of claims 1-9, the method comprising: S1: By extracting grayscale, edge and texture features from medical images, regions with closed contours and complex textures are selected as lesion candidate areas. The morphological change trend of regions in consecutive frames is calculated to determine the degree of structural mutation. The differences in image structure in the frame sequence are optimized, and image frames with prominent structural changes are selected to generate dominant image nodes. S2: Based on the dominant node of the image, obtain the annotation information of the corresponding image frame and the symptom keywords in the case text, analyze the semantic matching relationship between the keywords and lesion labels, filter out keywords that do not appear in the semantic corresponding items in the labels, identify the missing positions in the text semantic chain, generate a set of keywords that cannot be associated, and construct semantically uncovered identification data. S3: Using the dominant node of the image and the semantically uncovered identifier data, locate the contextual logical position of the keywords, calculate the topic association direction in the symptom expression sequence, filter sentence structure templates that match the semantic direction, adjust the word order and position combination of keywords in the sentence, analyze the grammatical integrity of the generated sentences, construct a set of questions, and output a set of physician inquiry statements. S4: Through the physician's question set, monitor the patient's voice, physiological signals, and facial expressions in real time, analyze the speech rate change trend between voice keywords, compare the positions of frequency abrupt changes and waveform changes in physiological signals, screen for abnormal tension transition areas in image sequences, integrate abnormal synchronization events, and generate a set of abnormal co-occurrence factors. S5: Using the set of abnormal co-occurrence factors, mark the time points of abnormal statements and the label sequence in the physician inquiry statement set, analyze the correspondence between the symptoms guided by the statements and the labels of the dominant nodes in the images, analyze the consistency between semantic direction and label logic, adjust the sorting structure of statements in the inquiry process, optimize the interaction path order, and construct a dynamic interactive decision instruction set for medical scenarios.