Intelligent central control terminal control method, device and equipment based on multi-modal large model
By processing users' electronic medical records and thyroid cancer knowledge graphs in a multimodal manner, a target follow-up plan is generated. The intelligent central control terminal actively assists users in completing rehabilitation tasks, solving the problem of operational convenience for the elderly or blind and other groups with limited physiological functions, and improving operational convenience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-23
- Publication Date
- 2026-04-17
AI Technical Summary
Existing intelligent central control terminal control methods are not very convenient for people with limited physiological functions, such as the elderly or the blind, and require complex touch interaction modes, resulting in cumbersome operation processes.
A smart central control terminal control method based on a multimodal large model is adopted. By performing node fusion, feature node mapping, node attention aggregation and gating enhancement processing on the user's electronic medical record and thyroid cancer knowledge graph, a target follow-up plan is generated, and the smart central control terminal is controlled to actively assist the user in completing rehabilitation tasks.
It reduces user interaction steps, improves operational convenience, and enables the intelligent central control terminal to proactively assist users in completing rehabilitation plans, thereby enhancing the ease with which users can control the intelligent central control terminal for rehabilitation training.
Smart Images

Figure CN121885140A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of computer technology, and specifically to a control method, apparatus, and device for an intelligent central control terminal based on a multimodal large model. Background Technology
[0002] The intelligent central control terminal control method assists users in home rehabilitation by controlling the intelligent central control terminal. While intelligent central control terminals can assist users in home rehabilitation, methods relying on visual observation, text input, or complex touch controls pose significant barriers to use for the elderly or blind, and other groups with limited physiological functions. Therefore, there is an urgent need for an intelligent assistance solution that breaks through traditional interaction modes, enabling the system to assist users in home rehabilitation through barrier-free functionality. Currently, the common approach to controlling intelligent central control terminals is for users to actively interact with the terminal based on a generated rehabilitation plan to complete the corresponding rehabilitation tasks.
[0003] However, when using the above method to control the intelligent central control terminal, the following technical problems often occur: Users need to use complex touch and other interaction modes to control the smart central control terminal to perform auxiliary tasks. The operation process is relatively cumbersome, resulting in poor convenience for users when controlling the smart central control terminal for rehabilitation training.
[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion that follows. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0006] Some embodiments of this disclosure propose intelligent central control terminal control methods, devices, electronic devices, and computer-readable media based on multimodal large models to solve one or more of the technical problems mentioned in the background section above.
[0007] In a first aspect, some embodiments of this disclosure provide a control method for an intelligent central control terminal based on a multimodal large model. The method includes: responding to a task execution signal sent by the intelligent central control terminal; performing node fusion processing on pre-acquired user electronic medical records and a thyroid cancer knowledge graph to update the thyroid cancer knowledge graph; performing feature node mapping processing on the user electronic medical records based on the updated thyroid cancer knowledge graph to obtain a user knowledge subgraph; performing node attention aggregation processing on the user knowledge subgraph to obtain each meta-path subgraph and each subgraph node feature information; performing gating enhancement processing on the user electronic medical records, the meta-path subgraphs, and the subgraph node feature information to obtain user feature information; performing scheme matching and adaptive optimization processing on the user feature information based on a multimodal large model used for processing thyroid cancer data to obtain a target follow-up scheme; and controlling the intelligent central control terminal to perform task assistance operations based on the target follow-up scheme.
[0008] Secondly, some embodiments of this disclosure provide a control device for an intelligent central control terminal based on a multimodal large model. The device includes: a node fusion unit configured to, in response to receiving a task execution signal sent by the intelligent central control terminal, perform node fusion processing on pre-acquired user electronic medical records and a thyroid cancer knowledge graph to update the thyroid cancer knowledge graph; a feature node mapping unit configured to, based on the updated thyroid cancer knowledge graph, perform feature node mapping processing on the aforementioned user electronic medical records to obtain a user knowledge subgraph; a node attention aggregation unit configured to, perform node attention aggregation processing on the aforementioned user knowledge subgraph to obtain various meta-path subgraphs and node feature information of each subgraph; a gating enhancement unit configured to, perform gating enhancement processing on the aforementioned user electronic medical records, the aforementioned various meta-path subgraphs, and the aforementioned node feature information of each subgraph to obtain user feature information; an adaptive optimization unit configured to, based on a multimodal large model used for processing thyroid cancer data, perform scheme matching and adaptive optimization processing on the aforementioned user feature information to obtain a target follow-up scheme; and a control unit configured to, based on the aforementioned target follow-up scheme, control the aforementioned intelligent central control terminal to perform task auxiliary operations.
[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.
[0011] The above embodiments of this disclosure have the following beneficial effects: the intelligent central control terminal control method based on a multimodal large model in some embodiments of this disclosure can improve operational convenience. The reason for poor operational convenience is that users need to use complex touch and other interactive modes to control the intelligent central control terminal to perform auxiliary tasks, making the operation process cumbersome and resulting in poor convenience for users controlling the intelligent central control terminal for rehabilitation training. Based on this, the intelligent central control terminal control method based on a multimodal large model in some embodiments of this disclosure firstly, in response to receiving a task execution signal sent by the intelligent central control terminal, performs node fusion processing on the pre-acquired user electronic medical records and thyroid cancer knowledge graph to update the thyroid cancer knowledge graph. This allows for the updating of the thyroid cancer knowledge graph. Secondly, based on the updated thyroid cancer knowledge graph, the user electronic medical records undergo feature node mapping processing to obtain a user knowledge subgraph. This allows for the obtaining of a user knowledge subgraph. Then, the user knowledge subgraph undergoes node attention aggregation processing to obtain each meta-path subgraph and each subgraph node feature information. Therefore, the feature information of each meta-path subgraph and each subgraph node can be obtained. Then, gating enhancement processing is performed on the aforementioned user electronic medical record, the aforementioned meta-path subgraph, and the aforementioned subgraph node feature information to obtain user feature information. Then, based on a multimodal large model used to process thyroid cancer data, the aforementioned user feature information is subjected to scheme matching and adaptive optimization processing to obtain a target follow-up scheme. Finally, based on the aforementioned target follow-up scheme, the aforementioned intelligent central control terminal is controlled to perform task-assisted operations. Because the intelligent central control terminal can actively assist the user in completing the rehabilitation plan based on the generated target follow-up scheme, rather than passively waiting for user interaction, the number of interaction steps required by the user is reduced, improving the ease of operation for the user when controlling the intelligent central control terminal for rehabilitation training. Attached Figure Description
[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0013] Figure 1 This is a flowchart of some embodiments of the intelligent central control terminal control method based on a multimodal large model according to the present disclosure; Figure 2 These are schematic diagrams of some embodiments of the intelligent central control terminal control device based on a multimodal large model according to this disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure; Figure 4 This can be a flowchart for training a large multimodal model. Detailed Implementation
[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0020] Figure 1 A flowchart 100 is shown, illustrating some embodiments of the intelligent central control terminal control method based on a multimodal large model according to this disclosure. This intelligent central control terminal control method based on a multimodal large model includes the following steps: Step 101: In response to receiving the task execution signal sent by the intelligent central control terminal, perform node fusion processing on the pre-acquired user electronic medical records and thyroid cancer knowledge graph to update the thyroid cancer knowledge graph.
[0021] In some embodiments, the execution subject (e.g., a computing device) of the intelligent central control terminal control method based on a multimodal large model can, in response to receiving a task execution signal sent by the intelligent central control terminal, perform node fusion processing on the pre-acquired user electronic medical records and thyroid cancer knowledge graph to update the thyroid cancer knowledge graph.
[0022] The aforementioned intelligent central control terminal can be a device capable of communicating with various terminal devices and controlling these terminal devices to perform specific operations. For example, the intelligent central control terminal can be a smart gateway. Each of the aforementioned terminal devices can be a terminal capable of performing specific functions according to instructions. The aforementioned terminal devices can include, but are not limited to, smart bracelets and smart cameras. The aforementioned task execution signal can be a signal used to prompt the executing entity to begin processing data.
[0023] The aforementioned user electronic medical records can be the electronic medical records of the target user. These user electronic medical records may include surgical records, user pathology reports, pathology slide images, and preoperative laboratory data.
[0024] The aforementioned surgical records may be the surgical records of thyroid cancer surgery performed on the target user. These surgical records may include, but are not limited to, the surgery date. The target user may be a user who has undergone thyroid cancer surgery. The aforementioned thyroid cancer surgery may be a procedure performed for thyroid cancer. For example, the aforementioned thyroid cancer surgery may be a thyroidectomy.
[0025] The aforementioned user pathology reports can be the pathology reports of the target users mentioned above.
[0026] The aforementioned pathological section images can be micrographs obtained after pathological sectioning of thyroid lesions removed from the target user's body.
[0027] The aforementioned preoperative laboratory data can be JSON formatted data corresponding to the physiological or biochemical indicators of the target user. This preoperative laboratory data may include complete blood count data, thyroid function data, tumor marker detection data, and liver and kidney function data.
[0028] The above-mentioned blood routine test data can be the data obtained after the user undergoes a blood routine test. The above-mentioned blood routine test data may include, but is not limited to: red blood cell count, white blood cell count, and platelet count.
[0029] The thyroid function data mentioned above can be used to determine the hormone levels of the thyroid gland. This thyroid function data may include, but is not limited to, the concentrations of thyroid-stimulating hormone (TSH), free thyroxine (FTH), and free triiodothyronine (TTI).
[0030] The tumor marker detection data mentioned above can be used to conveniently reflect the levels of thyroid cancer markers. These thyroid cancer markers can be biomarkers used to reflect the presence, type, burden, or risk of recurrence of thyroid cancer. The tumor marker detection data may include, but is not limited to: the concentrations of thyroglobulin, calcitonin, and carcinoembryonic antigen.
[0031] The aforementioned liver and kidney function data can be used to characterize the user's metabolic status. This data may include, but is not limited to, at least one of the following: alanine aminotransferase (ALT) concentration, aspartate aminotransferase (AST) concentration, serum creatinine concentration, and blood urea nitrogen concentration.
[0032] The aforementioned thyroid cancer knowledge graph can be a knowledge graph recording content related to thyroid cancer. The aforementioned thyroid cancer-related content can be text recording thyroid cancer information for each user diagnosed with thyroid cancer. The aforementioned thyroid cancer information can be information related to thyroid cancer. This thyroid cancer information may include, but is not limited to: the user's gender, age, surgery name, laboratory data, and anatomical location. The aforementioned thyroid cancer knowledge graph can include various nodes and edges. Each node in the aforementioned knowledge graph corresponds to a label, attribute, and node feature information. The aforementioned label can be a tag used to represent the category of the node. The aforementioned label can be, but is not limited to: gender, age, surgery name, laboratory data, and anatomical location. The aforementioned node feature information can be the feature vector of the entity corresponding to the node in the aforementioned thyroid cancer knowledge graph. The aforementioned laboratory data can be data obtained after human testing.
[0033] In some optional implementations of certain embodiments, the aforementioned executing entity may perform node fusion processing on the pre-acquired user electronic medical records and thyroid cancer knowledge graph through the following steps to update the thyroid cancer knowledge graph: The first step is to perform text preprocessing on the aforementioned surgical records and user pathology reports to obtain preprocessed text data. This preprocessed text data can be the text data obtained after preprocessing the aforementioned surgical records and user pathology reports.
[0034] In practice, firstly, the executing entity uses regular expressions to remove special characters from the surgical records and user pathology reports to perform text preprocessing, resulting in preprocessed surgical records and user pathology reports, respectively. The special characters can be Unicode characters other than letters, numbers, and Chinese characters. Secondly, the preprocessed surgical records and preprocessed user pathology reports can be combined into preprocessed text data.
[0035] The second step involves quantifying the microscopic features of the aforementioned pathological slide images to obtain image microscopic feature data. This image microscopic feature data can be data used to describe the morphological characteristics of cell nuclei in the pathological slide images. The image microscopic feature data may include area, perimeter, major axis length, minor axis length, eccentricity, and convex hull area ratio.
[0036] In practice, firstly, the aforementioned implementing entity can perform staining normalization processing on the aforementioned pathological slide images using staining normalization techniques, obtaining the staining-normalized pathological slide images as stained slide images. The aforementioned staining normalization techniques can be those capable of separating and reconstructing the dye components in the image. For example, the aforementioned staining normalization technique can be the Vahadane method.
[0037] Secondly, the stained slice image can be converted to the HSV color space using an image conversion function, resulting in the converted stained slice image as the image to be segmented. This image conversion function can be any function capable of converting an image to the HSV color space. For example, the image conversion function could be cv2.cvtColor().
[0038] Then, the image to be segmented can be processed using a threshold segmentation method to obtain a binary image corresponding to the image to be segmented, which serves as the tissue mask image. The threshold segmentation method can be any algorithm capable of converting an image into a binary image. For example, the threshold segmentation method could be the maximum inter-class variance method.
[0039] Then, a preset sliding window can be used to slide across the stained section image from left to right and from top to bottom with a preset step size to crop the stained section image into image blocks of the same size. The preset sliding window can be a pre-defined sliding window. The preset step size can be a pre-defined unit length that the preset sliding window moves across the stained section image each time.
[0040] Then, for each of the aforementioned slice image patches, the slice image patch can be input into a pre-trained pathological cell nucleus recognition model to obtain cell nucleus segmentation results, a horizontal distance channel matrix, and a vertical distance channel matrix. The cell nucleus segmentation result can be a two-dimensional matrix of the same size as the slice image patch, where each element can be used to represent the category of a pixel. Each element value in the cell nucleus segmentation result can be either 0 or 1. When an element value is "0", it indicates that the pixel category corresponding to that element value in the slice image patch is background. When an element value is "1", it indicates that the pixel category corresponding to that element value in the slice image patch is cell nucleus.
[0041] The aforementioned horizontal distance channel matrix can be the same size as the sliced image patch, where each element represents the horizontal distance of the corresponding pixel relative to the center of the cell nucleus to which the pixel belongs. For example, when the element value is "0.5", it can represent that the pixel corresponding to the element value in the sliced image patch is located to the right of the center of the cell nucleus to which the pixel belongs, at a distance of 0.5 times the radius of the cell nucleus. The aforementioned horizontal distance can be a distance in the horizontal direction. The aforementioned center of the cell nucleus can be the center point of the cell nucleus.
[0042] The aforementioned vertical distance channel matrix can be a matrix of the same size as the sliced image patch, where each element represents the vertical distance of the corresponding pixel relative to the center of the cell nucleus to which the pixel belongs. The aforementioned vertical distance can be a distance in the vertical direction.
[0043] The aforementioned pathological cell nucleus recognition model can be a neural network model that takes sliced image patches as input and outputs cell nucleus segmentation results, horizontal distance channel matrices, and vertical distance channel matrices. For example, the aforementioned pathological cell nucleus recognition model can be a pre-trained HoVer-Net model. The pre-training process involves fine-tuning the HoVer-Net model using an annotated set of sliced image patches and a composite loss function including cross-entropy loss and mean squared error loss. Then, the executing entity can use the watershed algorithm, utilizing the aforementioned horizontal distance channel matrices and vertical distance channel matrices, to perform instance segmentation processing on the cell nucleus segmentation results, obtaining a cell nucleus instance segmentation mask.
[0044] The cell nucleus instance segmentation mask can be a two-dimensional matrix of the same size as the sliced image patch, where each element can be used to represent the category of a pixel. The element values in the cell nucleus segmentation result can include, but are not limited to, at least one of the following: 0, 1, 2, 3. Specifically, when an element value is "0", it indicates that the pixel corresponding to that element value in the sliced image patch is classified as background. When an element value is "1", it indicates that the pixel corresponding to that element value in the sliced image patch is classified as the first cell nucleus. When an element value is "2", it indicates that the pixel corresponding to that element value in the sliced image patch is classified as the second cell nucleus, and so on.
[0045] Next, for each of the aforementioned slice image patches, the slice image patch and its corresponding cell nucleus instance segmentation mask can be input into a morphological function to obtain the morphological feature data corresponding to the slice image patch. Each of the aforementioned morphological feature data can be data used to characterize the morphological structure of the cell nucleus in the slice image patch. Each of the aforementioned morphological feature data corresponds to one cell nucleus. Each of the aforementioned morphological feature data can include area, perimeter, major axis length, minor axis length, eccentricity, and convex hull area ratio. The aforementioned morphological function can be a function capable of identifying the attributes of an image region. For example, the aforementioned morphological function can be the regionprops() function.
[0046] Finally, the morphological feature data with the largest eccentricity among the obtained morphological feature data can be identified as the image micro-feature data.
[0047] The third step involves data parsing and processing the aforementioned preoperative laboratory data to obtain structured laboratory data. This structured data can be derived from the preoperative laboratory data after transformation. The structured data may include various field names, each corresponding to a specific field value.
[0048] For example, when the structured data of the test is "red blood cell count: When using the field name "red blood cell count", the field value can be "". ".
[0049] In practice, the aforementioned implementing entity can use data parsing tools to convert the preoperative laboratory data into structured data for use as structured laboratory data. These data parsing tools can be those capable of converting semi-structured data into structured data. For example, Apache NiFi could be such a tool.
[0050] The fourth step involves performing entity recognition processing on the preprocessed text data, the image microscopic feature data, and the structured test data to obtain the various text entities.
[0051] Each of the aforementioned text entities can be an entity related to thyroid cancer extracted from the aforementioned preprocessed text data, the aforementioned image microscopic feature data, and the aforementioned structured laboratory data. Each of the aforementioned text entities corresponds to an entity category label. The aforementioned entity category label can be a label used to characterize the category of the text entity. The aforementioned entity category label can be, but is not limited to: gender, age, surgical name, laboratory data, and anatomical location.
[0052] In practice, firstly, the aforementioned execution entity can perform entity recognition processing on the preprocessed text data using a pre-trained entity recognition model to obtain various entities to be processed. The entity recognition model can be a neural network model that takes the preprocessed text data as input and outputs the various entities to be processed. For example, the entity recognition model can be a pre-trained BERT (Bidirectional Encoder Representations from Transformers) model. The pre-training process involves fine-tuning the BERT model using an annotated preprocessed text dataset and a cross-entropy loss function. Each entity to be processed can be an entity identified from the preprocessed text data. Each entity to be processed corresponds to an entity category label.
[0053] Secondly, the area, perimeter, major axis length, minor axis length, eccentricity, and convex hull area ratio of the aforementioned image micro-feature data can be identified as individual image entities. Then, for each of these image entities, the category to which the image entity belongs can be determined as the entity category label. For example, when the area is 20, "20" is the image entity, and "area" is the corresponding entity category label.
[0054] Then, the field values of the structured data from the above tests can be identified as individual test entities. Then, for each of these test entities, the corresponding field name can be identified as the entity category label.
[0055] Finally, the above-mentioned entities to be processed, the above-mentioned image entities, and the above-mentioned test entities can be combined into text entities.
[0056] The fifth step involves performing a standardization mapping process on the aforementioned text entities to obtain standardized text entities. These standardized text entities can be entities obtained by mapping the original text entities to a standard data table. Each standardized text entity corresponds to node feature information.
[0057] The aforementioned standard data table can be a data table used to represent the standard names of the aforementioned text entities. The aforementioned standard data table can include various standard names. Each of the aforementioned standard names can be a name of the standard corresponding to the aforementioned text entity, pre-defined by a technician. Each of the aforementioned standard names corresponds to a standard feature vector. The aforementioned standard feature vector can be the feature vector corresponding to the standard name.
[0058] In practice, for each of the aforementioned text entities, in response to the determination that the text entity is not numerical, a feature extraction network can be used to extract features from the text entity, obtaining a feature vector of a preset length as node feature information. The feature extraction network can be a neural network model that takes non-numerical text entities as input and outputs the node feature information of the input text entities. For example, the feature extraction network can be a pre-trained convolutional neural network. Pre-training can be a process of fine-tuning the convolutional neural network using a labeled text training dataset and a cross-entropy loss function. Each text training data in the text training dataset can be text data that does not contain numerical values. The preset length can be a value pre-set by an engineer. Here, the specific setting of the preset length is not limited.
[0059] Secondly, the cosine similarity between the aforementioned node feature information and the corresponding standard feature vectors in the aforementioned standard data table can be determined as the vector similarity. Then, the vector with the highest similarity value among these vectors can be determined as the target similarity. Next, the standard name corresponding to the target similarity can be determined as the standardized text entity corresponding to the aforementioned text entity. Finally, the node feature information corresponding to the aforementioned text entity can be determined as the node feature information corresponding to the standardized text entity.
[0060] Then, in response to determining that the text entity is a numerical value, the text entity can be identified as a standardized text entity. Then, the standardized text entity and its corresponding entity category label can be input to the target technical terminal. The target technical terminal can be a terminal used by a technician. Then, the feature vector of the preset length sent by the target technical terminal can be received as the node feature information of the text entity. For example, when the text entity is 20, the corresponding entity category label is "area", and the preset length is 3, the node feature information can be [0, 0, 20].
[0061] Step 6: Based on the above standardized text entities, perform entity alignment processing on the thyroid cancer knowledge graph to obtain the alignment results of each entity.
[0062] Each of the above entity alignment results can be a label used to characterize whether the thyroid cancer knowledge graph contains standardized text entities. For example, the entity alignment result can be "contains" or "does not contain".
[0063] In practice, for each of the aforementioned standardized text entities, firstly, a lookup function can be used to search for the standardized text entity in the thyroid cancer knowledge graph, and the return value of the lookup function can be used as the search result. The lookup function can be the MATCH function. Then, in response to determining that the search result is not a preset error character, "contains" can be determined as the entity alignment result. In response to determining that the search result is the preset error character, "does not contain" can be determined as the entity alignment result. The preset error character can be a character used to characterize a search failure. For example, the preset error character can be "#N / A".
[0064] Step 7: Based on the alignment results of the above entities, the standardized text entities and the thyroid cancer knowledge graph are fused to update the thyroid cancer knowledge graph.
[0065] In practice, for each of the above entity alignment results, the first step is to identify the standardized text entity corresponding to the above entity alignment result as the standardized entity to be processed.
[0066] Secondly, in response to the determination that the above entity alignment result is "not included", a new node can be created in the thyroid cancer knowledge graph using a creation statement to update the thyroid cancer knowledge graph. The label of the new node can be the entity category label corresponding to the above-mentioned standardized entity to be processed. The new node can be the above-mentioned standardized entity to be processed. Simultaneously, the node feature information corresponding to the above-mentioned standardized entity to be processed can be determined as the node feature information corresponding to the new node. The above creation statement can be a statement capable of creating a new node in the knowledge graph. For example, the above creation statement can be a CREATE statement.
[0067] Then, in response to determining that the above entity alignment result is "contains", the above lookup function can be used to search for nodes in the thyroid cancer knowledge graph whose attributes are the same as or correspond to the above-mentioned standardized entity to be processed, and these nodes are designated as nodes to be updated. Then, an update statement can be used to merge the above-mentioned standardized entity to be processed into the attributes of the nodes to be updated, thereby updating the thyroid cancer knowledge graph. The above update statement can be any statement capable of updating nodes in the knowledge graph. For example, the above update statement can be a SET statement.
[0068] Step 102: Based on the updated thyroid cancer knowledge graph, perform feature node mapping processing on the user's electronic medical record to obtain the user knowledge subgraph.
[0069] In some embodiments, the aforementioned execution entity may perform feature node mapping processing on the aforementioned user electronic medical records based on the updated thyroid cancer knowledge graph to obtain a user knowledge subgraph. The aforementioned user knowledge subgraph may be a subgraph obtained from the updated thyroid cancer knowledge graph. The aforementioned user knowledge subgraph may include various nodes and various edges. Each of the aforementioned nodes corresponds to node feature information, labels, and attributes.
[0070] In some optional implementations of certain embodiments, the aforementioned execution entity may perform feature node mapping processing on the aforementioned user electronic medical records based on the updated thyroid cancer knowledge graph to obtain a user knowledge subgraph: The first step is to extract key entities from the aforementioned user's electronic medical record to obtain the user's identity entity. This user identity entity can be the ID card number in the user's electronic medical record. In practice, the executing entity can use a body entity recognition model to extract key entities from the user's electronic medical record to obtain the user's identity entity. This body entity recognition model can be a neural network model that takes the user's electronic medical record as input and the user's identity entity as output. For example, this body entity recognition model can be a pre-trained BERT model. Pre-training can be the process of fine-tuning the BERT model using an annotated training text set and a cross-entropy loss function. Each training text in the training text set can be text data containing an ID card number.
[0071] The second step is to standardize the aforementioned user identity entities to obtain standardized identity entities. In practice, the aforementioned user identity entities can be identified as standardized identity entities.
[0072] The third step involves performing a breadth-first traversal of the updated thyroid cancer knowledge graph based on the standardized identity entities described above, to obtain various seed nodes. Each of these seed nodes can be a node directly connected to the target identity node. The target identity node can be a node in the thyroid cancer knowledge graph that is identical to the standardized identity entity described above.
[0073] In practice, the aforementioned executing entity can use the search function to find nodes in the updated thyroid cancer knowledge graph that are identical to the standardized identity entities mentioned above, as target identity nodes. Secondly, a search algorithm can be used to identify each node in the updated thyroid cancer knowledge graph directly connected to the target identity node as a seed node. This search algorithm can be one capable of traversing the graph structure. For example, a breadth-first search algorithm could be used.
[0074] Fourth, for each of the aforementioned seed nodes, based on the updated thyroid cancer knowledge graph, the path of each seed node is expanded to obtain an expanded node set. Each expanded node in the expanded node set can be a node in the updated thyroid cancer knowledge graph whose hop count to the aforementioned seed node is less than or equal to a preset value. The hop count can be the number of edges connecting two nodes. The hop count is greater than or equal to 1. The preset value can be a pre-defined value. For example, the preset value can be 3.
[0075] In practice, for each of the above seed nodes, the executing entity can determine the nodes in the updated thyroid cancer knowledge graph that have a hop count less than or equal to the above preset value as extended nodes, thereby obtaining each extended node as an extended node set.
[0076] The fifth step involves performing subgraph query and relationship reconstruction on the updated thyroid cancer knowledge graph based on the obtained extended node sets, to obtain the user knowledge subgraph.
[0077] In practice, firstly, the various extended node sets mentioned above can be combined into a new set using a set creation function, which then serves as the updated extended node set. This removes duplicate extended nodes from the various extended node sets. The set creation function can be any function capable of creating sets. For example, the set() function could be used.
[0078] Then, a database query can be used to retrieve a subgraph from the updated thyroid cancer knowledge graph that contains the aforementioned updated extended node set and the edges between each updated extended node, serving as the user knowledge subgraph. This database query can be a Cypher statement.
[0079] Step 103: Perform node attention aggregation processing on the user knowledge subgraph to obtain each meta-path subgraph and the node feature information of each subgraph.
[0080] In some embodiments, the aforementioned execution entity may perform node attention aggregation processing on the aforementioned user knowledge subgraph to obtain each meta-path subgraph and each subgraph node feature information.
[0081] In some optional implementations of certain embodiments, the aforementioned execution entity can perform node attention aggregation processing on the aforementioned user knowledge subgraph through the following steps to obtain each meta-path subgraph and each subgraph node feature information: For each path template in the pre-acquired path templates, perform the following steps: The first step involves performing pattern matching on the path template based on the aforementioned user knowledge subgraph to obtain a meta-path subgraph. The path template can be a sequence composed of the labels corresponding to each node. For example, the path template could be "{ID number, surgery name, lab data}". The meta-path subgraph can be a graph structure where the labels corresponding to each node found in the user knowledge subgraph are the same as those included in the path template. The meta-path subgraph can include each node and each edge. Each node corresponds to node feature information, attributes, and labels.
[0082] In practice, firstly, the first label included in the aforementioned path template can be designated as the first label. Then, any node in the aforementioned user knowledge subgraph whose label is the first label can be designated as the first node. Secondly, the second label included in the aforementioned path template can be designated as the second label. Then, a node connected to the first node whose corresponding label is the second label can be designated as the second node. Then, the third label included in the aforementioned path template can be designated as the third label. Then, a node connected to the second node whose corresponding label is the third label can be designated as the third node. This process continues until the node corresponding to the last label is obtained. Then, the obtained first, second, third, and even last nodes can be combined into a set as the label node set. Then, the aforementioned database query statement can be used to query the aforementioned user knowledge subgraph for a subgraph containing the aforementioned label node set and the edges between each label node, which can then be used as the meta-path subgraph.
[0083] The second step is to identify the nodes in the aforementioned meta-path subgraph that meet the preset instance conditions as instance nodes. The instance condition can be that the corresponding label is the first label in the aforementioned path template.
[0084] The third step is to determine the neighboring nodes corresponding to the aforementioned instance nodes as the instance's neighboring nodes. In practice, the executing entity can determine the neighboring nodes corresponding to the aforementioned instance nodes in the aforementioned meta-path subgraph as the instance's neighboring nodes.
[0085] Fourth, for each instance neighbor node among the aforementioned instance neighbor nodes, the semantic features of the aforementioned instance neighbor node and the aforementioned instance node are concatenated to obtain concatenated node feature information. This concatenated node feature information can be a feature vector obtained by concatenating the node feature information of the instance neighbor node and the instance node.
[0086] In practice, for each of the aforementioned instance neighbor nodes, the executing entity can input the node feature information corresponding to the instance neighbor node and the node feature information corresponding to the instance node into the concatenation function to obtain concatenated node feature information. The concatenation function can be any function capable of concatenating features. For example, the concatenation function could be the `concat()` function.
[0087] The fifth step involves performing feature dimension mapping processing on the obtained feature information of each splicing node to obtain various mapped feature information. Each mapped feature information can be a splicing node feature information mapped to a preset feature length. The preset feature length can be a pre-defined length. For example, the preset feature length can be 3.
[0088] In practice, for each of the splicing node feature information mentioned above, the length of the splicing node feature information can be determined as the vector length. Then, a matrix with the number of rows equal to the vector length, the number of columns equal to the preset feature length, and all elements being normally random numbers can be created using a matrix creation function as a mapping matrix. The product of the splicing node feature information and the mapping matrix can then be used to determine the mapped feature information. Here, the normally random numbers can be random numbers following a standard normal distribution. The matrix creation function can be any function capable of creating matrices. For example, the matrix creation function could be the `torch.randn()` function.
[0089] Step 6: Based on the preset attention vector and the aforementioned mapping feature information, generate various attention weights. The attention vector can be a preset vector of the same size as the mapping feature information. Each attention weight can be a weight obtained from weighted data. Each attention weight corresponds to an instance neighbor node. The weighted data can be the dot product of the mapping feature information and the attention vector.
[0090] In practice, firstly, for each of the aforementioned mapping features, the dot product of the mapping feature and the attention vector can be used to determine the weighted data. Secondly, for each of the determined weighted data, the weighted data can be input into a linear rectified function, and the return value of the linear rectified function can be used as the positive weighted data. Then, the attention weight can be determined as the power of the positive weighted data with a preset base. The preset base can be the natural constant e. For example, when the positive weighted data is 1.1, the attention weight can be e raised to the power of 1.1.
[0091] Step 7: Based on the attention weights and neighbor nodes of each instance mentioned above, generate subgraph node feature information. This subgraph node feature information can be the feature vector corresponding to the user knowledge subgraph.
[0092] In practice, for each of the attention weights mentioned above, firstly, the instance neighbor node corresponding to the attention weight can be determined as the target instance neighbor node. Then, the node feature information corresponding to the target instance neighbor node can be determined as the target node feature information. Finally, the product of the attention weight and the target node feature information can be determined as the weighted node feature information.
[0093] Next, the sum of the determined weighted node feature information can be used to determine the subgraph feature information. Finally, the subgraph feature information can be input into the activation function, and the return value of the activation function can be used as the subgraph node feature information. The activation function can be a linear rectified function.
[0094] Step 104: Perform gating enhancement processing on the user's electronic medical record, each metapath subgraph, and the feature information of each subgraph node to obtain the user feature information.
[0095] In some embodiments, the execution entity may perform gating enhancement processing on the user's electronic medical record, the various meta-path subgraphs, and the feature information of the nodes of the various subgraphs to obtain user feature information. The user feature information may be a feature vector obtained by fusing the features of the user's electronic medical record and the feature information of the nodes of the various subgraphs.
[0096] In some optional implementations of certain embodiments, the aforementioned execution entity may perform gating enhancement processing on the aforementioned user electronic medical record, the aforementioned meta-path subgraphs, and the aforementioned subgraph node feature information through the following steps to obtain user feature information: The first step is to determine the path weight data for each of the aforementioned meta-path subgraphs. This path weight data can be a decimal used to represent the weight of the meta-path subgraph.
[0097] In practice, for each of the aforementioned meta-path subgraphs, firstly, the number of nodes included in the meta-path subgraph can be determined as the number of graph nodes. Secondly, the sum of the determined number of graph nodes can be determined as the total number of nodes. Then, for each of the aforementioned meta-path subgraphs, the ratio of the number of graph nodes corresponding to the aforementioned meta-path subgraph to the total number of nodes can be determined as the path weight data.
[0098] The second step is to perform weighted fusion processing on the feature information of each subgraph node based on the determined path weight data to obtain fused feature information.
[0099] The aforementioned fused feature information can be a feature vector obtained by weighted fusion of the feature information of each subgraph node.
[0100] In practice, for each path weight data point in the aforementioned path weight data, the product of the path weight data and the feature information of the corresponding subgraph node can be determined as the weighted feature information. Then, the sum of the determined weighted feature information can be determined as the fusion feature information.
[0101] The third step involves extracting and fusing features from the user's electronic medical record and the fused feature information to obtain clinical splicing feature information. This clinical splicing feature information can be a feature vector obtained by splicing the feature vector extracted from the user's electronic medical record with the fused feature information.
[0102] In practice, firstly, for the preoperative laboratory data included in the aforementioned user electronic medical records, the aforementioned data parsing tool can be used to parse the preoperative laboratory data into structured data, which is then used as structured laboratory data. This structured laboratory data can include various field names. Each field name corresponds to a specific field value.
[0103] Secondly, the field values of the structured test data can be arranged into a row vector according to a preset order to serve as the test vector. This preset order can be a pre-defined order by the technicians for arranging the field values in the structured test data. Then, the test vector and the fusion feature information can be input into the splicing function to obtain the spliced feature vector, which serves as the clinical splicing feature information.
[0104] The fourth step involves performing parameter weighting on the aforementioned clinical splicing feature information to obtain a parameter-weighted vector. This parameter-weighted vector can be clinical splicing feature information mapped to a preset vector length. The preset vector length can be a pre-defined value. For example, the preset vector length can be 3.
[0105] In practice, the aforementioned executing entity can determine the parameter-weighted vector by multiplying the aforementioned clinical splicing feature information with a preset parameter-weighted matrix. The parameter-weighted matrix can be a matrix with the same number of rows as the aforementioned clinical splicing feature information, the same number of columns as the preset vector length, and each element containing a random number following a standard normal distribution.
[0106] Fifth, the sum of the above-mentioned parameter weighted vector and the preset bias vector is determined as the bias weighted vector. The bias vector can be a preset vector with the same size as the above-mentioned parameter weighted vector.
[0107] The sixth step involves performing variable mapping on the aforementioned bias weighted vector to obtain a gated value vector. This gated value vector can be a vector obtained by mapping the element values of the aforementioned bias weighted vector to a range between 0 and 1.
[0108] In practice, the aforementioned execution entity can input the bias-weighted vector into a mapping function, obtaining the output vector of the mapping function as the gate value vector. The mapping function can be any function capable of mapping values in the vector to the range 0-1. For example, the mapping function could be the sigmoid function.
[0109] Step 7: Based on the above gate value vector, generate a gate value negative vector. The gate value negative vector can be the opposite of the above gate value vector.
[0110] Step 8: Determine the first linear vector by combining the above-mentioned gated value vector with the above-mentioned fused feature information.
[0111] The ninth step is to determine the Hadamard product of the above-mentioned gated negative vector and the above-mentioned parameter weighted vector as the second linear vector.
[0112] Step 10: The sum of the first linear vector and the second linear vector is determined as the user feature information.
[0113] Step 105: Based on the multimodal large model used to process thyroid cancer data, the user feature information is matched and adaptively optimized to obtain the target follow-up plan.
[0114] In some embodiments, the aforementioned execution entity may perform scheme matching and adaptive optimization processing on the aforementioned user feature information based on a multimodal large model used for processing thyroid cancer data, to obtain a target follow-up scheme.
[0115] In some optional implementations of certain embodiments, the aforementioned execution entity may perform scheme matching and adaptive optimization processing on the aforementioned user feature information based on a multimodal large model used for processing thyroid cancer data through the following steps to obtain a target follow-up scheme: The first step involves generating a follow-up template plan based on the aforementioned user characteristic information and a pre-defined standard protocol knowledge base. This standard protocol knowledge base can be a database storing various rehabilitation template plans. Each of these rehabilitation template plans can be a pre-designed rehabilitation plan for users in the thyroid cancer recovery period. For example, a rehabilitation template plan could be "First week post-surgery: Daily monitoring of blood calcium, daily uploading of wound photos." Each of these rehabilitation template plans corresponds to template characteristic information. This template characteristic information can be the user characteristic information applicable to the user for whom the rehabilitation template plan is applied.
[0116] In practice, for each rehabilitation template scheme included in the aforementioned standard scheme knowledge base, firstly, the cosine similarity between the template feature information corresponding to the rehabilitation template scheme and the user feature information can be determined as the user feature similarity. Then, the user feature similarity with the largest value among the determined user feature similarities can be determined as the target user feature similarity. Finally, the rehabilitation template scheme corresponding to the target user feature similarity can be determined as the follow-up template scheme.
[0117] The second step involves determining standard health data based on the aforementioned follow-up template and the pre-defined standard data table. The standard data table can be a data table used to characterize the correspondence between standard data and the follow-up template. The standard data table can include various standard data points. Each standard data point can be preoperative laboratory data corresponding to a healthy individual. Each standard data point corresponds to a standard template. The standard template can be the follow-up template corresponding to the standard data.
[0118] In practice, firstly, the aforementioned executing entity can encode the aforementioned follow-up template scheme using text encoding technology to obtain the feature vector corresponding to the follow-up template scheme as the follow-up feature vector. The aforementioned text encoding technology can be any technology capable of encoding text. For example, the aforementioned text encoding technology can be TF-IDF (termfrequency–inverse document frequency).
[0119] Secondly, for each standard template scheme in the standard template scheme corresponding to the standard data table of the above scheme, the above-mentioned text encoding technology can be used to encode the above-mentioned standard template scheme to obtain the feature vector corresponding to the above-mentioned standard template scheme as the template feature vector.
[0120] For each template feature vector obtained, the cosine similarity between the template feature vector and the follow-up feature vector can be determined as the template similarity. Then, the template similarity with the largest value among the determined template similarities can be determined as the target template similarity. Finally, the standard data corresponding to the target template similarity can be determined as standard health data.
[0121] The third step involves generating a long-term follow-up plan based on preset adjustment prompts, the aforementioned follow-up template, the aforementioned standard health data, and the aforementioned user's electronic medical record. This long-term follow-up plan can be an adjusted version of the aforementioned follow-up template, representing a rehabilitation plan that a user in the thyroid cancer recovery phase needs to follow.
[0122] The aforementioned adjustment prompts can be used to guide technical personnel to adjust the follow-up template plan. For example, the adjustment prompts could be: "Based on the aforementioned standard health data and the aforementioned user electronic medical records, a plan for adjusting the follow-up template plan is provided."
[0123] In practice, the aforementioned implementing entity can send the adjustment prompts, follow-up template, standard health data, and user electronic medical records to the target terminal. Then, it can receive the data sent by the target terminal as a long-term follow-up plan. The target terminal can be the terminal corresponding to the technical personnel.
[0124] The fourth step is to standardize the aforementioned user feature information to obtain standard feature information. This standard feature information can be the standardized user feature information itself.
[0125] In practice, the aforementioned implementing entity can use a standardization tool to convert the user feature information into a standard normal distribution with a mean of 0 and a standard deviation of 1, thereby standardizing the user feature information and obtaining the processed user feature information as standard feature information. The standardization tool can be any tool capable of converting features into a standard normal distribution with a mean of 0 and a standard deviation of 1. For example, the standardization tool could be sklearn.preprocessing.StandardScaler.
[0126] The fifth step involves performing word segmentation and mapping on the aforementioned long-term follow-up scheme to obtain a follow-up word segmentation data sequence. Each follow-up word segmentation data in the sequence can be a numerical value corresponding to a word in the aforementioned long-term follow-up scheme.
[0127] In practice, firstly, the aforementioned execution entity can use a tokenizer to perform word segmentation on the long-term follow-up scheme and map the segmented long-term follow-up scheme to a lexicon, obtaining the segmented and mapped long-term follow-up scheme as a long-term word segmentation data sequence. The tokenizer can be a tool capable of segmenting text and mapping the segmented data to numerical values. For example, the tokenizer could be a Tokenizer. The lexicon can be a lexicon included with the tokenizer tool. For example, the lexicon could be a lexicon included with the Tokenizer. The lexicon can include individual words and numbers. Each word in the lexicon has a unique number. Each long-term segmented data in the long-term word segmentation data sequence can be a number from the lexicon. Then, a preset start string can be added to the beginning position of the long-term word segmentation data sequence using an add function, obtaining the added long-term word segmentation data sequence as a follow-up word segmentation data sequence. The add function can be a function capable of adding a specified string to the sequence. For example, the add function could be the insert() function. The preset start string can be a pre-defined string. For example, the preset start string mentioned above could be "[BOS]".
[0128] The sixth step is to input the above standard feature information and the above follow-up word segmentation data sequence into the above multimodal large model to obtain the feature score matrix.
[0129] Step 7: Based on the aforementioned feature score matrix, generate the target follow-up plan. This target follow-up plan can be the most suitable follow-up plan for the target users.
[0130] In practice, for each row in the aforementioned feature score matrix, the element with the largest value among all element values in that row can be determined as the target row element value. Then, the corresponding index of the target row element value can be determined as the target index. Next, the determined target indices can be arranged into a sequence according to the order from the first row to the last row of the feature score matrix, forming the target index sequence.
[0131] Then, for each target number in the above target number sequence, the word corresponding to that target number in the above vocabulary can be identified as the target word. Finally, the target words corresponding to the above target number sequence can be arranged into a sequence as a target follow-up plan according to the order in the above target number sequence.
[0132] In some alternative implementations of certain embodiments, the aforementioned multimodal large model is trained through the following steps: The first step is to obtain the sample set. Each sample in the sample set includes standard feature information, a follow-up word segmentation data sequence, and a sample label sequence. The standard feature information can be the standard feature information used when training the multimodal large model. The follow-up word segmentation data sequence can be the follow-up word segmentation data sequence used when training the multimodal large model. Each sample number in the sample number sequence can be a number from the vocabulary.
[0133] The second step, based on the sample set, is to perform the following training steps: The first sub-step involves inputting the standard feature information and follow-up word segmentation data sequence of each sample in the sample set into the input layer of the initial network model, thereby obtaining the follow-up word segmentation matrix, feature embedding information, and text embedding matrix corresponding to each sample. The input layer can be a neural network layer that takes the standard feature information and follow-up word segmentation data sequence as input and outputs the feature embedding information and text embedding matrix. This input layer may include a vector processing layer and a text processing layer.
[0134] The aforementioned vector processing layer may include a fully connected layer and a feature position encoding module. The fully connected layer may be a neural network layer that takes sample standard feature information as input and transform feature information as output. The transform feature information may be sample standard feature information mapped to a preset feature vector length by the fully connected layer. The preset feature vector length may be a pre-defined length of the feature vector. For example, the preset feature vector length may be 128. The feature position encoding module may be a neural network module that takes transform feature information as input and outputs feature embedding information. For example, the feature position encoding module may be Positional Encoding. The feature embedding information may be the transformed feature information after position encoding.
[0135] The text processing layer described above can include a feature embedding module and a text position encoding module.
[0136] The aforementioned feature embedding module can be a neural network module that takes the sample follow-up word segmentation data sequence as input and the follow-up word segmentation matrix as output. For example, the aforementioned feature embedding module can be Token Embeddings. The aforementioned follow-up word segmentation matrix can be a matrix in which each row corresponds to a sample follow-up word segmentation data in the aforementioned sample follow-up word segmentation data sequence, and each row of data is a feature vector of the corresponding sample follow-up word segmentation data.
[0137] The aforementioned text position encoding module can be a neural network module that takes a follow-up word segmentation matrix as input and a text embedding matrix as output. For example, the aforementioned text position encoding module can be Positional Encoding. The aforementioned text embedding matrix can be a follow-up word segmentation matrix after position encoding processing. The length of the aforementioned feature embedding information is the same as the number of columns in the aforementioned text embedding matrix; that is, the number of columns in the aforementioned text embedding matrix is equal to the aforementioned preset feature vector length.
[0138] The second sub-step involves generating at least one decoding feature matrix based on the multimodal fusion coding layer, feature-aware encoder, conditional decoder, follow-up word segmentation matrix corresponding to each of the at least one sample in the initial network model, feature embedding information, and text embedding matrix.
[0139] In practice, firstly, for each of the at least one sample mentioned above, the feature embedding information and text embedding matrix corresponding to the sample can be input into the multimodal fusion coding layer in the initial network model to obtain the multimodal fusion matrix corresponding to the sample. The multimodal fusion coding layer can be a neural network layer that takes the feature embedding information and text embedding matrix as input and the multimodal fusion matrix as output.
[0140] The aforementioned multimodal fusion coding layer may include a shape expansion function, a cross-modal attention mechanism, a tensor broadcasting mechanism, and a feature matrix concatenation function. The processing procedure of the multimodal fusion coding layer is as follows: First, the shape of the text embedding matrix can be determined as the target extended shape. Second, the multimodal fusion coding layer can input the feature embedding information into the shape extension function to extend the shape of the feature embedding information to the target extended shape, obtaining the extended feature embedding information as the feature embedding matrix. The shape extension function can be the `unsqueeze()` function. For example, when the shape of the text embedding matrix is (1, 25, 128) and the shape of the feature embedding information is 128, the resulting feature embedding matrix has the shape (1, 1, 128).
[0141] Then, the feature embedding matrix can be used as the query, and the text embedding matrix as the key and value. A cross-modal attention mechanism can be used to perform attention-weighted aggregation on the text embedding matrix to obtain the cross-modal feature matrix. This cross-modal attention mechanism can be Cross-Attention. The cross-modal feature matrix can be the matrix obtained after performing attention-weighted aggregation on the text embedding matrix. For example, when the shape of the feature embedding matrix is (1,1,128) and the shape of the text embedding matrix is (1,25,128), the shape of the resulting cross-modal feature matrix is (1,1,128).
[0142] Then, the cross-modal feature matrix and the text embedding matrix can be added together using the tensor broadcasting mechanism described above, resulting in a cross-modal broadcasting matrix. The tensor broadcasting mechanism can be a tensor broadcasting mechanism. For example, when the cross-modal feature matrix has a shape of (1,1,128) and the text embedding matrix has a shape of (1,25,128), the cross-modal broadcasting matrix will have a shape of (1,25,128).
[0143] Finally, the aforementioned feature embedding matrix and cross-modal broadcast matrix can be input into the aforementioned feature matrix concatenation function to obtain the multimodal fusion matrix as the output matrix of the feature matrix concatenation function. The aforementioned feature matrix concatenation function can be the torch.cat() function. For example, when the shape of the feature embedding matrix is (1,1,128) and the shape of the cross-modal broadcast matrix is (1,25,128), the shape of the multimodal fusion matrix can be (1,26,128).
[0144] Second, for each of the at least one sample mentioned above, the multimodal fusion matrix corresponding to the sample can be input into the feature-aware encoder in the initial network model to obtain the encoded feature matrix corresponding to the sample. The feature-aware encoder can be an encoder that takes the multimodal fusion matrix as input and the encoded feature matrix as output. For example, the feature-aware encoder can be a 6-layer Transformer. The encoded feature matrix can be the multimodal fusion matrix processed by the 6-layer Transformer.
[0145] Third, the follow-up segmentation matrix and encoding feature matrix corresponding to each sample in at least one of the above samples can be input into the conditional decoder in the initial network model to obtain at least one decoded feature matrix. The conditional decoder can be a decoder that takes the follow-up segmentation matrix and encoding feature matrix as input and the decoded feature matrix as output.
[0146] The conditional decoder described above can include a masked self-attention mechanism, a cross-attention mechanism, and a feature processing network. The processing procedure of the conditional decoder is as follows: The aforementioned masked self-attention mechanism uses the follow-up word segmentation matrix as the Query, Key, and Value, and performs attention-weighted aggregation processing on the follow-up word segmentation matrix to obtain a weighted follow-up word segmentation matrix. This masked self-attention mechanism can be termed Masked Self-Attention.
[0147] The aforementioned cross-attention mechanism uses the weighted follow-up word segmentation matrix as the query and the encoded feature matrix as the key and value. It queries the encoded feature matrix using the weighted follow-up word segmentation matrix to obtain the cross-attention matrix. Specifically, the cross-attention matrix can be the matrix obtained by selectively extracting the encoded feature matrix from the weighted follow-up word segmentation matrix using the cross-attention mechanism.
[0148] The feature processing network described above can be a neural network that takes a cross-attention matrix as input and a decoded feature matrix as output. For example, the feature processing network described above can be a feed-forward network (FFN). The decoded feature matrix described above can be the cross-attention matrix after feature transformation and nonlinear mapping processing by the feed-forward neural network.
[0149] The third sub-step involves inputting at least one decoded feature matrix corresponding to at least one sample into the output layer of the initial network model to obtain at least one feature score matrix. The output layer can be a linear projection layer that takes the decoded feature matrix as input and the feature score matrix as output. The feature score matrix can be a matrix where each row corresponds to a decoding position, each column corresponds to a number in the vocabulary, and each element represents the score of the data at the decoding position corresponding to the row as the score of the number corresponding to the column. The decoding position can be the position of the currently generated data in the sequence during sequence generation. The number of rows in the feature score matrix is the same as the number of sample numbers included in the sample number sequence.
[0150] The fourth sub-step involves performing a difference comparison between the feature score matrix corresponding to each of the at least one sample and the corresponding sample number sequence to obtain at least one comparison result. Each of these at least one comparison result can be the cross-entropy loss value between the feature score matrix and the corresponding sample number sequence.
[0151] In practice, for the feature score matrix corresponding to each sample in at least one of the above samples, the cross-entropy loss value of the feature score matrix and the corresponding sample number sequence can be used as the alignment result to obtain at least one alignment result.
[0152] The fifth sub-step involves determining that at least one of the above alignment results meets the optimization objective, and then defining the initial network model as a multimodal large model. The optimization objective can be that the average value of at least one of the above alignment results is less than a preset loss threshold. This preset loss threshold can be a pre-defined value. The specific setting of this preset loss threshold is not limited here.
[0153] Optionally, in response to determining that at least one of the above alignment results has not achieved the above optimization objective, the network parameters of the initial network model are adjusted, and a sample set is formed using unused samples. The adjusted initial network model is then used as the initial network model, and the above training steps are executed again. In practice, the above execution entity can use the backpropagation algorithm (BP algorithm) and gradient descent methods (such as mini-batch gradient descent algorithm) to adjust the network parameters of the initial network model.
[0154] Figure 4 This can be a flowchart for training a large multimodal model.
[0155] Step 106: Based on the target follow-up plan, control the intelligent central control terminal to perform task assistance operations.
[0156] In some embodiments, the aforementioned execution entity may control the aforementioned intelligent central control terminal to perform task assistance operations based on the aforementioned target follow-up scheme.
[0157] In addressing the aforementioned technical problems in the process of adopting technical solutions, and considering the application scenario of assisting elderly people living alone with rehabilitation, the following technical issues often arise: the operation process of the smart control terminal is relatively cumbersome, and users are prone to making incorrect interactions due to unfamiliarity with the operation, resulting in poor convenience for users to control the smart control terminal for rehabilitation training. Considering the following requirements for this application scenario: elderly people living alone are often unfamiliar with the operation of smart devices and have high requirements for the ease of operation of controlling the smart control terminal, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned execution entity can control the aforementioned intelligent central control terminal to perform task assistance operations based on the aforementioned target follow-up scheme through the following steps: The first step, in response to determining that the triggering conditions for a subjective feeling inquiry operation are met, is to play the target audio segment to the user. This subjective feeling inquiry operation can be an action that proactively asks the user about their current discomfort symptoms. The target audio segment can be audio data used to ask the user what discomfort symptoms they currently experience. For example, the target audio segment could be the audio message, "What physical discomfort symptoms do you currently experience?"
[0158] In practice, the triggering condition for a subjective feeling inquiry operation can be determined by receiving an active inquiry message from the target technology terminal. The target technology terminal can be a terminal used by a technician. The active inquiry message can be information used to prompt the executing entity to perform the subjective feeling inquiry operation. For example, the active inquiry message could be "Start performing the subjective feeling inquiry operation."
[0159] The second step is to respond to the user's voice response to the target voice segment and determine the user's status information based on the voice response.
[0160] The aforementioned response voice can be a voice message issued by the user. The aforementioned user status information can be the name representing the disease the user has. For example, the aforementioned user status information can be, but is not limited to, "no disease" or "thyroid cancer".
[0161] In practice, firstly, the aforementioned executing entity can input preset output prompts and the surgical records included in the user's electronic medical record into the large language model, obtaining the text output by the large language model as the user's status information. The aforementioned output prompts can be information used to instruct the large language model to output user status information based on the surgical records included in the user's electronic medical record. For example, the output prompts could be, "Please directly output the name of the user's disease based on the surgical record; if not, please reply 'No disease'." The aforementioned large language model can be DeepSeek.
[0162] The third step is to perform text conversion processing on the above response speech to obtain speech-text data. This speech-text data can be the text data corresponding to the above response speech.
[0163] In practice, the aforementioned implementing entity can use speech-to-text technology to convert the aforementioned speech information into text data as speech-text data. Specifically, the speech-to-text technology can be any technology capable of converting speech into text. For example, the speech-to-text technology can be speech-to-text conversion.
[0164] The fourth step is to perform text preprocessing on the above-mentioned speech-text data to obtain preprocessed speech-text data. This preprocessed speech-text data can be the speech-text data that has already undergone preprocessing. In practice, firstly, consecutive punctuation marks can be removed from the speech-text data using a data cleaning function to perform text preprocessing and obtain preprocessed speech-text data. This data cleaning function can be any function capable of removing consecutive punctuation marks from the text. For example, the data cleaning function could be the `re.sub()` function.
[0165] The fifth step involves extracting keywords from the preprocessed speech-text data and the user status information to obtain a text keyword sequence. Each keyword in this sequence can be a keyword extracted from the preprocessed speech-text data that characterizes the user's current symptoms. For example, keywords could be, but are not limited to, "wound suppuration," "cough," or "fever."
[0166] In practice, the aforementioned execution entity can use a pre-trained named entity recognition model to extract keywords from the preprocessed speech-text data, obtaining a text keyword sequence. This named entity recognition model can be a neural network model that takes the preprocessed speech-text data as input and the text keyword sequence as output. For example, the named entity recognition model can be a pre-trained BERT (Bidirectional Encoder Representations from Transformers). The pre-training process involves fine-tuning BERT using an annotated preprocessed speech-text dataset and a cross-entropy loss function.
[0167] Step 6: Based on the aforementioned user status information and the aforementioned text keyword sequence, generate status risk data. This status risk data can be a label representing whether the user's current physical condition requires close monitoring. For example, the status risk data could be "requires close monitoring" or "does not require close monitoring."
[0168] In practice, firstly, in response to determining that the user's status information is "no disease," "no need for high-level monitoring" can be identified as status risk data. Then, in response to determining that the user's status information is not "no disease," the aforementioned lookup function can be used to search for a disease name identical to the user's status information in a pre-defined disease symptom data table as the target disease name. This disease symptom data table can be a table representing the relationship between disease names and various symptom keywords. The disease symptom data table can include various disease names. Each disease name corresponds to various symptom keywords. Each symptom keyword can be a keyword representing abnormal symptoms exhibited by the user after disease treatment. For example, a symptom keyword could be "cough." Then, the symptom keywords corresponding to the target disease name can be identified as the target symptom keyword set. In response to determining that any text keyword in the aforementioned text keyword sequence is identical to any target symptom keyword in the target symptom keyword set, "needs high-level monitoring" can be identified as status risk data.
[0169] Step 7: In response to the determination that the above-mentioned status risk data does not meet the preset risk warning conditions, the following steps are executed: The first sub-step involves obtaining relevant text data based on the aforementioned text keyword sequence and a pre-defined professional text knowledge base. The aforementioned risk warning condition can be set to "requires high-level monitoring" for the aforementioned status risk data.
[0170] The aforementioned specialized text knowledge base can be a database storing text data on how to deal with various symptoms. Each piece of text data on how to deal with symptoms can be text data recording symptoms and corresponding coping measures. For example, the text data on how to deal with symptoms could be "When a user has a fever, they should take antipyretic medication promptly." The aforementioned specialized text knowledge base can include text data on how to deal with various symptoms.
[0171] In practice, for each symptom response text data item and each text keyword in the aforementioned text keyword sequence, a string search function can be used to search for the aforementioned text keyword in the symptom response text data. The return value of the string search function is taken as the keyword search result. The string search function can be any function capable of finding a specified string in text data. For example, the string search function can be the `find()` function in Python. The keyword search result can be the index of the first occurrence of the target search string or -1. The target search string can be a string in the symptom response text data that is identical to the text keyword. Then, in response to determining that the obtained keyword search results do not meet the preset keyword search conditions, the aforementioned symptom response text data can be identified as relevant text data. The keyword search condition can be that all obtained keyword search results are -1. Thus, the relevant text data can be obtained.
[0172] The second sub-step involves generating a text similarity score for each of the aforementioned related text data, based on the preprocessed speech-text data and the related text data. This text similarity score can be a numerical value representing the degree of similarity between the preprocessed speech-text data and the related text data.
[0173] In practice, for each of the aforementioned related text data, firstly, the executing entity can use a text segmentation tool to segment the preprocessed speech text data and the related text data to obtain speech text segmentation data and related text segmentation data. The speech text segmentation data can be the data obtained after segmenting the preprocessed speech text data. For example, when the preprocessed speech text data is "I like the sea," the obtained speech text segmentation data could be "["I", "like", "sea"]". The related text segmentation data can be the data obtained after segmenting the related text data. The text segmentation tool can be any tool capable of segmenting text. For example, the text segmentation tool could be jieba. Secondly, feature processing techniques can be used to convert the speech text segmentation data into vectors as speech text vectors. Then, the aforementioned feature processing techniques can be used to convert the related text data into vectors as related text vectors. The feature processing techniques can be any techniques capable of converting text into vectors. For example, the feature processing technique could be TF-IDF technology. Then, the cosine similarity between the above-mentioned speech text vector and the above-mentioned related text vector can be determined as the text similarity.
[0174] The third sub-step involves identifying the text similarities among the generated text similarities that meet preset similarity criteria as the target text similarities. The aforementioned similarity criteria can be maximizing the corresponding text similarity value.
[0175] The fourth sub-step involves identifying the relevant text data corresponding to the target text similarity as the target-related text data.
[0176] The fifth sub-step involves generating a text response based on the aforementioned user status information, the aforementioned voice-text data, and the aforementioned target-related text data. This text response can be information used to reply to the aforementioned voice information. In practice, the executing entity can determine the aforementioned target-related text data as the text response.
[0177] The sixth sub-step involves performing speech conversion processing on the aforementioned text response information to obtain speech response data. This speech response data can be the audio data of the aforementioned text response information. In practice, the executing entity can use text-to-speech technology to convert the text response information into speech data as the speech response data. This text-to-speech technology can be any technology capable of converting text into speech. For example, the text-to-speech technology can be TextToSpeech technology.
[0178] The seventh sub-step involves playing the aforementioned voice response data.
[0179] The above technical solution and its related content, combined with step 106, serve as an inventive point of this disclosure, solving the problem of "poor operational convenience of controlling the intelligent central control terminal for rehabilitation training." Factors leading to wasted power and computing resources often include: the operation process of the intelligent central control terminal is relatively cumbersome, and users are prone to making incorrect interactions due to unfamiliarity with the operation, resulting in poor user convenience in controlling the intelligent central control terminal for rehabilitation training. Solving these factors can improve the operational convenience of controlling the intelligent central control terminal for rehabilitation training. To achieve this effect, this disclosure firstly, in response to determining that the triggering condition for a subjective feeling inquiry operation is met, a target audio segment is played to the user. This allows for proactively initiating an inquiry to the user. Secondly, in response to receiving the user's response audio to the target audio segment, user status information is determined based on the response audio. This allows for obtaining user status information. Then, the response audio is processed into text to obtain audio-text data. This allows for obtaining audio-text data. Then, the audio-text data is preprocessed to obtain preprocessed audio-text data. This allows for obtaining preprocessed audio-text data. Next, keyword extraction is performed on the preprocessed speech-text data and the user status information to obtain a text keyword sequence. Then, based on the user status information and the text keyword sequence, status risk data is generated. Then, in response to the determination that the status risk data does not meet the preset risk warning conditions, the following steps are performed: First, based on the text keyword sequence and a preset professional text knowledge base, various related text data are obtained. Thus, when it is determined that the user's symptoms are normal, various related text data can be obtained. Second, for each related text data, a text similarity is generated based on the preprocessed speech-text data and the related text data. Then, various text similarities are obtained. Then, the text similarities that meet the preset similarity conditions among the generated text similarities are determined as target text similarities. Then, the related text data corresponding to the target text similarities are determined as target related text data. Then, target related text data are obtained. Next, based on the aforementioned user status information, voice-text data, and target-related text data, a text response is generated. This allows for the extraction of information from the target-related text data to reply to the user. Then, the text response is processed through speech conversion to obtain voice response data. Finally, the voice response data is played back.Because the intelligent central control terminal can proactively inquire about the user's current symptoms to obtain the user's physical condition in real time, and then take appropriate action based on the user's voice response, without waiting for the user to perform complex operations, it can reduce the number of times the intelligent central control terminal frequently executes incorrect instructions due to user operation errors, thereby improving the convenience for users to control the intelligent central control terminal for rehabilitation training.
[0180] In addressing the aforementioned technical problems in the application scenario—a rehabilitation center for the disabled—the following technical issues often arise: operating a smart control terminal to perform rehabilitation tasks typically involves cumbersome steps, and users often find it difficult to actively interact with the smart control terminal due to mobility limitations, resulting in poor user convenience when operating the terminal. Considering the following requirements for this application scenario: users in rehabilitation centers for the disabled often have limited mobility and need the smart control terminal to actively interact to assist them in executing rehabilitation plans, thereby improving the convenience of interaction with the smart control terminal, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned execution entity can control the aforementioned intelligent central control terminal to perform task assistance operations based on the aforementioned target follow-up scheme through the following steps: The first step is to generate today's task information based on the real-time acquired time data and the aforementioned target follow-up plan. The time data can be the date representing today. For example, the time data could be "2025-10-16", representing October 16, 2025. The today's task information can be information representing the tasks the user needs to complete today. For example, today's task information could be "Today's Tasks: Complete the postoperative symptom questionnaire, take and upload photos of the wound."
[0181] In practice, firstly, the implementing entity can determine the surgical time in the surgical records included in the user's electronic medical record as the target surgical time. Secondly, the number of days between the aforementioned time data and the target surgical time can be determined as the difference in days data. For example, when the time data is "2025-10-16" and the target surgical time is "2025-10-10", the difference in days is 6 days, and the resulting difference in days data is 6 days.
[0182] Then, the aforementioned difference in days data, the aforementioned target follow-up plan, and the preset today's task generation information can be sent to the target server, and the text data returned by the target server can be used as today's task information. The aforementioned today's task generation information can be used to prompt the target server to generate today's task information based on the aforementioned difference in days data and the aforementioned target follow-up plan. For example, the aforementioned today's task generation information could be "Please generate the plan to be executed today based on the aforementioned difference in days data and the aforementioned target follow-up plan." The aforementioned target server can be a server capable of generating today's task information based on the aforementioned difference in days data and the aforementioned target follow-up plan.
[0183] The second step, in response to determining that the above-mentioned task information for today meets the preset risk detection conditions, is to execute the following steps: The first sub-step involves controlling the aforementioned intelligent central control terminal to perform a body indicator monitoring task and obtain the user's body indicator data. The risk detection condition can be that the aforementioned today's task information simultaneously includes a first preset string and a second preset string. The first preset string can be a pre-defined string. For example, the first preset string could be "photograph wound". The second preset string can be a pre-defined string that is different from the first preset string. For example, the second preset string could be "monitoring indicators". The user's body indicator data can be data used to characterize the user's body indicators. The user's body indicator data can include: blood pressure, body temperature, heart rate, blood glucose, and blood oxygen saturation.
[0184] In practice, firstly, the aforementioned executing entity can send a preset indicator monitoring command to the aforementioned smart central control terminal. This indicator monitoring command can be a prompt to the smart central control terminal to begin monitoring the user's health indicator data. For example, the indicator monitoring command could be "text:on," indicating that monitoring the user's health indicator data needs to begin. In response to receiving the indicator monitoring command, the aforementioned smart central control terminal can send a preset start monitoring command to the smart monitoring device worn by the user to monitor the user's health indicator data and obtain the user's health indicator data. This smart monitoring device can be a device capable of monitoring the wearer's blood pressure, body temperature, heart rate, blood glucose, and blood oxygen saturation. For example, the smart monitoring device could be a smart bracelet. Then, the aforementioned smart central control terminal can send the user's health indicator data to the aforementioned executing entity. The aforementioned executing entity can then receive the user's health indicator data sent by the aforementioned smart central control terminal.
[0185] The second sub-step involves controlling the aforementioned intelligent central control terminal to perform a wound imaging task and obtain a wound image. This wound image can be an image corresponding to the surgically sutured wound. In practice, the executing entity can send a preset wound imaging command to the intelligent central control terminal to control it to perform the wound imaging task and obtain a wound image. This wound imaging command can be a command used to prompt the intelligent central control terminal to perform the wound imaging task. For example, the wound imaging command could be "takephoto:on".
[0186] The third sub-step involves performing image detection processing on the aforementioned wound images to obtain image detection results. These results can be labels used to characterize whether the wound images meet the requirements. For example, the image detection results can be "meets requirements" or "does not meet requirements."
[0187] In practice, firstly, the aforementioned executing entity can convert the wound image into a grayscale image using a grayscale conversion function. This grayscale conversion function can be any function capable of converting a color image to grayscale. For example, the grayscale conversion function could be `cv2.cvtColor()`.
[0188] Secondly, the pixel values in the aforementioned grayscale wound image can be converted to 0 or 255 using a binarization function to perform binarization processing on the grayscale wound image, resulting in a binarized grayscale wound image. The binarization function can be any function capable of performing image binarization processing. For example, the binarization function could be cv2.adaptiveThreshold().
[0189] Then, the pixel values with a value of 0 among the pixel values included in the above binarized image can be determined as the minimum pixel values. Then, the number of these minimum pixel values can be determined as the minimum number of pixels.
[0190] Then, the pixel values with a value of 255 among the pixel values included in the above binarized image can be determined as the maximum pixel values. Then, the number of these maximum pixel values can be determined as the maximum number of pixels.
[0191] Then, the ratio of the maximum number of pixels to the minimum number of pixels can be determined as the pixel ratio data. In response to determining that the pixel ratio data is greater than a preset ratio threshold, "meets requirements" can be determined as the image detection result. The ratio threshold can be a preset value. Here, the specific setting of the ratio threshold is not limited. In response to determining that the pixel ratio data is less than or equal to the ratio threshold, "does not meet requirements" can be determined as the image detection result.
[0192] The fourth sub-step involves controlling the intelligent central control terminal to perform a wound imaging guidance task in response to the determination that the above image detection results do not meet the preset wound image conditions, thereby obtaining wound image data.
[0193] The aforementioned wound image conditions can be defined as an image detection result that "meets requirements." The aforementioned wound imaging guidance task can be a task that prompts the user to re-image the wound. The aforementioned wound image data can be image data of the wound re-imaged after the aforementioned wound imaging guidance task. The aforementioned wound imaging guidance task can be a task that guides the user to re-image the wound. For example, the aforementioned wound imaging guidance task can be: a voice prompt to the user to re-image the wound.
[0194] In practice, the aforementioned executing entity can send a preset shooting guidance instruction to the aforementioned intelligent central control terminal to control the intelligent central control terminal to perform the wound shooting guidance task. Then, it can receive image data sent by the aforementioned intelligent central control terminal and identify the image data sent by the intelligent central control terminal as wound image data. The aforementioned shooting guidance instruction can be a command prompting the intelligent central control terminal to guide the user to take wound images again.
[0195] The fifth sub-step involves performing region of interest (ROI) identification processing on the aforementioned wound image data to obtain the wound image region. This wound image region can be the image area of the wound within the aforementioned wound image data.
[0196] In practice, the aforementioned execution entity can input the wound image data into a pre-trained region recognition model to obtain wound bounding box data. The region recognition model can be a neural network model that takes wound image data as input and outputs wound bounding box data. For example, the region recognition model can be a pre-trained Faster R-CNN. Pre-training can be a process of fine-tuning Faster R-CNN using an annotated wound image dataset and a cross-entropy loss function. The wound bounding box data can be data used to characterize the position of the wound bounding box in the wound image data. The wound bounding box can be a bounding box used to label the image region where the wound is located. The wound bounding box data can include the coordinates of the top-left corner of the wound, the coordinates of the bottom-right corner of the wound, the wound category label, and the confidence score. The coordinates of the top-left corner of the wound can be the pixel coordinates of the top-left vertex of the wound bounding box in the wound image data. The coordinates of the bottom-right corner of the wound can be the pixel coordinates of the bottom-right vertex of the wound bounding box in the wound image data. The wound category label mentioned above can be a label used to characterize the category of the image region marked by the wound bounding box. For example, the wound category label could be "wound". The confidence level mentioned above can characterize the reliability of the wound bounding box data.
[0197] Then, the coordinates of the upper left corner and the lower right corner of the wound can be combined into the trimming parameters. For example, when the coordinates of the upper left corner of the wound are (1,3) and the coordinates of the lower right corner of the wound are (5,6), the corresponding trimming parameters can be "(1,3,5,6)".
[0198] Next, the cropping parameters can be input into the cropping function to crop the image region corresponding to the cropping parameters from the wound image data, and the cropped image region can be identified as the wound image region. The cropping function can be any function capable of cropping a specified image region from the image data. For example, the cropping function could be the crop() function.
[0199] The sixth sub-step involves feature extraction processing of the aforementioned wound image region to obtain wound feature information. This wound feature information can be a feature vector corresponding to the wound image region. In practice, the executing entity can use a pre-trained feature extraction model to perform feature extraction processing on the aforementioned wound image region to obtain wound feature information. This feature extraction model can be a neural network model that takes the wound image region as input and outputs the wound feature information. For example, the feature extraction model can be a pre-trained convolutional neural network. The pre-training process can be a fine-tuning of the convolutional neural network using a labeled set of wound image regions and a cross-entropy loss function.
[0200] The seventh sub-step involves generating fused feature information based on the aforementioned user physical indicator data and wound feature information. This fused feature information can be a feature vector representing the features of the aforementioned user physical indicator data and wound feature information.
[0201] In practice, firstly, the aforementioned executing entity can combine the user's physical indicators, including blood pressure, body temperature, heart rate, blood glucose, and blood oxygen saturation, into a row vector as the indicator vector. Then, the indicator vector and the wound feature information can be input into a feature concatenation function to obtain the output data of the feature concatenation function as the fused feature information. The feature concatenation function can be any function capable of concatenating two vectors. For example, the feature concatenation function could be the `concatenate()` function.
[0202] The eighth sub-step involves generating user physical status information based on the aforementioned fused feature information. This user physical status information can be a label representing the user's physical health status. For example, the user physical status information could be "good," "moderate," or "poor."
[0203] In practice, firstly, the aforementioned executing entity can generate the target physical health data of the aforementioned fused feature information through a pre-defined correspondence table. The target physical health data can be the physical health data corresponding to the aforementioned fused feature information. This physical health data can be a numerical value used to characterize the user's physical condition. The aforementioned correspondence table can be a table pre-defined by technicians based on statistics of a large number of fused feature vectors and physical health data, storing the correspondence between each fused feature vector and the physical health data. The aforementioned correspondence table can include various fused feature vectors. Each of the aforementioned fused feature vectors can be pre-generated fused feature information. Each of the aforementioned fused feature vectors corresponds to one piece of physical health data.
[0204] As an example, firstly, for each fused feature vector in the aforementioned correspondence table, the cosine similarity between the fused feature information and the fused feature vector can be determined as the feature similarity. Then, the feature similarity with the largest value among the determined feature similarities can be determined as the target feature similarity. Next, the fused feature vector corresponding to the target feature similarity can be determined as the target fused feature vector. Finally, the health data corresponding to the target fused feature vector can be determined as the target health data corresponding to the aforementioned fused feature information.
[0205] Then, in response to determining that the target health data is greater than a first health threshold, "good" can be defined as the user's health status information. The first health threshold can be a pre-set value. Here, the specific setting of the first health threshold is not limited. In response to determining that the target health data is less than or equal to the first health threshold and greater than a second health threshold, "moderate" can be defined as the user's health status information. The second health threshold can be a pre-set value. Here, the specific setting of the second health threshold is not limited. In response to determining that the target health data is less than or equal to the second health threshold, "poor" can be defined as the user's health status information.
[0206] The ninth sub-step involves controlling the intelligent central control terminal to execute a status warning task in response to determining that the user's physical condition information meets a preset status warning condition. The status warning condition can be that the user's physical condition information is "poor." In practice, the executing entity can send a preset warning instruction to the intelligent central control terminal to control it to suspend the execution of the target follow-up plan and send the user's physical condition information to the target terminal, thereby controlling the intelligent central control terminal to execute the status warning task. The warning instruction can be an instruction used to prompt the intelligent central control terminal to issue a warning and temporarily stop the execution of the target follow-up plan. For example, the warning instruction could be "warning:on," indicating that a warning has been initiated and the execution of the target follow-up plan has been temporarily suspended.
[0207] The above technical solution and its related content, combined with step 106, serve as an inventive point of this disclosure, solving the problem of "poor operational convenience." Factors leading to poor operational convenience often include: operating the intelligent central control terminal to perform rehabilitation tasks typically involves cumbersome steps, and users often find it difficult to actively interact with the intelligent central control terminal due to mobility limitations, resulting in poor convenience for users when operating the intelligent central control terminal to perform rehabilitation tasks. Solving these factors can improve operational convenience. To achieve this effect, this disclosure first generates today's task information based on real-time acquired time data and the aforementioned target follow-up plan. This allows for the determination of the rehabilitation plan the user needs to complete today. Second, in response to determining that the above-mentioned today's task information meets preset risk detection conditions, the following steps are executed: Then, the intelligent central control terminal is controlled to perform a body indicator monitoring task to obtain the user's body indicator data. This allows for the monitoring of the user's body indicator data. Then, the intelligent central control terminal is controlled to perform a wound imaging task to obtain wound imaging images. This allows for the imaging of the user's wound. Finally, the wound imaging images are subjected to image detection processing to obtain image detection results. Therefore, it can be determined whether the wound image captured meets the requirements. Then, in response to the determination that the image detection result does not meet the preset wound image conditions, the intelligent central control terminal is controlled to execute the wound imaging guidance task to obtain wound image data. Thus, wound image data after re-capture can be obtained. Then, region of interest identification processing is performed on the wound image data to obtain the wound image region. Thus, the image region where the wound is located can be identified. Next, feature extraction processing is performed on the wound image region to obtain wound feature information. Thus, feature information corresponding to the wound image region can be obtained. Then, based on the user's physical indicator data and the wound feature information, fused feature information is generated. Thus, fused feature information can be obtained. Next, based on the fused feature information, user physical status information is generated. Thus, user physical status information can be obtained. Then, in response to the determination that the user's physical status information meets the preset status warning conditions, the intelligent central control terminal is controlled to execute the status warning task. Thus, when the user's physical status information indicates that the user's current physical condition is poor, the intelligent central control terminal can be controlled to issue a warning. Because it can control the intelligent central control terminal to actively monitor the user's physical indicators and actively capture wound images, thereby generating the user's physical status information, and actively perform corresponding operations based on the user's physical status information, without waiting for the user to perform cumbersome operation steps, it can significantly reduce the operating threshold of the intelligent central control terminal and significantly improve the convenience for users when operating the intelligent central control terminal to perform rehabilitation tasks.
[0208] The above embodiments of this disclosure have the following beneficial effects: the intelligent central control terminal control method based on a multimodal large model in some embodiments of this disclosure can improve operational convenience. The reason for poor operational convenience is that users need to use complex touch and other interactive modes to control the intelligent central control terminal to perform auxiliary tasks, making the operation process cumbersome and resulting in poor convenience for users controlling the intelligent central control terminal for rehabilitation training. Based on this, the intelligent central control terminal control method based on a multimodal large model in some embodiments of this disclosure firstly, in response to receiving a task execution signal sent by the intelligent central control terminal, performs node fusion processing on the pre-acquired user electronic medical records and thyroid cancer knowledge graph to update the thyroid cancer knowledge graph. This allows for the updating of the thyroid cancer knowledge graph. Secondly, based on the updated thyroid cancer knowledge graph, the user electronic medical records undergo feature node mapping processing to obtain a user knowledge subgraph. This allows for the obtaining of a user knowledge subgraph. Then, the user knowledge subgraph undergoes node attention aggregation processing to obtain each meta-path subgraph and each subgraph node feature information. Therefore, the feature information of each meta-path subgraph and each subgraph node can be obtained. Then, gating enhancement processing is performed on the aforementioned user electronic medical record, the aforementioned meta-path subgraph, and the aforementioned subgraph node feature information to obtain user feature information. Then, based on a multimodal large model used to process thyroid cancer data, the aforementioned user feature information is subjected to scheme matching and adaptive optimization processing to obtain a target follow-up scheme. Finally, based on the aforementioned target follow-up scheme, the aforementioned intelligent central control terminal is controlled to perform task-assisted operations. Because the intelligent central control terminal can actively assist the user in completing the rehabilitation plan based on the generated target follow-up scheme, rather than passively waiting for user interaction, the number of interaction steps required by the user is reduced, improving the ease of operation for the user when controlling the intelligent central control terminal for rehabilitation training.
[0209] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an intelligent central control terminal control device based on a multimodal large model. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0210] like Figure 2As shown, some embodiments of the intelligent central control terminal control device 200 based on a multimodal large model include: a node fusion unit 201, a feature node mapping unit 202, a node attention aggregation unit 203, a gating enhancement unit 204, an adaptive optimization unit 205, and a control unit 206. The system includes the following components: a node fusion unit 201, configured to perform node fusion processing on pre-acquired user electronic medical records and a thyroid cancer knowledge graph in response to a task execution signal received from the intelligent central control terminal, thereby updating the thyroid cancer knowledge graph; a feature node mapping unit 202, configured to perform feature node mapping processing on the aforementioned user electronic medical records based on the updated thyroid cancer knowledge graph, to obtain a user knowledge subgraph; a node attention aggregation unit 203, configured to perform node attention aggregation processing on the aforementioned user knowledge subgraph, to obtain each meta-path subgraph and each subgraph node feature information; a gating enhancement unit 204, configured to perform gating enhancement processing on the aforementioned user electronic medical records, the aforementioned meta-path subgraphs, and the aforementioned subgraph node feature information, to obtain user feature information; an adaptive optimization unit 205, configured to perform scheme matching and adaptive optimization processing on the aforementioned user feature information based on a multimodal large model used for processing thyroid cancer data, to obtain a target follow-up scheme; and a control unit 206, configured to control the aforementioned intelligent central control terminal to perform task assistance operations based on the aforementioned target follow-up scheme.
[0211] It is understandable that the units described in the device 200 are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.
[0212] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device (such as a computing device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0213] like Figure 3 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0214] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0215] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0216] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0217] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0218] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: respond to receiving a task execution signal sent by the intelligent central control terminal, perform node fusion processing on the pre-acquired user electronic medical records and thyroid cancer knowledge graph to update the thyroid cancer knowledge graph; based on the updated thyroid cancer knowledge graph, perform feature node mapping processing on the aforementioned user electronic medical records to obtain a user knowledge subgraph; perform node attention aggregation processing on the aforementioned user knowledge subgraph to obtain each meta-path subgraph and each subgraph node feature information; perform gating enhancement processing on the aforementioned user electronic medical records, the aforementioned meta-path subgraphs, and the aforementioned subgraph node feature information to obtain user feature information; based on a multimodal large model used for processing thyroid cancer data, perform scheme matching and adaptive optimization processing on the aforementioned user feature information to obtain a target follow-up scheme; and based on the aforementioned target follow-up scheme, control the aforementioned intelligent central control terminal to perform task assistance operations.
[0219] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0220] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0221] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a node fusion unit, a feature node mapping unit, a node attention aggregation unit, a gating enhancement unit, an adaptive optimization unit, and a control unit. The names of these units do not necessarily limit the specific unit; for example, a control unit may also be described as "a unit that controls the intelligent central control terminal to perform task-assisted operations based on the aforementioned target follow-up scheme."
[0222] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0223] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A control method for an intelligent central control terminal based on a multimodal large model, comprising: In response to receiving the task execution signal sent by the intelligent central control terminal, the node fusion processing is performed on the pre-acquired user electronic medical records and thyroid cancer knowledge graph to update the thyroid cancer knowledge graph; Based on the updated thyroid cancer knowledge graph, feature node mapping processing is performed on the user's electronic medical record to obtain a user knowledge subgraph. The user knowledge subgraph is subjected to node attention aggregation processing to obtain each meta-path subgraph and the node feature information of each subgraph; Gated enhancement processing is performed on the user's electronic medical record, each metapath subgraph, and the feature information of each subgraph node to obtain user feature information; Based on a multimodal large model for processing thyroid cancer data, the user feature information is subjected to scheme matching and adaptive optimization to obtain a target follow-up scheme. Based on the target follow-up plan, the intelligent central control terminal is controlled to perform task assistance operations.
2. The method of claim 1, wherein, Based on the updated thyroid cancer knowledge graph, feature node mapping processing is performed on the user's electronic medical record to obtain a user knowledge subgraph, including: Key entities are extracted from the user's electronic medical record to obtain the user's identity entity; The user identity entity is standardized to obtain a standardized identity entity; Based on the standardized identity entities, a breadth-first traversal is performed on the updated thyroid cancer knowledge graph to obtain each seed node; For each of the seed nodes, based on the updated thyroid cancer knowledge graph, the path of the seed node is expanded to obtain an expanded node set; Based on the obtained sets of extended nodes, subgraph query and relationship reconstruction are performed on the updated thyroid cancer knowledge graph to obtain the user knowledge subgraph.
3. The method of claim 1, wherein, The user's electronic medical record includes surgical records, user pathology reports, pathology slide images, and preoperative laboratory data; The process of performing node fusion processing on pre-acquired user electronic medical records and thyroid cancer knowledge graphs to update the thyroid cancer knowledge graph includes: The surgical record and the user pathology report are preprocessed to obtain preprocessed text data. The pathological slide images are subjected to microscopic feature quantization processing to obtain image microscopic feature data; The preoperative laboratory data is parsed and processed to obtain structured laboratory data; Entity recognition processing is performed on the preprocessed text data, the image microscopic feature data, and the laboratory structured data to obtain various text entities, wherein each text entity corresponds to an entity category label; The text entities are standardized and mapped to obtain standardized text entities, wherein each standardized text entity corresponds to node feature information. Based on the standardized text entities, entity alignment processing is performed on the thyroid cancer knowledge graph to obtain the entity alignment results. Based on the alignment results of each entity, the standardized text entities and the thyroid cancer knowledge graph are fused to update the thyroid cancer knowledge graph.
4. The method according to claim 1, wherein, The user knowledge subgraph includes nodes and edges, and each node has corresponding node feature information. The node attention aggregation processing performed on the user knowledge subgraph to obtain each meta-path subgraph and each subgraph node feature information includes: For each path template in the pre-acquired path templates, perform the following steps: Based on the user knowledge subgraph, pattern matching is performed on the path template to obtain a meta-path subgraph, wherein the meta-path subgraph includes various nodes, and each node in each node corresponds to node feature information; Nodes that meet the preset instance conditions among the nodes included in the meta-path subgraph are identified as instance nodes. Each neighbor node corresponding to the instance node is determined as the instance neighbor node; For each instance neighbor node in the various instance neighbor nodes, the instance neighbor node and the instance node are concatenated with node semantic features to obtain concatenated node feature information; The feature information of each splicing node is processed by feature dimension mapping to obtain the mapped feature information. Based on the preset attention vector and the various mapping feature information, various attention weights are generated, wherein each attention weight corresponds to an instance neighbor node; Based on the attention weights and the neighboring nodes of each instance, subgraph node feature information is generated.
5. The method according to claim 1, wherein, The gating enhancement process is applied to the user's electronic medical record, each metapath subgraph, and the feature information of each subgraph node to obtain user feature information, including: For each metapath subgraph in the various metapath subgraphs, determine the path weight data corresponding to the metapath subgraph; Based on the determined path weight data, the feature information of each subgraph node is weighted and fused to obtain fused feature information. The user's electronic medical record and the fused feature information are subjected to feature extraction and fusion processing to obtain clinical splicing feature information; The clinical splicing feature information is subjected to parameter weighting processing to obtain a parameter weighting vector; The sum of the parameter weighted vector and the preset bias vector is determined as the bias weighted vector; The bias weighted vector is subjected to variable mapping to obtain the gated value vector; Based on the gated value vector, generate a gated value negative vector; The Hadamard product of the gated value vector and the fused feature information is determined as the first linear vector; The Hadamard product of the negative gate value vector and the parameter weighted vector is determined as the second linear vector; The sum of the first linear vector and the second linear vector is determined as the user feature information.
6. The method according to claim 1, wherein, The multimodal large model used for processing thyroid cancer data performs scheme matching and adaptive optimization on the user feature information to obtain the target follow-up scheme, including: Based on the user characteristic information and the preset standard scheme knowledge base, a follow-up template scheme is generated; Based on the aforementioned follow-up template scheme and the preset scheme standard data table, standard health data are determined; Based on the preset adjustment prompts, the follow-up template, the standard health data, and the user's electronic medical record, a long-term follow-up plan is generated. The user feature information is standardized to obtain standard feature information; The long-term follow-up scheme is segmented and mapped to obtain a follow-up segmented data sequence; The standard feature information and the follow-up word segmentation data sequence are input into the multimodal large model to obtain the feature score matrix; Based on the feature score matrix, a target follow-up plan is generated.
7. The method according to claim 6, wherein, The multimodal large model is obtained through the following steps: Obtain a sample set, wherein each sample in the sample set includes sample standard feature information, sample follow-up word segmentation data sequence, and sample label sequence; Based on the sample set, perform the following training steps: The standard feature information of each sample and the follow-up word segmentation data sequence of each sample in the sample set are input into the input layer of the initial network model to obtain the follow-up word segmentation matrix, feature embedding information and text embedding matrix corresponding to each sample; Based on the multimodal fusion coding layer, feature-aware encoder, conditional decoder, follow-up word segmentation matrix, feature embedding information and text embedding matrix corresponding to each of the at least one sample in the initial network model, at least one decoding feature matrix is generated; The at least one decoded feature matrix corresponding to the at least one sample is respectively input into the output layer of the initial network model to obtain at least one feature score matrix; The feature score matrix corresponding to each sample in the at least one sample is compared with the corresponding sample number sequence to obtain at least one comparison result; In response to determining that at least one alignment result has reached the optimization objective, the initial network model is determined as a multimodal large model.
8. A smart central control terminal control device based on a multimodal large model, comprising: The node fusion unit is configured to perform node fusion processing on the pre-acquired user electronic medical records and thyroid cancer knowledge graph in response to receiving a task execution signal sent by the intelligent central control terminal, so as to update the thyroid cancer knowledge graph. The feature node mapping unit is configured to perform feature node mapping processing on the user's electronic medical record based on the updated thyroid cancer knowledge graph to obtain a user knowledge subgraph. The node attention aggregation unit is configured to perform node attention aggregation processing on the user knowledge subgraph to obtain each meta-path subgraph and each subgraph node feature information. The gating enhancement unit is configured to perform gating enhancement processing on the user's electronic medical record, the various metapath subgraphs, and the feature information of the nodes of the various subgraphs to obtain user feature information. An adaptive optimization unit is configured to perform scheme matching and adaptive optimization processing on the user feature information based on a multimodal large model used for processing thyroid cancer data, to obtain a target follow-up scheme. The control unit is configured to control the intelligent central control terminal to perform task assistance operations based on the target follow-up scheme.
9. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 7.
10. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Interpretable smart medical auxiliary diagnosis system based on text retrieval
CN112687388A
Traffic flow prediction method based on multimode dynamic memory graph convolutional network
CN119107798A
Feature fusion method based on multi-modal medical data
CN119557840A
Medical decision support system based on knowledge graph
CN121215159A
Multi-task collaborative prediction medical follow-up visit method and system, terminal and medium
CN121237339A