Method and device for training large model in field of health care and value system identification application
By constructing a high-dimensional health and wellness knowledge tensor and generating a low-dimensional core tensor and factor matrix, the training data retrieval of the large-scale health and wellness model is optimized, solving the problems of data redundancy and computational resource consumption, and improving training efficiency and performance.
Patent Information
- Application Number
- CN202510974735.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing large-scale models face challenges in training in the health and wellness field, including high data redundancy, enormous storage and computing resource consumption, and difficulty in learning accurate, specialized, and structured knowledge from general text data.
A high-dimensional health and wellness knowledge tensor is constructed, and a low-dimensional core tensor and multi-factor matrix are generated through iterative projection to form a compressed knowledge base component. Furthermore, training data retrieval is optimized by sampling weights of task-related data to achieve on-demand decompression data retrieval.
This reduces storage and computational load, improves the training efficiency and overall performance of the health and wellness model, and ensures that each subtask obtains training samples with the highest information density.
Smart Images

Figure CN120874990A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a training method, apparatus, and value system recognition application for a large-scale model in the field of health and wellness. Background Technology
[0002] The field of large-scale model technology specifically refers to the collection of technologies for developing and applying artificial intelligence models with massive parameters (typically reaching billions or even trillions) and complex structures. The core characteristic of this field is that, through self-supervised pre-training on massive and diverse unlabeled data, the models are able to learn general world knowledge and deep-level pattern representation capabilities.
[0003] Existing technologies have limitations when training large-scale models in specific domains such as health and wellness. Current large-scale models typically rely on general, unlabeled text data for pre-training. While this approach endows the model with broad global knowledge, its depth and accuracy are insufficient when facing highly specialized, structured, and knowledge-intensive scenarios like health and wellness. For example, when dealing with complex interactions between drugs, genes, and diseases, general text corpora contain a large amount of scattered and even contradictory information, making it difficult for the model to directly learn accurate and reliable causal chains. Directly using massive amounts of raw medical literature or database records as training input not only results in high data redundancy and enormous storage and I / O costs, but also requires the model to expend significant computational resources to sift through noise and learn effective relationships. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a training method, device, and value system identification application for a large-scale model in the field of health and wellness.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a training method for a large-scale model in the field of health and wellness, comprising the following steps: Based on UMLS entity relationship data, DrugBank drug target data and clinical symptom records, a high-dimensional health and wellness knowledge tensor is constructed by assigning a unique index to each drug, gene and disease entity and assigning dimensional coordinates to each relationship type, mapping the existence and credibility values of the relationship between entities to the target position in the tensor space composed of the index and coordinates. Based on the high-dimensional health and wellness knowledge tensor, iterative projection is used to project the tensor sequentially along each dimension of drug, gene, disease, and relationship type to generate a low-dimensional factor matrix corresponding to each dimension. The low-dimensional core tensor connecting the multi-factor matrices is calculated, and the core tensor and the multi-factor matrices are combined into a compressed knowledge base component.
[0006] Preferably, the method further includes: Construct a correlation matrix between health and wellness sub-tasks and data features. Fill the data feature correlation matrix with expert-annotated values to quantify the correlation strength of data features. When a training task is specified, extract the correlation strength row vector corresponding to the task from the data feature correlation matrix as the sampling weight of the task-related data. The system receives training query requests and converts them into tensor slice coordinates. It then applies the sampling weights of the task-related data to adjust the selection probability of the factor matrix row vectors. Through matrix multiplication, it calculates the target index range of the core tensor on the selected factor matrix row vectors. It extracts the core tensor elements and associated factor matrix row vectors within the target index range from the compressed knowledge base component, performs tensor shrinking and reconstructs the required data, and generates an instant training data subset.
[0007] Preferably, the steps for obtaining the high-dimensional health and wellness knowledge tensor are as follows: Based on UMLS entity relationship data, DrugBank drug target data and clinical symptom records, we integrate drug entity, gene entity and disease entity lists from the three data sources, remove duplicate entities, assign each cleaned entity an integer sequence number starting from zero and incrementing, and build an entity index table. Based on the entity index table, coordinate axes are set for four dimensions: drug, gene, disease, and relationship type. The index value of each entity in the entity index table and the preset relationship type identifier are combined into a quadruple to form a multidimensional space coordinate set. Based on the multidimensional spatial coordinate set, the credibility values of the relationships between entities in the original data are read, and the credibility values are directly written into the tensor space position specified by the corresponding quadruple in the multidimensional spatial coordinate set. For coordinate positions without records, zero values are filled in to construct a high-dimensional health and wellness knowledge tensor.
[0008] Preferably, the step of obtaining the compressed knowledge base component specifically includes: Based on the high-dimensional health and wellness knowledge tensor, a random factor matrix is initialized. Through iterative optimization using the alternating least squares method, the modal expansion matrix of the high-dimensional health and wellness knowledge tensor and the Khatri-Rao product of the factor matrices other than the current dimension are solved. The factor matrix is updated dimension by dimension, generating a set of four converged low-dimensional factor matrices. Based on the set of low-dimensional factor matrices, the Moore-Penrose pseudo-inverse of each matrix in the set of low-dimensional factor matrices is calculated. The high-dimensional health and wellness knowledge tensor is then subjected to tensor inner product operation with the four calculated pseudo-inverse matrices one by one to obtain the low-dimensional core tensor. Based on the low-dimensional core tensor and the set of low-dimensional factor matrices, a data structure is created. The data structure contains four matrices from the low-dimensional core tensor and the set of low-dimensional factor matrices. The data structure is then serialized and written to a disk file to establish a compressed knowledge base component.
[0009] Preferably, the step of obtaining the sampling weights of the task-related data specifically includes: Construct a correlation matrix between health and wellness sub-tasks and data features, where rows represent tasks and columns represent data features. Domain experts score the correlation strength of each task-feature pair to generate a task-data correlation strength matrix. Based on the task data association strength matrix, when a specified cardiovascular risk assessment training task is received, the row index of the task data association strength matrix is scanned to locate the row that completely matches the task name, and the values of all columns in that row are extracted to obtain the association strength row vector.
[0010] Preferably, the step of obtaining the task-related data sampling weights further includes: Based on the association strength row vector, each value in the vector is directly used as the sampling probability of the corresponding data feature to generate task association data sampling weights.
[0011] Preferably, the step of obtaining the subset of training data in real time specifically includes: Receive training query requests, parse the drug, gene and disease entities specified in the request, query the entity index table to obtain the corresponding index, and combine the index and the relationship type identifier of the request into a slice range in tensor space to establish a tensor slice coordinate set. Based on the tensor slice coordinate set, the task-related data sampling weights are applied to the row vector selection of the factor matrix. By performing matrix multiplication between the tensor slice coordinate set and the weighted factor matrix, the index range of the corresponding elements in the core tensor is calculated, and the target index range is obtained.
[0012] Preferably, the step of obtaining the subset of training data in real time further includes: Based on the target index range, the core tensor element blocks within the target index range are indexed and extracted from the compressed knowledge base component. At the same time, the factor matrix row vectors associated with the query are extracted. Tensor product operation is performed to reconstruct the core tensor element blocks and factor matrix row vectors into a dense subtensor, generating an instant training data subset.
[0013] The present invention also provides a training device for a large-scale model in the field of health and wellness. The training device includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, a training method for the large-scale model in the field of health and wellness is implemented.
[0014] This invention also provides a method for identifying the value system of a large-scale model in the field of health and wellness. The large-scale model in the field of health and wellness is obtained based on the training method of a large-scale model in the field of health and wellness. The corpus text includes one or more of the following: health and wellness materials related to value understanding, basic value system theoretical content, questionnaires related to the value system, and interview materials related to the value system.
[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: This invention unifies heterogeneous health and wellness knowledge from multiple sources, such as UMLS entity relationships, DrugBank drug targets, and clinical symptom records, into a high-dimensional health and wellness knowledge tensor containing dimensions of drugs, genes, diseases, and relationship types, thus constructing a structured knowledge carrier. Furthermore, iterative projection is used to decompose this high-dimensional tensor into a low-dimensional core tensor and a multi-factor matrix, forming a compressed knowledge base component. This operation reduces the physical overhead of storing large-scale relational data and optimizes the data access structure. During training data retrieval, query requests are converted into tensor slice coordinates, and calculations are performed directly on the compressed factor matrix. Only the core tensor regions relevant to the query are reconstructed using inverse operations, achieving on-demand decompression of data retrieval. This query-as-decompression mechanism avoids loading and decompressing the entire dataset, reducing I / O load and memory consumption during training. Furthermore, by constructing a correlation matrix between tasks and data features, data sampling weights are dynamically generated for different health and wellness sub-tasks. This enables the system to prioritize the retrieval of the most relevant data based on the specific training task, such as cardiovascular risk assessment or depression screening, ensuring that each sub-task obtains training samples with the highest information density. This, in turn, improves the overall performance and training efficiency of the health and wellness big data model. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the steps of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0018] Please see Figure 1 This invention provides a technical solution, a training method for a large-scale model in the field of health and wellness, comprising the following steps: Based on UMLS entity relationship data, DrugBank drug target data and clinical symptom records, a high-dimensional health and wellness knowledge tensor is constructed by assigning a unique index to each drug, gene and disease entity and assigning dimensional coordinates to each relationship type, mapping the existence and credibility values of the relationship between entities to the target position in the tensor space composed of the index and coordinates. Based on the high-dimensional health and wellness knowledge tensor, iterative projection is used to project the tensor along each dimension of drug, gene, disease and relationship type in sequence to generate a low-dimensional factor matrix corresponding to each dimension. The low-dimensional core tensor connecting the multi-factor matrix is calculated, and the core tensor and the multi-factor matrix are combined into a compressed knowledge base component. Construct a correlation matrix between health and wellness sub-tasks and data features. Fill the data feature correlation matrix with expert-annotated values to quantify the correlation strength of data features. When a training task is specified, extract the correlation strength row vector corresponding to the task from the data feature correlation matrix as the sampling weight of the task-related data. The system receives training query requests and converts them into tensor slice coordinates. It applies task-related data sampling weights to adjust the selection probability of factor matrix row vectors. It calculates the target index range of the core tensor on the selected factor matrix row vectors through matrix multiplication. It extracts the core tensor elements and associated factor matrix row vectors within the target index range from the compressed knowledge base component, performs tensor shrinking and reconstructs the required data, and generates an instant training data subset.
[0019] The specific steps for obtaining the high-dimensional health and wellness knowledge tensor are as follows: Based on UMLS entity relationship data, DrugBank drug target data and clinical symptom records, we integrate drug entity, gene entity and disease entity lists from the three data sources, remove duplicate entities, assign each cleaned entity an integer sequence number starting from zero and incrementing, and build an entity index table. Based on the entity index table, coordinate axes are set for four dimensions: drug, gene, disease, and relationship type. The index value of each entity in the entity index table and the preset relationship type identifier are combined into a quadruple to form a multidimensional space coordinate set. Based on a multidimensional spatial coordinate set, the credibility values of the relationships between entities in the original data are read, and the credibility values are directly written into the tensor space location specified by the corresponding quadruples in the multidimensional spatial coordinate set. For coordinate locations without records, zero values are filled in to construct a high-dimensional health and wellness knowledge tensor.
[0020] Specifically, based on UMLS (Unified Medical Language System) entity relation data, DrugBank drug target data, and clinical symptom records, the process first extracts Concept Unique Identifiers (CUIs) and their corresponding drug, gene, and disease entity names from UMLS. It then parses approved drugs, experimental drugs, and their known gene targets from the DrugBank database. Finally, it extracts symptom entities from structured or unstructured clinical symptom record text using natural language processing techniques (e.g., using a BERT-based named entity recognition model, pre-trained on general medical corpora and fine-tuned on task-specific corpora such as MIMIC-III; input is symptom description text, output is identified symptom entities). The collected three categories of... The data source integrates lists of drug entities (e.g., aspirin, imatinib), gene entities (e.g., PTGS1, BCR-ABL), and disease entities (e.g., pain, chronic myeloid leukemia). The integrated entity list undergoes standardization, for example, mapping generic and brand names of drugs to unique identifiers, and normalizing different disease descriptions (e.g., "heart attack" and "myocardial infarction") by mapping them to UMLS CUIs. Duplicate entities are then removed by comparing unique identifiers (e.g., CUI or DrugBankID) to ensure the uniqueness of each entity in the list. Finally, each drug entity, gene entity, and disease entity in the cleaned list is independently assigned a unique index, starting from zero and incrementing sequentially. For example, if a drug entity has... If there are 1, then its index range is 1. arrive Similarly, it is a genetic entity ( One, index arrive ) and disease entity ( One, index arrive Allocate indexes, store these entities and their corresponding indexes, and create an entity index table.
[0021] Based on an entity index table, logical coordinate axes are set for four dimensions: drug, gene, disease, and relation type. The scale of the drug axis corresponds to the index value of the drug entity, the scale of the gene axis corresponds to the index value of the gene entity, and the scale of the disease axis corresponds to the index value of the disease entity. The relation type dimension is set according to a predefined set of relations and their corresponding numerical identifiers. For example, the preset relation types include "treatment" (identifier 0), "cause" (identifier 1), "drug-targeted gene" (identifier 2), "gene-disease related" (identifier 3), and "drug-inhibited gene" (identifier 4). These relation types and their identifiers are manually organized and solidified into a mapping table through domain expert knowledge and analysis of existing medical knowledge bases. This mapping table records information such as "relationship text description." The process involves mapping the relation to the "relation type identifier (integer)" and then iterating through each specific relation instance parsed from the original data, such as "drug A treats disease B" or "drug C targets gene D". It searches the entity index table for the integer index values of drug A, disease B, drug C, and gene D, and then searches the preset relation type mapping table for the relation type identifiers corresponding to "treatment" and "targeting". These index values and relation type identifiers are combined into a quadruple, specifically (drug entity index, gene entity index, disease entity index, relation type identifier). If a relation instance does not directly involve a certain dimension, such as "drug A treats disease B" not directly involving genes, then the index for that gene dimension can use a preset special placeholder index, for example... (i.e., the total number of gene entities, representing a value that is outside the range of the regular index) or -1, to indicate that this dimension is not applicable or not specified in the relation. By performing such a transformation on all the relations in the original data, a multidimensional space coordinate set consisting of these quadruples is formed.
[0022] Based on a multidimensional spatial coordinate set and raw relation data extracted from UMLS entity relation data, DrugBank drug target data, and clinical symptom records, the confidence value is first determined for each raw relation (e.g., "aspirin treats headache"). The method for obtaining this confidence value depends on the characteristics of the data source. For example, for drug-target relations in DrugBank, confidence can be quantified based on evidence level (e.g., preclinical studies = 0.3, Phase I clinical trials = 0.5, marketed = 0.9). For relations in UMLS, if a relevant confidence score exists, it is directly used; otherwise, a default value (e.g., 0.7) is set based on the authority of the relation source or its frequency of occurrence. For relations obtained through text mining in clinical symptom records (e.g., co-occurrence of symptom A and disease B), their confidence can be obtained by calculating Normalized Pointwise Mutual Information (NPMI) or by using the probability value output by a machine learning-based link prediction model. The value range is typically normalized to [value range missing]. The interval is defined, where higher values indicate higher confidence. Then, for each quadruplet coordinate in the multidimensional space coordinate set... The coordinate corresponds to a specific relation instance in the original data. The confidence value of the previously determined relation instance (e.g., 0.85) is directly written into a four-dimensional tensor initialized to all zeros (i.e., the high-dimensional health and wellness knowledge tensor) at the corresponding position specified by the quadruple, i.e., Tensor[idx_drug, idx_gene, idx_disease, idx_relation_type]=confidence_value. For coordinate positions that exist in the multidimensional coordinate set but are not explicitly recorded in the original data as corresponding relations or confidence data, their default padding value of zero is maintained. By performing the above assignment operation on the coordinates in all coordinate sets, the high-dimensional health and wellness knowledge tensor is constructed.
[0023] The specific steps for obtaining the compressed knowledge base component are as follows: Based on the high-dimensional health and wellness knowledge tensor, a random factor matrix is initialized. Through iterative optimization using the alternating least squares method, the modal expansion matrix of the high-dimensional health and wellness knowledge tensor and the Khatri-Rao product of the factor matrices other than the current dimension are solved. The factor matrix is updated dimension by dimension, generating a set of low-dimensional factor matrices after four convergence terms. Based on the set of low-dimensional factor matrices, the Moore-Penrose pseudo-inverse of each matrix in the set of low-dimensional factor matrices is calculated. The high-dimensional health and wellness knowledge tensor is then subjected to tensor inner product operation with the four calculated pseudo-inverse matrices one by one to obtain the low-dimensional core tensor. A data structure is created based on a low-dimensional core tensor and a set of low-dimensional factor matrices. The data structure contains four matrices from the low-dimensional core tensor and the set of low-dimensional factor matrices. The data structure is then serialized and written to a disk file to create a compressed knowledge base component.
[0024] Specifically, based on the high-dimensional health and wellness knowledge tensor (in (These represent the number of entities or types related to drugs, genes, diseases, and relationship types, respectively). First, determine the target low-dimensional rank of each factor matrix. For example, if there are 10,000 drugs, 20,000 genes, 5,000 diseases, and 20 types of relationships, then the corresponding low-dimensional rank can be empirically set to... The selection of these ranks is based on a trade-off between data sparsity, computational resources, and desired compression ratio and accuracy. They are typically selected through cross-validation, and then a four-factor matrix is initialized. , , , Typically, a standard normal distribution (mean 0, standard deviation 0.01) is used. The values are randomly sampled from the matrix and filled with the matrix. Then, iterative optimization is performed using the Alternating Least Squares (ALS) method. In each iteration, each factor matrix is updated sequentially. ( ), specific updates At the same time, keep the other three factor matrices intact. Fix, calculate intermediate tensors This intermediate tensor Along its first Each mode (corresponding to the original tensor) The (each mode) is expanded into a matrix Then, it is obtained through singular value decomposition (SVD). ,Pick The former Column as updated The process of "solving the Khatri-Rao product of the modal expansion matrix of the high-dimensional health and wellness knowledge tensor with the factor matrices other than the current dimension" is a typical description of updating the factor matrix in CP decomposition. In ALS of Tucker decomposition (such as the HOOI algorithm), the factor matrix is updated through SVD as mentioned above, and the iteration continues until the change in the factor matrix is less than a preset convergence threshold (e.g., the sum of the changes in the Frobenius norm of all factor matrices is less than 10 ... This threshold is set based on historical experience. When the change is less than this value, the model performance usually no longer improves significantly. Alternatively, it may reach the maximum number of iterations (e.g., 200 times, to prevent infinite loops and ensure that the algorithm ends within a reasonable time). Finally, a set of four converged low-dimensional factor matrices is generated.
[0025] Based on the set of low-dimensional factor matrices after four convergences ,in First, for each factor matrix in this set... Calculate its Moore-Penrose pseudoinverse The standard method for calculating the pseudo-inverse is to use singular value decomposition (SVD). ,but ,in It is The non-zero singular values in the matrix are obtained by taking the reciprocal and then transposing them. For example, for a factor matrix... Calculate its pseudo-inverse Subsequently, the original high-dimensional health and wellness knowledge tensor The four pseudo-inverse matrices calculated Perform tensor model one by one The product (TensorTimesMatrix, TTM) operation is mathematically expressed as follows: In practice, first calculate Then calculate Then calculate Finally, calculate The final tensor obtained This is the low-dimensional core tensor we are looking for.
[0026] Based on low-dimensional core tensor and low-dimensional factor matrix set First, a composite data structure is created in memory to store these components. This data structure can be a dictionary or a custom object, for example, a Python dictionary knowledge_components={'core_tensor': \mathcal{G}_{numpy\_array}, 'factor\_drug': \mathbf{A}^{(1)}_{numpy\_array}, 'factor\_gene': \mathbf{A}^{(2)}_{numpy\_array}, 'factor\_disease': \mathbf{A}^{(3)}_{numpy\_array}, 'factor\_relation': \mathbf{A}^{(4)}_{numpy_array}\}\), where the _numpy_array suffix indicates that these tensors and matrices are stored as NumPy arrays, ensuring that the data structure clearly identifies the core tensor and the four factor matrices corresponding to the dimensions of drug, gene, disease, and relation type, respectively. Subsequently, this complete data structure containing the low-dimensional core tensor and the four low-dimensional factor matrices is serialized. For example, in the Python environment, the dump method of the pickle library can be used to convert the dictionary object into a byte stream. Alternatively, if all components are NumPy arrays, the numpy.savez_compressed function can be used to efficiently save them to a .npz compressed archive file. The serialized byte stream is then completely written to a binary file under the specified disk path, for example, named compressed_knowledge_base.dat or tucker_model.npz, thereby establishing the compressed knowledge base components.
[0027] The specific steps for obtaining the sampling weights of task-related data are as follows: Construct a correlation matrix between health and wellness sub-tasks and data features, where rows represent tasks and columns represent data features. Domain experts score the correlation strength of each task-feature pair to generate a task-data correlation strength matrix. Based on the task data association strength matrix, when a specified cardiovascular risk assessment training task is received, the row index of the task data association strength matrix is scanned to locate the row that completely matches the task name, and the values of all columns in that row are completely extracted to obtain the association strength row vector. Based on the association strength row vector, each value in the vector is directly used as the sampling probability of the corresponding data feature to generate the task association data sampling weight.
[0028] Specifically, a correlation matrix between health and wellness sub-tasks and data features is constructed. First, the scope of the health and wellness sub-tasks is defined, such as "cardiovascular risk assessment," "diabetes management," "early screening for Alzheimer's disease," and "postoperative rehabilitation guidance." These task lists are predefined based on current research or application needs and form the row index of the matrix. Second, data features are defined. These features correspond to the potential information dimensions of the high-dimensional health and wellness knowledge tensor or their derived features, such as "efficacy data of specific drug classes (e.g., statins)," "disease association data of specific genes (e.g., APOE4)," "frequency of specific symptoms (e.g., chest tightness) and drug response," "drug-drug interaction data," and "lifestyle factors (e.g., smoking history, exercise levels)." These data features form the column index of the matrix. Then, at least three domain experts with more than five years of clinical or research experience in the corresponding health and wellness sub-task area are invited, such as cardiovascular... Internal medicine doctors, endocrinologists, neurologists, and rehabilitation physicians scored the association strength of each "task-feature" pair using a five-point Likert scale, where 1 represents "no association," 2 represents "weak association," 3 represents "moderate association," 4 represents "strong association," and 5 represents "very strong association." For example, for the "cardiovascular risk assessment" task and the "statin efficacy data" feature, experts gave 4 or 5 points, while for the "cardiovascular risk assessment" task and the "antidepressant side effect data" feature, experts gave 2 points. After collecting all experts' scores, the arithmetic mean of the multiple expert scores for each "task-feature" pair was calculated and rounded to one decimal place. For example, if three experts scored a certain combination as 3, 4, and 4, the average score would be 3.7. These average scores were filled into the corresponding cells of the matrix to form the task data association strength matrix.
[0029] Based on the task data association strength matrix, when the system receives a specific training task instruction, such as a request containing the task name "cardiovascular risk assessment" selected through the user interface or passed through an API call, it first performs a precise match search on the received task name. That is, it traverses the row labels (predefined task list) of the task data association strength matrix, and completely compares the input task name "cardiovascular risk assessment" with the task name string of each row in the matrix. Once a completely matching row is found, for example, there is a row in the matrix labeled "cardiovascular risk assessment", the system locks the row index of that row. Then, along the locked row index, it completely extracts the values corresponding to all columns in that row. These values are the average score of the association strength between the specific task (cardiovascular risk assessment) and each data feature, which was previously assessed by experts. For example, if the data feature columns include "statin drug efficacy data", "blood pressure monitoring data", "specific gene (such as KIF6) variation information", "lifestyle questionnaire results", etc., the extracted values are [4.5, 4.8, 3.5, 4.0, ...]. The extracted value sequence constitutes a one-dimensional array or vector, and the association strength row vector is obtained.
[0030] Based on the association strength row vector, for example, the association strength row vector obtained for the "cardiovascular risk assessment" task is: ,in Is the task and the first The association strength scores for each data feature (e.g., within the range of 1 to 5) are first normalized to convert them into an effective probability distribution. Specifically, if the score itself cannot be directly used as a probability (e.g., its sum is not 1, or it contains negative values, or the range is inappropriate), the softmax function is used for transformation. The calculation formula is as follows: ,in It is the first The sampling probability of each data feature. It is the original association strength score. It is the total number of data features. It is a temperature parameter used to adjust the smoothness of the probability distribution, for example, Based on experience, it can be set to 1.0. If the original scores are already between 0 and 1 and the sum is approximately 1 (e.g., obtained through a certain proportional allocation), then softmax can be skipped and used directly. Alternatively, a simple linear normalization method can be adopted, such as dividing each score into 1.0. Divide by the sum of all ratings ,Right now Ensure all Non-negative and summing to 1, for example, if the association strength row vector is [4, 5, 1] and its sum is 10, then the normalized sampling probabilities are 0.4, 0.5, and 0.1 respectively. These normalized numerical sequences This is the final task-related data sampling weight.
[0031] The specific steps for obtaining a subset of training data in real time are as follows: Receive training query requests, parse the drug, gene and disease entities specified in the request, query the entity index table to obtain the corresponding index, and combine the index and the relationship type identifier of the request into a slice range in tensor space to establish a tensor slice coordinate set. Based on the tensor slice coordinate set, the sampling weight of the task-related data is applied to the row vector selection of the factor matrix. By performing matrix multiplication between the tensor slice coordinate set and the weighted factor matrix, the index range of the corresponding element in the core tensor is calculated, and the target index range is obtained. Based on the target index range, the core tensor element blocks within the target index range are indexed and extracted from the compressed knowledge base component. At the same time, the factor matrix row vectors associated with the query are extracted. Tensor product operation is performed to reconstruct the core tensor element blocks and factor matrix row vectors into a dense subtensor, generating an instant training data subset.
[0032] Specifically, the system receives a structured training query request, which can be a JSON object or an XML document. This request explicitly specifies one or more drug entity names (e.g., "atorvastatin"), gene entity names (e.g., "PCSK9"), and disease entity names (e.g., "hypercholesterolemia") involved in the query, along with one or more predefined relation type identifiers (e.g., identifier 0 represents "treatment," identifier 2 represents "drug-targeted gene"). The system first parses this request, extracting these entity names and relation type identifiers. Then, for each extracted entity name, the system queries the entity index table built in the previous steps (this table stores the mapping from entity names to their unique integer indices) to obtain the unique integer index corresponding to each entity; for example, the drug index corresponding to "atorvastatin" is... "PCSK9" corresponding gene index Index of diseases corresponding to "hypercholesterolemia" If an entity for a certain dimension is not specified in the request, a wildcard or an index range representing all entities in that dimension can be used. For example, if a specific gene is not specified, the gene index range can be... arrive Then, these obtained entity indexes (or index ranges) are compared with the relation type identifier specified in the request (e.g., These can be combined to form one or more quadruples, each quadruple defining a slice range or specific coordinate point in the high-dimensional health and wellness knowledge tensor. For example, a specific query forms... Such coordinates, or This slice range description, summarizing all these combinations, establishes a tensor slice coordinate set.
[0033] Based on the tensor slice coordinate set (which defines the indexes of drug, gene, and disease entities involved in the query, as well as identifiers of relation types), and using previously acquired task-related data sampled for the current training task, the selection process of row vectors in the factor matrix is first adjusted. Specifically, for each entity index specified in the tensor slice coordinate set... (belonging to the dimension) ), and its corresponding factor matrix The OK It will be selected, and at the same time, the weight value in the task-related data sampling weight corresponding to this entity (or its feature category) will be selected. (If applicable, e.g., weights are applied to feature categories rather than individual entities) will be used to adjust the contribution or probability of selection of that row vector (if probability-based secondary sampling occurs in subsequent steps). If the sampling weights are applied to the latent feature dimensions of the factor matrix (i.e., the columns of the factor matrix), these weights will be used to weight the columns of the factor matrix, forming a weighted factor column. ,in It is the first The weights of each latent feature are then used to compute the core tensor. The system will specify the index range of the corresponding element in the tensor slice coordinate set, and will use the entity index (e.g., drug index) explicitly given in the tensor slice coordinate set. Gene Index Disease Index Relational type index The row vector of the factor matrix corresponding to ) Using these selected row vectors as input, the target index range in each dimension of the core tensor is determined by analyzing their numerical properties. For example, for each dimension You can select the row vector The median value exceeds a certain dynamic threshold (for example, the mean of the absolute values of the row vector elements plus one standard deviation; let this threshold be ). ,but If the mean is 0.5, the standard deviation is 0.2, and the threshold is 0.7, then the column indexes are collected, and these column indexes constitute the core tensor in dimension 1. target subset Its minimum and maximum index values define and This approach, for example, allows larger values in the row vectors of the factor matrix to point to more relevant regions in the core tensor, thus obtaining the target index range.
[0034] Based on the target index range, which specifies the low-dimensional core tensor The starting and ending indices of the sub-blocks to be extracted, for example, if the target index range is determined to be the first indices of the core tensor corresponding to the drug dimension. arrive The column, the gene dimension corresponds to the first core tensor. arrive Similarly, the system first deserializes and loads data from a previously established compressed knowledge base component (e.g., an .npz file or serialized object containing the core tensor and factor matrices), then precisely indexes and extracts specific element blocks of the core tensor based on this target index range. Simultaneously, the system also extracts the complete factor matrix row vectors from the compressed knowledge base component, corresponding to the drug, gene, and disease entity indexes and relation type identifiers specified in the original training query request. For example, if the query involves the drug index... Then extract the factor matrix. The OK and trim or select it with the core quantum block. The part that matches the dimension, i.e. Similar processing is performed on the row vectors of the factor matrices in other dimensions to obtain... , , Then, the Tucker product operation is performed to extract the core tensor element blocks. Reconstructing these processed factor matrix row vectors (now considered as single-row matrices) is performed as follows: The result of this calculation It is a dense subtensor containing information highly relevant to the query request, generating an immediate subset of training data.
[0035] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A training method for a large-scale model in the field of health and wellness, characterized in that, Includes the following steps: Based on UMLS entity relationship data, DrugBank drug target data and clinical symptom records, a high-dimensional health and wellness knowledge tensor is constructed by assigning a unique index to each drug, gene and disease entity and assigning dimensional coordinates to each relationship type, mapping the existence and credibility values of the relationship between entities to the target position in the tensor space composed of the index and coordinates. Based on the high-dimensional health and wellness knowledge tensor, iterative projection is used to project the tensor sequentially along each dimension of drug, gene, disease, and relationship type to generate a low-dimensional factor matrix corresponding to each dimension. The low-dimensional core tensor connecting the multi-factor matrices is calculated, and the core tensor and the multi-factor matrices are combined into a compressed knowledge base component.
2. The training method for a large-scale model in the health and wellness field according to claim 1, characterized in that, The method further includes: Construct a correlation matrix between health and wellness sub-tasks and data features. Fill the data feature correlation matrix with expert-annotated values to quantify the correlation strength of data features. When a training task is specified, extract the correlation strength row vector corresponding to the task from the data feature correlation matrix as the sampling weight of the task-related data. The system receives training query requests and converts them into tensor slice coordinates. It then applies the sampling weights of the task-related data to adjust the selection probability of the factor matrix row vectors. Through matrix multiplication, it calculates the target index range of the core tensor on the selected factor matrix row vectors. It extracts the core tensor elements and associated factor matrix row vectors within the target index range from the compressed knowledge base component, performs tensor shrinking and reconstructs the required data, and generates an instant training data subset.
3. The training method for a large-scale model in the health and wellness field according to claim 1, characterized in that, The specific steps for obtaining the high-dimensional health and wellness knowledge tensor are as follows: Based on UMLS entity relationship data, DrugBank drug target data and clinical symptom records, we integrate drug entity, gene entity and disease entity lists from the three data sources, remove duplicate entities, assign each cleaned entity an integer sequence number starting from zero and incrementing, and build an entity index table. Based on the entity index table, coordinate axes are set for four dimensions: drug, gene, disease, and relationship type. The index value of each entity in the entity index table and the preset relationship type identifier are combined into a quadruple to form a multidimensional space coordinate set. Based on the multidimensional spatial coordinate set, the credibility values of the relationships between entities in the original data are read, and the credibility values are directly written into the tensor space position specified by the corresponding quadruple in the multidimensional spatial coordinate set. For coordinate positions without records, zero values are filled in to construct a high-dimensional health and wellness knowledge tensor.
4. The training method for a large-scale model in the health and wellness field according to claim 1, characterized in that, The specific steps for obtaining the compressed knowledge base component are as follows: Based on the high-dimensional health and wellness knowledge tensor, a random factor matrix is initialized. Through iterative optimization using the alternating least squares method, the modal expansion matrix of the high-dimensional health and wellness knowledge tensor and the Khatri-Rao product of the factor matrices other than the current dimension are solved. The factor matrix is updated dimension by dimension, generating a set of four converged low-dimensional factor matrices. Based on the set of low-dimensional factor matrices, the Moore-Penrose pseudo-inverse of each matrix in the set of low-dimensional factor matrices is calculated. The high-dimensional health and wellness knowledge tensor is then subjected to tensor inner product operation with the four calculated pseudo-inverse matrices one by one to obtain the low-dimensional core tensor. Based on the low-dimensional core tensor and the set of low-dimensional factor matrices, a data structure is created. The data structure contains four matrices from the low-dimensional core tensor and the set of low-dimensional factor matrices. The data structure is then serialized and written to a disk file to establish a compressed knowledge base component.
5. The training method for a large-scale model in the health and wellness field according to claim 2, characterized in that, The specific steps for obtaining the sampling weights of the task-related data are as follows: Construct a correlation matrix between health and wellness sub-tasks and data features, where rows represent tasks and columns represent data features. Domain experts score the correlation strength of each task-feature pair to generate a task-data correlation strength matrix. Based on the task data association strength matrix, when a specified cardiovascular risk assessment training task is received, the row index of the task data association strength matrix is scanned to locate the row that completely matches the task name, and the values of all columns in that row are extracted to obtain the association strength row vector.
6. The training method for a large-scale model in the field of health and wellness according to claim 5, characterized in that, The step of obtaining the task-related data sampling weights further includes: Based on the association strength row vector, each value in the vector is directly used as the sampling probability of the corresponding data feature to generate task association data sampling weights.
7. The training method for a large-scale model in the field of health and wellness according to claim 2, characterized in that, The specific steps for obtaining the subset of instant training data are as follows: Receive training query requests, parse the drug, gene and disease entities specified in the request, query the entity index table to obtain the corresponding index, and combine the index and the relationship type identifier of the request into a slice range in tensor space to establish a tensor slice coordinate set. Based on the tensor slice coordinate set, the task-related data sampling weights are applied to the row vector selection of the factor matrix. By performing matrix multiplication between the tensor slice coordinate set and the weighted factor matrix, the index range of the corresponding elements in the core tensor is calculated, and the target index range is obtained.
8. The training method for a large-scale model in the field of health and wellness according to claim 7, characterized in that, The step of obtaining the subset of training data in real time also includes: Based on the target index range, the core tensor element blocks within the target index range are indexed and extracted from the compressed knowledge base component. At the same time, the factor matrix row vectors associated with the query are extracted. Tensor product operation is performed to reconstruct the core tensor element blocks and factor matrix row vectors into a dense subtensor, generating an instant training data subset.
9. A training device for a large-scale model in the field of health and wellness, characterized in that, The training device includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the training method for a large-scale model in the field of health and wellness as described in any one of claims 1-8.
10. A method for identifying the value system of a large-scale model in the health and wellness field, wherein the large-scale model in the health and wellness field is obtained based on the training method of the large-scale model in the health and wellness field according to any one of claims 1-8, characterized in that, The corpus texts include one or more of the following: health and wellness materials related to value understanding, basic value system theory content, questionnaires related to the value system, and interview materials related to the value system.