Association mining method, device, equipment, medium and computer program product

By obtaining the importance of reference factors and using the target processing model to calculate the association mining results, the problems of low efficiency and insufficient reliability in the existing technology are solved, and efficient and reliable association mining results are obtained.

CN114283887BActive Publication Date: 2026-03-31TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies, the association mining process is inefficient and unreliable. It requires hypothesis testing for each reference factor and places high demands on the user's domain background knowledge, making it difficult to obtain reliable association mining results.

Method used

By obtaining the importance of each reference factor in the processing task corresponding to the research objective, the association mining results are directly obtained using the objective processing model, avoiding human hypothesis testing. Machine learning models such as decision trees or deep learning models are used to train the representation features of sample objects, calculate the usage frequency and importance of reference factors, and obtain association mining results based on the importance.

Benefits of technology

It improves the efficiency and reliability of association mining results, directly obtaining association results based on the target processing model without the need for manual hypothesis testing, thus enhancing the efficiency and reliability of association mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283887B_ABST
    Figure CN114283887B_ABST
Patent Text Reader

Abstract

The application discloses a kind of association mining method, device, equipment, medium and computer program product, belong to computer technical field.The method comprises: obtaining the importance degree that each reference factor has respectively in the process of realizing the processing task corresponding to research target, and the importance degree is obtained based on target processing model, and target processing model is used to realize processing task based on each reference factor;Based on the importance degree, the association mining result between each reference factor and research target is obtained.In this way, the association mining result is directly obtained based on the importance degree that each reference factor has respectively in the process of realizing the processing task corresponding to research target, without needing to carry out hypothesis test for each reference factor respectively, and the efficiency of obtaining association mining result is higher.In addition, the importance degree is directly obtained based on target processing model, the whole association mining process does not need human participation, and the reliability of the obtained association mining result is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, medium, and computer program product for correlation mining. Background Technology

[0002] With the development of computer technology, association mining is being applied in more and more scenarios. For example, it can be used to mine the associations between various reference factors and research objectives to obtain association mining results, which can indicate which reference factors(s) are associated with the research objectives.

[0003] In related technologies, research subjects are divided into multiple groups according to research objectives. For any reference factor among various reference factors, the hypothesis testing method is used to statistically analyze whether there are significant statistical differences between the corresponding factor characteristics of subjects in different groups under any reference factor. If there are significant statistical differences, it is determined that any reference factor is associated with the research objective. After determining whether each reference factor is associated with the research objective by hypothesis testing, the association mining results are obtained.

[0004] The above method requires hypothesis testing for each reference factor to obtain association mining results, which is inefficient. In addition, the process of determining whether a reference factor is related to the research objective by hypothesis testing requires a high level of domain background knowledge from the user, making it difficult to obtain highly reliable association mining results based on this method. Summary of the Invention

[0005] This application provides a method, apparatus, device, medium, and computer program product for association mining, which can be used to improve the efficiency and reliability of obtaining association mining results. The technical solution is as follows:

[0006] On the one hand, embodiments of this application provide an association mining method, the method comprising:

[0007] The importance of each reference factor in achieving the processing task corresponding to the research objective is obtained. The importance is based on the target processing model, which is used to achieve the processing task based on the reference factors. The reference factors are factors whose relationship with the research objective is to be explored.

[0008] Based on the stated importance, the correlation mining results between each reference factor and the research objective are obtained.

[0009] On the other hand, an association mining device is provided, the device comprising:

[0010] The first acquisition unit is used to acquire the importance of each reference factor in the process of achieving the processing task corresponding to the research objective. The importance is obtained based on the target processing model, which is used to achieve the processing task based on the reference factors. The reference factors are factors that are to be explored to have a correlation with the research objective.

[0011] The second acquisition unit is used to acquire the correlation mining results between each reference factor and the research objective based on the importance level.

[0012] In one possible implementation, the first acquisition unit is used to acquire the target processing model, which is trained based on the representation features and object labels corresponding to the sample objects. The representation features corresponding to a sample object are composed of the factor features corresponding to the sample object under each of the various reference factors. Based on the target processing model, the unit acquires the importance of each of the reference factors in the process of achieving the processing task corresponding to the research objective.

[0013] In one possible implementation, the second acquisition unit is used to take reference factors that meet the reference conditions in terms of importance as correlation factors; and based on the correlation factors, to acquire the correlation mining results between each reference factor and the research objective.

[0014] In one possible implementation, the second acquisition unit is further configured to acquire a performance metric of the target processing model; and determine the confidence level of the association mining result based on the performance metric.

[0015] In one possible implementation, each reference factor is a reference gene fragment, and the factor characteristic is usage frequency. The first acquisition unit is further configured to acquire the usage frequency of the sample object under each reference gene fragment; and to construct the characterization features corresponding to the sample object based on the usage frequency of the sample object under each reference gene fragment.

[0016] In one possible implementation, the first acquisition unit is further configured to acquire a biological tissue sample of the sample object; sequence the lymphocyte receptors in the biological tissue sample to obtain gene fragment usage information; and calculate the usage frequency of the sample object under each reference gene fragment based on the gene fragment usage information.

[0017] In one possible implementation, the usage frequency of the sample object under each reference gene fragment is represented using structured data.

[0018] In one possible implementation, the processing task is a classification task based on the various reference factors; or, the processing task is a regression task based on the various reference factors.

[0019] In one possible implementation, the target processing model is a decision tree model.

[0020] In one possible implementation, the reference factors are reference gene fragments, the research target is a disease-related target, the processing task is used to predict the processing result corresponding to the disease-related target based on the reference gene fragments, and the association mining result is used to indicate the association between the reference gene fragments and the disease-related target.

[0021] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to enable the computer device to implement any of the association mining methods described above.

[0022] On the other hand, a computer-readable storage medium is also provided, wherein at least one computer program is stored therein, the at least one computer program being loaded and executed by a processor to enable a computer to implement any of the association mining methods described above.

[0023] On the other hand, a computer program product is also provided, which includes a computer program or computer instructions, which are loaded and executed by a processor to enable a computer to implement any of the association mining methods described above.

[0024] The technical solution provided in this application has at least the following beneficial effects:

[0025] The technical solution provided in this application directly obtains association mining results based on the importance of each reference factor in achieving the processing task corresponding to the research objective, without the need for hypothesis testing for each reference factor, thus achieving high efficiency in obtaining association mining results. Furthermore, the importance is directly derived from the target processing model, and the entire association mining process requires no human intervention, resulting in high reliability of the obtained association mining results. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a schematic diagram of the implementation environment of an association mining method provided in an embodiment of this application;

[0028] Figure 2 This is a flowchart of an association mining method provided in an embodiment of this application;

[0029] Figure 3 This is a schematic diagram illustrating the distribution of usage frequency provided in an embodiment of this application;

[0030] Figure 4 This is a schematic diagram illustrating the distribution of usage frequency provided in an embodiment of this application;

[0031] Figure 5 This is a schematic diagram illustrating the distribution of usage frequency provided in an embodiment of this application;

[0032] Figure 6 This is a schematic diagram illustrating the distribution of usage frequency provided in an embodiment of this application;

[0033] Figure 7 This is a schematic diagram of an association mining process provided in an embodiment of this application;

[0034] Figure 8 This is a schematic diagram illustrating a gene fragment and its degree of importance provided in an embodiment of this application;

[0035] Figure 9 This is a schematic diagram illustrating a gene fragment and its importance and performance metrics provided in an embodiment of this application;

[0036] Figure 10 This is a schematic diagram of an associated mining device provided in an embodiment of this application;

[0037] Figure 11 This is a schematic diagram of the structure of a server provided in an embodiment of this application;

[0038] Figure 12 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0040] In an exemplary embodiment, the association mining method provided in this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.

[0041] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science. AI attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0042] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and intelligent transportation.

[0043] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.

[0044] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, and intelligent transportation. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0045] Figure 1 A schematic diagram of the implementation environment of the association mining method provided in this application embodiment is shown. The implementation environment includes: terminal 11 and server 12.

[0046] The association mining method provided in this application embodiment can be executed by terminal 11, server 12, or jointly by terminal 11 and server 12; this application embodiment does not limit this. In the case where the association mining method provided in this application embodiment is jointly executed by terminal 11 and server 12, server 12 undertakes the main computational work, and terminal 11 undertakes the secondary computational work; or, server 12 undertakes the secondary computational work, and terminal 11 undertakes the main computational work; or, server 12 and terminal 11 use a distributed computing architecture for collaborative computation.

[0047] In one possible implementation, terminal 11 can be any electronic product capable of human-computer interaction with a user through one or more methods such as a keyboard, touchpad, touchscreen, remote control, voice interaction, or handwriting device. Examples include PCs (Personal Computers), mobile phones, smartphones, PDAs (Personal Digital Assistants), wearable devices, PPCs (Pocket PCs), tablets, smart car systems, smart TVs, smart speakers, smart voice interaction devices, smart home appliances, and in-vehicle terminals. Server 12 can be a single server, a server cluster consisting of multiple servers, or a cloud computing service center. Terminal 11 and server 12 establish a communication connection via wired or wireless network.

[0048] Those skilled in the art should understand that the above-described terminal 11 and server 12 are merely examples. Other existing or future terminals or servers that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0049] Based on the above Figure 1 The implementation environment shown in this application embodiment provides an association mining method, which is executed by a computer device. This computer device can be a terminal 11 or a server 12; this application embodiment does not limit the specific type of device. Figure 2 As shown, the association mining method provided in this application embodiment includes the following steps 201 and 202.

[0050] In step 201, the importance of each reference factor in the process of achieving the processing task corresponding to the research objective is obtained. The importance is obtained based on the target processing model, which is used to achieve the processing task based on each reference factor. Each reference factor is a factor that is to be explored to explore the relationship between the research objective and the target.

[0051] The association mining method provided in this application is used to mine the associations between various reference factors and research objectives to obtain association mining results between the various reference factors and research objectives. That is, each reference factor is a factor whose association with the research objective is to be mined. Each reference factor and the research objective constitute a pair of subjects whose association needs to be mined. The specific details of this pair of subjects can be flexibly adjusted according to the application scenario, and this application does not limit this.

[0052] In an exemplary embodiment, each reference factor refers to a reference gene fragment, and the research target refers to a disease-related target. Each reference gene fragment is a gene fragment for which the association with the disease-related target needs to be explored. The method of setting each reference gene fragment is related to the actual application scenario, and this embodiment does not limit this. In an exemplary embodiment, each reference gene fragment includes at least one of V (variable) gene fragments, D (diversity) gene fragments, and J (joining) gene fragments. Wherein, V gene fragments, D gene fragments, and J gene fragments refer to the corresponding gene fragments constituting the lymphocyte receptor sequence.

[0053] For example, lymphocytes include B cells and T cells. B cells and T cells are core participants in the adaptive immune system, which plays a crucial role in the body's control and clearance of viral infections and influences clinical outcomes. B cells and T cells primarily recognize antigens through receptors on their cell surface (BCRs / TCRs) and then initiate a specific immune response. Specifically, BCRs directly recognize specific antigens, while TCRs recognize specific antigens presented by MHC (Major Histocompatibility Complex) molecules.

[0054] Structurally, both BCR and TCR consist of two strands. Taking TCR as an example, TCR is mainly composed of two strands, α and β, each containing a variable domain and a constant domain. The variable domain has three hypervariable regions: CDR (Complementarity Determining Region) 1, CDR 2, and CDR 3. CDR 3 has the greatest variation and mainly determines the antigen-binding specificity of TCR. The CDR 3 region of the α strand is composed of V gene fragments and J gene fragments, while the CDR 3 region of the β strand is composed of V gene fragments, D gene fragments, and J gene fragments. For example, the gene fragments that make up the strands of BCR and TCR can be represented as V(D)J gene fragments. The diversity of the permutations and combinations of V(D)J gene fragments results in the high variability of the CDR 3 region, which contributes to the diversification of TCR and ensures that the human body can successfully recognize different antigens.

[0055] The choice of which of the following gene fragments (V, D, and J) is included in each reference gene fragment depends on the specific application scenario. For example, if the application scenario is derived from sequencing a strand of BCR or TCR composed of V and J gene fragments as the target strand, then each reference gene fragment includes at least one of V and J gene fragments; if the application scenario is derived from sequencing a strand of BCR or TCR composed of V, D, and J gene fragments as the target strand, then each reference gene fragment includes at least one of V, D, and J gene fragments.

[0056] For example, the V gene fragment, D gene fragment, and J gene fragment each have multiple specific variations. For instance, the V gene fragment constituting the β chain of the TCR has approximately 47 variations, the D gene fragment constituting the β chain of the TCR has 2 variations, and the J gene fragment constituting the β chain of the TCR has 13 variations. A reference gene fragment refers to one specific variation of the V gene fragment, D gene fragment, or J gene fragment. Taking the inclusion of the V gene fragment as an example, each reference gene fragment may include all variations of the V gene fragment, or it may include only some variations of the V gene fragment. This depends on the actual application scenario, and the embodiments of this application do not limit this.

[0057] Disease-related targets can be flexibly adjusted according to the actual application scenario, and this application embodiment does not limit them. For example, disease-related targets include, but are not limited to, whether an object suffers from a certain disease, what stage of a certain disease the object is in, whether the object has a certain medical history, and the required treatment time for the object. For example, based on the target of whether an object suffers from a certain disease, objects can be divided into two categories: sick and healthy, or into four categories: healthy, mild, severe, and recovering. Based on the target of what stage of a certain disease the object is in, objects can be divided into three categories: in the early stage of a certain disease, in the development stage of a certain disease, and in the recovery stage of a certain disease. Based on the target of whether an object has a certain medical history, objects can be divided into two categories: with a certain medical history and without a certain medical history. For example, a disease-related target can refer to a specific disease-related target, which is a disease that researchers want to further study. For example, a disease-related target can also be called a disease characteristic.

[0058] For example, when each reference factor refers to a reference gene fragment and the research target is a disease-related target, the association mining method provided in this application can assist in the analysis of the immune repertoire. The immune repertoire refers to the sum of all functionally diverse T cells and B cells in the circulatory system of an individual at a given time. Analyzing the immune repertoire and exploring the association between BCR / TCR and disease characteristics is of great significance for the treatment of cancer and autoimmune diseases, exploring tumor immune mechanisms, discovering disease treatment targets, developing antibodies, and evaluating vaccine efficacy. Therefore, mining the association between the V(D)J gene fragment and disease characteristics helps medical personnel explore the immune process of B cell / T cell recognition of disease-related antigens, helps in finding biomarkers and immunotherapy targets, conducting immunotherapy, developing vaccines, and evaluating their efficacy.

[0059] It should be noted that the above description only uses reference gene fragments as reference factors and disease-related targets as examples. The specific reference factors and research targets are not limited to this and can be flexibly adjusted according to the actual application scenario. For example, the reference factors could be user attributes, and the research target could be whether a user has a need to purchase financial products, etc.

[0060] The processing task corresponding to the research objective refers to the task defined according to the research objective for mining associations. This task is used to predict the processing results corresponding to the research objective based on various reference factors. This application does not limit the type of processing task corresponding to the research objective; it depends on the purpose of the task. For example, if the processing task corresponding to the research objective is used to predict the classification result corresponding to the research objective based on various reference factors, then the processing task is a classification task based on various reference factors; if the processing task corresponding to the research objective is used to predict the regression result corresponding to the research objective based on various reference factors, then the processing task is a regression task based on various reference factors.

[0061] In an exemplary embodiment, taking each reference factor as a reference gene fragment and the research objective as a disease-related objective as an example, the processing task is used to predict the processing result corresponding to the disease-related objective based on each reference gene fragment. For example, the processing result corresponding to the disease-related objective may be a classification result or a regression result, depending on the specific circumstances of the disease-related objective.

[0062] For example, if the disease-related objective is whether an individual has a certain disease, then based on the disease-related objective, two disease-related categories can be distinguished: diseased and healthy. In this case, the processing result corresponding to the disease-related objective can refer to the classification result indicating the probability corresponding to each of the two disease-related categories: diseased and healthy. As another example, if the disease-related objective is the required cure time for a diseased individual, then the processing result corresponding to the disease-related objective can refer to the regression result indicating the cure time.

[0063] The importance of each reference factor in achieving the processing task corresponding to the research objective reflects the magnitude of its role in achieving that task. The higher the importance of a reference factor in achieving the processing task, the greater its role and the more likely it is to be associated with the research objective. Therefore, based on the importance of each reference factor in achieving the processing task, relatively reliable association mining results can be obtained to indicate which reference factors(s) are associated with the research objective.

[0064] The relative importance of each reference factor in achieving the processing task corresponding to the research objective is derived from the target processing model used to achieve the processing task based on each reference factor. For example, if the processing task corresponding to the research objective is a classification task based on each reference factor, then the target processing model is a target classification model; if the processing task corresponding to the research objective is a regression task based on each reference factor, then the target processing model is a regression model. The target processing model is a machine learning model, thus making the association mining method provided in this application an association mining method based on machine learning.

[0065] This application does not limit the type of target processing model, as long as it can achieve the processing task corresponding to the research objective. For example, the target processing model is a decision tree model. Decision tree models have good performance in scenarios involving the processing of small to medium-sized structured or tabular data. For example, decision tree models include, but are not limited to, XGBoost (Extreme Gradient Boosting) models. In some embodiments, the target processing model can also be a deep learning-based model, such as TabNet (an interpretable model for structured or tabular data), which can effectively process large-scale structured or tabular data. Of course, in exemplary embodiments, the target processing model can also be an SVM (Support Vector Machine) model, a Random Forest model, a Deep Forest model, etc.

[0066] Before proceeding to step 201, the target processing model needs to be obtained. The target processing model is trained based on the representational features and object labels corresponding to the sample objects. Specifically, the representational features corresponding to a sample object consist of the factor features corresponding to that sample object under each reference factor.

[0067] A sample object refers to the object represented by the representational features upon which the target processing model is trained. This application does not limit the number of sample objects; exemplarily, the number of sample objects is multiple, to ensure the reliability of the trained target processing model.

[0068] Each sample object corresponds to a factor feature under each reference factor. The factor feature is used to characterize the reference factor, and the factor feature used to characterize different types of reference factors is different. For example, if the type of reference factor is a gene fragment, then the factor feature is the frequency of use of the gene fragment. For example, if the gene fragment is TRBJ2-6, the factor feature is 0.15. If the type of reference factor is a user attribute, then the factor feature is the value of the user attribute. For example, if the user attribute is age, the factor feature is 20 years old.

[0069] Different sample objects may have the same or different factor characteristics under the same reference factor, depending on the actual situation. This application does not limit this. The characterization feature corresponding to a sample object is used to characterize the sample object based on considering various reference factors. The characterization feature corresponding to a sample object is composed of the factor characteristics corresponding to that sample object under each reference factor.

[0070] This application does not limit the representation form of the characterization feature corresponding to a sample object. For example, the characterization feature can be represented as a multi-dimensional vector. In this case, the elements in the multi-dimensional vector correspond one-to-one with the factor features corresponding to the sample object under each reference factor. That is, the number of elements in the multi-dimensional vector is the same as the total number of reference factors. It should be noted that, for the case where the characterization feature is a multi-dimensional vector, elements at the same position in the characterization features corresponding to different sample objects correspond to the same reference factor.

[0071] The implementation method for obtaining the characterization features corresponding to the sample object is related to the type of each reference factor. In this embodiment, the reference factors are each reference gene fragment as an example to introduce the implementation method for obtaining the characterization features corresponding to the sample object.

[0072] When each reference factor is a reference gene fragment, the factor characteristic is its usage frequency. In this case, taking a single sample as an example, the representational features for that sample are obtained as follows: the usage frequency of the sample under each reference gene fragment is obtained; and based on the usage frequency of the sample under each reference gene fragment, the representational features for that sample are constructed. It should be noted that when there are multiple sample objects, the usage frequency of each sample under each reference gene fragment needs to be obtained separately, and then the representational features for each sample are constructed based on the usage frequency of each sample under each reference gene fragment.

[0073] In one possible implementation, the process of obtaining the usage frequency of the sample object under each reference gene fragment is as follows: obtain a biological tissue sample of the sample object; sequence the lymphocyte receptors in the biological tissue sample to obtain gene fragment usage information; and calculate the usage frequency of the sample object under each reference gene fragment based on the gene fragment usage information.

[0074] Biological tissue samples of a sample object refer to samples containing lymphocytes obtained from the sample object, such as peripheral blood. After obtaining the biological tissue sample of the sample object, the lymphocyte receptors in the biological tissue sample can be sequenced. It should be noted that the lymphocyte receptors here refer to lymphocyte receptors associated with each reference gene fragment. For example, if the reference gene fragments are gene fragments required to form one or more strands of the BCR, then the lymphocyte receptors here refer to the BCR; if the reference gene fragments are gene fragments required to form one or more strands of the TCR, then the lymphocyte receptors here refer to the TCR.

[0075] In an exemplary embodiment, sequencing of lymphocyte receptors in a biological tissue sample can refer to sequencing each strand of the lymphocyte receptor in the biological tissue sample as the target strand, or it can refer to sequencing a specific strand of the lymphocyte receptor in the biological tissue sample as the target strand. This is related to whether each reference gene fragment is a gene fragment required to form each strand or a gene fragment required to form a specific strand, and the embodiments of this application do not limit this.

[0076] For example, sequencing technology is used to sequence lymphocyte receptors in biological tissue samples. This application does not limit the type of sequencing technology used; for example, high-throughput sequencing technology can be used. High-throughput sequencing technology can obtain the sequence information of the TCR and BCR for antigen recognition in the human body. For example, this sequence information refers to gene sequences.

[0077] After sequencing lymphocyte receptors in biological tissue samples, gene fragment usage information can be obtained. This information indicates the presence of gene fragments within the lymphocyte receptor gene sequence. Based on this gene fragment usage information, the usage frequency of the sample object under each reference gene fragment can be calculated.

[0078] For example, the process of calculating the usage frequency of a sample object under any reference gene fragment based on gene fragment usage information is as follows: Based on the gene fragment usage information, the number of target gene sequences containing that reference gene fragment is counted in each of the measured gene sequences; the ratio of the number of target gene sequences to the total number of gene sequences is taken as the usage frequency of the sample object under that reference gene fragment. It should be noted that this explanation uses a single sample object as an example, meaning that the measured gene sequences refer to the gene sequences obtained by sequencing lymphocyte receptors in an individual's immune repertoire. This method can obtain the usage frequency of the sample object under each reference gene fragment. It should be noted that if the gene fragment usage information indicates that a certain reference gene fragment is not present in any of the measured gene sequences, then the usage frequency of the sample object under that reference gene fragment is 0.

[0079] In an exemplary embodiment, the usage frequency of the sample object under each reference gene fragment can be stored in a database (e.g., a relational database). In this case, the usage frequency of the sample object under each reference gene fragment can be directly extracted from the database.

[0080] In an exemplary embodiment, the usage frequency of a sample object under each reference gene fragment is represented using structured data. For example, the usage frequency of each sample object under each reference gene fragment is represented by a row of data in a usage frequency statistics table. Each column of data in this usage frequency statistics table represents the usage frequency of each sample object under a specific reference gene fragment. Using structured data to represent the usage frequency of sample objects under each reference gene fragment helps improve the efficiency of association mining.

[0081] For example, each reference gene fragment can be denoted as a V(D)J gene fragment. The process of obtaining the usage frequency of each reference gene fragment for a sample object can be called the immune repertoire V(D)J gene fragment usage frequency construction process. In practice, researchers first obtain biological tissue samples (e.g., peripheral blood) from the sample object and perform high-throughput sequencing on all TCRs (or BCRs) to obtain the sequences of all TCRs (or BCRs) and V(D)J gene fragment usage information to form a T cell (or B cell) receptor immune repertoire. Based on the above data, the usage frequency of the V(D)J gene fragment is statistically analyzed as the main research subject to obtain features (i.e., the representational features corresponding to the sample object) for subsequent input processing models.

[0082] For example, different categories of individuals may have different usage frequencies for the same gene segment. For example, the distribution of usage frequencies for individuals with disease A under each of the V gene segments (Vα) constituting the α chain of the TCR, and the distribution of usage frequencies for healthy individuals under each of the V gene segments (Vα) are as follows: Figure 3 As shown; the distribution of usage frequency under each J gene segment (Jα) constituting the α chain of the TCR in subjects with disease A, and the distribution of usage frequency under each J gene segment (Jα) in healthy subjects are as follows. Figure 4 As shown; the distribution of usage frequency under each V gene segment (Vβ) constituting the β chain of the TCR in subjects with disease A, and the distribution of usage frequency under each V gene segment (Vβ) in healthy subjects are as follows. Figure 5 As shown; the distribution of usage frequency under each J gene segment (Jβ) constituting the β chain of the TCR in subjects with disease A, and the distribution of usage frequency under each J gene segment (Jβ) in healthy subjects are as follows. Figure 6 As shown.

[0083] The object labels of sample objects are used to provide supervision signals for the training process of the processing model. The type of object label of a sample object is related to the type of processing task corresponding to the research objective. For example, if the type of processing task corresponding to the research objective is a classification task based on various reference factors, then the object label of the sample object is a category label; if the type of processing task corresponding to the research objective is a regression task based on various reference factors, then the label of the sample object is a regression value label. For example, the object labels of sample objects are determined by researchers according to the research objective in order to provide a more reliable supervision signal for the training process of the processing model used to achieve the processing task corresponding to the research objective. For example, the object labels of the same sample object determined according to different research objectives may be the same or different, and this embodiment of the application does not limit this.

[0084] The process of obtaining the target processing model can refer to directly extracting and storing the target processing model that has been trained in advance based on the representation features and object labels corresponding to the sample objects, or it can refer to training the target processing model in real time based on the representation features and object labels corresponding to the sample objects. This application does not limit this.

[0085] In an exemplary embodiment, the process of training the target processing model based on the representational features and object labels corresponding to the sample objects is as follows: inputting the representational features corresponding to the sample objects into the initial processing model to obtain the processing result output by the initial processing model; obtaining the loss function based on the difference between the processing result and the object labels; updating the parameters of the initial classification model using the loss function; and obtaining the target processing model in response to the training process meeting the termination condition. The process of training the target processing model is a supervised training process, which will not be elaborated here.

[0086] In this embodiment of the application, to achieve association mining, a surrogate problem is predefined, that is, the representation features composed of the factor features corresponding to the sample object under each reference factor are used as the input features of the model, and the model predicts the processing result. A machine learning model (i.e., the target processing model) is trained under this surrogate problem.

[0087] After obtaining the target processing model, the importance of each reference factor in achieving the processing task corresponding to the research objective is obtained based on the target processing model. For example, since the target processing model is used to achieve the processing task corresponding to the research objective based on each reference factor, the importance of each reference factor in achieving the processing task corresponding to the research objective can be obtained by interpreting the target processing model. For example, the representation of importance in this application does not limit the method of representation; for example, importance can be represented by a decimal in the range of 0 to 1, or importance can be represented by an importance score.

[0088] In the exemplary embodiments, each processing model has a corresponding model interpretation method to analyze the importance of each reference factor in achieving the processing task. That is, the way to obtain the importance of each reference factor in achieving the processing task corresponding to the research objective based on the target processing model is related to the type of the target processing model, and this application embodiment does not limit this.

[0089] In an exemplary embodiment, the target processing model itself has the attribute of the importance of each feature (i.e., each reference factor) in the calculation input during the process of realizing the processing task. After obtaining the target processing model, the importance of each reference factor in the process of realizing the processing task can be obtained directly by using the attribute of the target processing model itself.

[0090] In an exemplary embodiment, machine learning interpretation methods such as SHAP (Shapley Additive Explanations) are used to interpret the target processing model in order to obtain the importance of each reference factor in the process of achieving the processing task.

[0091] In an exemplary embodiment, the factor characteristics under each reference factor are randomly perturbed, and the changes in the indicators of the target processing model are observed. The importance of each reference factor is determined according to the magnitude of the rate of change; the greater the rate of change, the higher the importance.

[0092] Regardless of the method used, it is possible to obtain the degree of importance of each reference factor in the process of achieving the research objective and corresponding processing task, and then proceed to step 202.

[0093] In step 202, the correlation mining results between each reference factor and the research objective are obtained based on the degree of importance.

[0094] After determining the importance of each reference factor in achieving the processing task corresponding to the research objective, the association between each reference factor and the research objective can be analyzed based on their respective importance in achieving the processing task, thus obtaining association mining results. These association mining results indicate the association between each reference factor and the research objective, that is, which or more reference factors are associated with the research objective. For example, in the case where the reference factors are reference gene fragments and the research objective is a disease-related objective, the association mining results indicate the association between each reference gene fragment and the disease-related objective, that is, which or more reference gene fragments are associated with the disease-related objective.

[0095] In one possible implementation, the process of obtaining the association mining results between each reference factor and the research objective based on the importance of each reference factor in achieving the processing task corresponding to the research objective is as follows: Reference factors whose importance meets the reference conditions are taken as association factors; based on the association factors, the association mining results between each reference factor and the research objective are obtained. Reference factors whose importance meets the reference conditions are considered to have higher importance. The criteria for meeting the reference conditions are set based on experience or flexibly adjusted according to the actual application scenario; this application embodiment does not limit this.

[0096] In an exemplary embodiment, the importance level satisfying the reference condition means that the importance level is not less than the importance threshold. The importance threshold is set based on experience or can be flexibly adjusted according to the application scenario; this embodiment does not limit this. In an exemplary embodiment, the importance level satisfying the reference condition means that the importance level is among the top K (K is an integer not less than 1) most important. In this case, the number of identified related factors is K.

[0097] After identifying the relevant factors, the association mining results between each reference factor and the research objective are obtained based on these factors. The methods for obtaining these association mining results vary depending on the representation of the association mining results.

[0098] For example, if the association mining results are directly represented using association factors, then the association factors are directly used as the association mining results. For example, if the association mining results are represented using association factors and the importance of those factors, then the association factors and their importance are used as the association mining results. For example, if the association mining results are represented using association factors ranked by importance, then the association factors arranged in descending or ascending order of importance are used as the association mining results.

[0099] In an exemplary embodiment, in addition to obtaining the association mining results, the embodiments of this application can also obtain the confidence level of the association mining results, so as to use the confidence level to characterize the reliability of the obtained association mining results. The method for obtaining the confidence level of the association mining results is as follows: obtain the performance measurement index of the target processing model; and determine the confidence level of the association mining results based on the performance measurement index.

[0100] The performance metrics of the target processing model are used to measure its processing performance. This application does not limit the type of performance metrics, but includes, but is not limited to, at least one of F1 score, recall, precision, accuracy, and AUC (Area Under ROC Curve). ROC (Receiver Operating Characteristic) refers to the receiver operating characteristic curve. In an exemplary embodiment, the performance metrics of the target processing model are obtained based on test results, which are obtained by testing the target processing model using a test set. The process of obtaining the performance metrics of the target processing model based on the test results is related to the type of performance metrics, and this application does not limit this process. The test set includes the representational features and object labels corresponding to the test objects.

[0101] After obtaining the performance metrics of the target processing model, the confidence level of the association mining results is determined based on these metrics. In an exemplary embodiment, if there is only one performance metric, that single performance metric is directly used as the confidence level of the association mining results. If there are multiple performance metrics, each performance metric can be used as a confidence level for the association mining results, thus providing multifaceted confidence levels; alternatively, the various performance metrics can be weighted and fused to obtain a fused metric, which is then used as the confidence level of the association mining results.

[0102] The core idea of ​​this application is to use machine learning to discover the association between various reference factors and the research objective by solving a set proxy problem. If a reference factor (e.g., a gene fragment in an immune repertoire) plays a dominant role in the model's achievement of the processing task corresponding to the research objective, then there is a specific correlation between this reference factor and the research objective, that is, it means that the reference factor is associated with the research objective.

[0103] For example, taking the V(D)J gene fragment as an example, the association mining process in this embodiment of the application is as follows: Figure 7 As shown, this association mining process consists of two parts: feature construction and association mining, specifically divided into three steps: ① obtaining the usage frequency of V(D)J gene fragments in the immune repertoire to provide data support for obtaining the representational features of the input model; ② training the target processing model; ③ association mining. In the feature construction part, TCRs in the immune repertoire are sequenced to obtain TCR gene fragment usage information, and then the usage frequency of objects under each V(D)J gene fragment is determined based on the gene fragment usage information. In the association mining part, representational features of multiple sample objects are first constructed based on the usage frequency of sample objects under each V(D)J gene fragment, and then the target processing model is trained based on the representational features and object labels; the importance of V(D)J gene fragments is obtained based on the target processing model, and the association mining results are obtained based on the importance. In addition, the target processing model can be tested, and the performance evaluation index of the target processing model is obtained based on the prediction results output by the target processing model, and then the confidence level of the association mining results is obtained based on the performance evaluation index of the target processing model.

[0104] For example, each reference factor is taken as a reference gene fragment. In different combinations of reference gene fragments and research objectives, the top 10 gene fragments with the highest importance determined by the method provided in this application, and their respective importance levels, will vary.

[0105] In an exemplary embodiment, each reference gene fragment is a V gene fragment required to constitute the α chain of the TCR, and the research objective is to determine whether the subject suffers from disease A. In this case, based on the embodiments of this application, the top 10 gene fragments with the highest importance and their importance are as follows: Figure 8 As shown in (1) of the table.

[0106] In an exemplary embodiment, each reference gene fragment is a J gene fragment required to constitute the α chain of the TCR, and the research objective is to determine whether the subject suffers from disease A. In this case, based on the embodiments of this application, the top 10 gene fragments with the highest importance and their importance are as follows: Figure 8 As shown in (2) of the text.

[0107] In an exemplary embodiment, each reference gene fragment is a V gene fragment required to constitute the β chain of the TCR, and the research objective is to determine whether the subject suffers from disease A. In this case, based on the embodiments of this application, the top 10 gene fragments with the highest importance and their importance are as follows: Figure 8 As shown in (3) of the text.

[0108] In an exemplary embodiment, each reference gene fragment is a J gene fragment required to constitute the β chain of the TCR, and the research objective is to determine whether the subject suffers from disease A. In this case, based on the embodiments of this application, the top 10 gene fragments with the highest importance and their importance are as follows: Figure 8 As shown in (4) of the text.

[0109] In an exemplary embodiment, each reference gene fragment is a V, D, or J gene fragment required to constitute the β chain of the TCR, and the research target is the disease stage (early stage, development stage, recovery stage) of the subject. In this case, based on the embodiments of this application, the top 10 gene fragments with the highest importance and their importance are as follows: Figure 9 As shown in (1) of this application, the performance metrics of the target processing model obtained based on the embodiments of this application are as follows: Figure 9 As shown in (5) of the text.

[0110] In an exemplary embodiment, each reference gene fragment is a V, D, or J gene fragment required to constitute the β chain of the TCR. The research objective is to determine whether the subject has a history of cancer. In this case, based on the embodiments of this application, the top 10 gene fragments with the highest importance and their importance are as follows: Figure 9 As shown in (2) of this application, the performance metrics of the target processing model obtained based on the embodiments of this application are as follows: Figure 9 As shown in (6) of the document.

[0111] In an exemplary embodiment, each reference gene fragment is a V, D, or J gene fragment required to constitute the β chain of the TCR. The research objective is to determine whether the subject has a history of diabetes. In this case, based on the embodiments of this application, the top 10 gene fragments with the highest importance and their importance are as follows: Figure 9 As shown in (3) of this application, the performance metrics of the target processing model obtained based on the embodiments of this application are as follows: Figure 9 As shown in (7) of the document.

[0112] In an exemplary embodiment, each reference gene fragment is a V, D, or J gene fragment required to constitute the β chain of the TCR. The research objective is to determine whether the subject has a history of chronic hypertension. In this case, based on the embodiments of this application, the top 10 gene fragments with the highest importance and their importance are as follows: Figure 9 As shown in (4) of this application, the performance metrics of the target processing model obtained based on the embodiments of this application are as follows: Figure 9 As shown in (8) of the table.

[0113] In an exemplary embodiment, taking each reference factor as a reference gene fragment and the research target as a disease-related target as an example, the performance metrics of different types of target processing models (XGBoost model, TabNet model, SVM model) were tested under multiple disease-related targets (the disease stage of the object, whether the object has a history of cancer, whether the object has a history of chronic hypertension, and whether the object has a history of diabetes). The test results are shown in Table 1.

[0114]

[0115] Table 1 shows that, for four disease-related targets—disease stage, cancer history, chronic hypertension history, and diabetes history—the XGBoost model outperforms the TabNet model. This indicates that the XGBoost model is superior to the deep learning-based TabNet model in handling small to medium-sized structured or tabular data. A similar conclusion can be drawn by comparing the performance of the XGBoost model and the SVM model; that is, the XGBoost model outperforms the SVM model for the same four disease-related targets. The XGBoost model achieves better performance for all four disease-related targets—disease stage, cancer history, chronic hypertension history, and diabetes history—demonstrating its advantage in handling structured or tabular data.

[0116] The method provided in this application can realize multi-factor analysis. Several reference factors with high importance (i.e., the most important reference factors) are associated with the research objective. If multiple reference factors have high importance at the same time, they can be considered to be jointly associated with the research objective, thereby realizing multi-factor analysis.

[0117] Based on the embodiments of this application, a machine learning-based immune repertoire analysis method can be implemented to mine the association between the T / B cell receptor V(D)J gene fragment in the immune repertoire and disease characteristics, and provide the confidence level of the revealed V(D)J gene fragment association with the disease. Machine learning (deep learning) enables end-to-end association mining, has no strict preconditions regarding the statistical distribution of data, requires minimal domain background knowledge from users, and requires minimal data preprocessing. Machine learning (deep learning) can use multi-factor, multi-dimensional feature vectors (where the number of elements in the feature vector equals the number of factors) as model inputs, modeling multi-group analysis as a multi-classification problem, thus exhibiting significant advantages in multi-factor, multi-group analysis. While mining the association between the V(D)J gene fragment and disease characteristics, the machine learning-based method can use the model's accuracy to indicate the confidence level of the mined associations.

[0118] The methods provided in this application can offer new analytical tools for immune repertoire analysis, reducing the need for domain background knowledge and data preparation, saving time and learning costs for analysts, and enhancing the ability of immune repertoire analysis to discover new knowledge. These methods can advance the association analysis between V(D)J gene fragments and disease characteristics from a primarily statistical, single-factor approach to a multi-factor simultaneous analysis. Furthermore, the methods provided in this application, on a long-term timescale, are beneficial for the discovery of biomarkers and immunotherapy targets based on immune repertoire analysis, the implementation of immunotherapy, vaccine development, and efficacy evaluation. The association mining methods provided in this application can serve as a feature screening tool to identify associated features for building machine learning models, such as addressing questions like which tumor markers and patient clinical information are related to lymph node metastasis. The end-to-end association mining method based on machine learning provided in this application can be widely used in the processing of relevant data, extracting valuable information from the data to provide a reference for subsequent experimental verification, and significantly reducing the workload of relevant personnel in data processing.

[0119] The association mining method provided in this application directly obtains association mining results based on the importance of each reference factor in achieving the processing task corresponding to the research objective, without the need for hypothesis testing for each reference factor, thus achieving high efficiency in obtaining association mining results. Furthermore, the importance is directly derived from the target processing model, and the entire association mining process requires no human intervention, resulting in high reliability of the obtained association mining results.

[0120] See Figure 10 This application provides an association mining device, which includes:

[0121] The first acquisition unit 1001 is used to acquire the importance of each reference factor in the process of achieving the processing task corresponding to the research objective. The importance is obtained based on the target processing model. The target processing model is used to achieve the processing task based on each reference factor. Each reference factor is a factor that is to be explored and has a relationship with the research objective.

[0122] The second acquisition unit 1002 is used to acquire the correlation mining results between each reference factor and the research objective based on the importance level.

[0123] In one possible implementation, the first acquisition unit 1001 is used to acquire the target processing model, which is trained based on the representation features and object labels corresponding to the sample objects. The representation features corresponding to a sample object are composed of the factor features corresponding to the sample object under each reference factor. Based on the target processing model, the importance of each reference factor in the process of achieving the processing task corresponding to the research objective is acquired.

[0124] In one possible implementation, the second acquisition unit 1002 is used to take reference factors that meet the reference conditions in terms of importance as correlation factors; and based on the correlation factors, to acquire the correlation mining results between each reference factor and the research objective.

[0125] In one possible implementation, the second acquisition unit 1002 is further used to acquire performance metrics of the target processing model; and based on the performance metrics, to determine the confidence level of the association mining results.

[0126] In one possible implementation, each reference factor is a reference gene fragment, and the factor feature is the usage frequency. The first acquisition unit 1001 is also used to acquire the usage frequency of the sample object under each reference gene fragment. Based on the usage frequency of the sample object under each reference gene fragment, the characteristic features corresponding to the sample object are constructed.

[0127] In one possible implementation, the first acquisition unit 1001 is further configured to acquire a biological tissue sample of the sample object; sequence the lymphocyte receptors in the biological tissue sample to obtain gene fragment usage information; and calculate the usage frequency of the sample object under each reference gene fragment based on the gene fragment usage information.

[0128] In one possible implementation, the usage frequency of the sample object under each reference gene fragment is represented using structured data.

[0129] In one possible implementation, the processing task is a classification task based on various reference factors; or, the processing task is a regression task based on various reference factors.

[0130] In one possible implementation, the target processing model is a decision tree model.

[0131] In one possible implementation, each reference factor is a reference gene fragment, the research target is a disease-related target, the processing task is used to predict the processing results corresponding to the disease-related target based on each reference gene fragment, and the association mining results are used to indicate the association between each reference gene fragment and the disease-related target.

[0132] The association mining device provided in this application directly obtains association mining results based on the importance of each reference factor in achieving the processing task corresponding to the research objective, without the need for hypothesis testing for each reference factor, thus achieving high efficiency in obtaining association mining results. Furthermore, the importance is directly derived from the target processing model, and the entire association mining process requires no human intervention, resulting in high reliability of the obtained association mining results.

[0133] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional units. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the device can be divided into different functional units to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0134] Figure 11This is a schematic diagram of a server structure provided in an embodiment of this application. The server can vary significantly due to differences in configuration or performance. It may include one or more Central Processing Units (CPUs) 1101 and one or more memories 1102. The one or more memories 1102 store at least one computer program, which is loaded and executed by the one or more processors 1101 to enable the server to implement the association mining methods provided in the above-described method embodiments. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated upon here.

[0135] Figure 12 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application. The terminal can be: a PC, mobile phone, smartphone, PDA, wearable device, PPC (Pocket PC), tablet computer, smart car system, smart TV, smart speaker, smart voice interaction device, smart home appliance, in-vehicle terminal, etc. The terminal may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names.

[0136] Typically, a terminal includes a processor 1501 and a memory 1502.

[0137] Processor 1501 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1501 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1501 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1501 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1501 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0138] The memory 1502 may include one or more computer-readable storage media, which may be non-transitory. The memory 1502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1502 are used to store at least one instruction, which is executed by the processor 1501 to cause the terminal to implement the association mining method provided in the method embodiments of this application.

[0139] In some embodiments, the terminal may also optionally include: a peripheral device interface 1503 and at least one peripheral device. The processor 1501, memory 1502, and peripheral device interface 1503 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1503 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1504, a display screen 1505, a camera assembly 1506, an audio circuit 1507, and a power supply 1509.

[0140] Peripheral device interface 1503 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1501 and memory 1502. In some embodiments, processor 1501, memory 1502 and peripheral device interface 1503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1501, memory 1502 and peripheral device interface 1503 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0141] The radio frequency (RF) circuit 1504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1504 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1504 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1504 can communicate with other terminals via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1504 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0142] Display screen 1505 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1505 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1501 for processing. In this case, display screen 1505 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, display screen 1505 can be a single screen, disposed on the front panel of the terminal; in other embodiments, display screen 1505 can be at least two screens, disposed on different surfaces of the terminal or in a folded design; in other embodiments, display screen 1505 can be a flexible display screen, disposed on a curved or folded surface of the terminal. Furthermore, display screen 1505 can be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 1505 can be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0143] The camera assembly 1506 is used to acquire images or videos. Optionally, the camera assembly 1506 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1506 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0144] The audio circuit 1507 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1501 for processing, or input to the radio frequency circuit 1504 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1501 or the radio frequency circuit 1504 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1507 may also include a headphone jack.

[0145] Power supply 1509 is used to power the various components in the terminal. Power supply 1509 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1509 includes a rechargeable battery, the rechargeable battery can support wired or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0146] In some embodiments, the terminal further includes one or more sensors 1510. The one or more sensors 1510 include, but are not limited to: an acceleration sensor 1511, a gyroscope sensor 1512, a pressure sensor 1513, an optical sensor 1515, and a proximity sensor 1516.

[0147] Accelerometer 1511 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by the terminal. For example, accelerometer 1511 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1501 can control display screen 1505 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1511. Accelerometer 1511 can also be used for games or for acquiring user motion data.

[0148] The gyroscope sensor 1512 can detect the terminal's orientation and rotation angle. The gyroscope sensor 1512, in conjunction with the accelerometer sensor 1511, can collect the user's 3D movements on the terminal. Based on the data collected by the gyroscope sensor 1512, the processor 1501 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0149] The pressure sensor 1513 can be disposed on the side bezel of the terminal and / or on the lower layer of the display screen 1505. When the pressure sensor 1513 is disposed on the side bezel of the terminal, it can detect the user's grip signal on the terminal, and the processor 1501 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1513. When the pressure sensor 1513 is disposed on the lower layer of the display screen 1505, the processor 1501 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1505. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0150] An optical sensor 1515 is used to collect ambient light intensity. In one embodiment, the processor 1501 can control the display brightness of the display screen 1505 based on the ambient light intensity collected by the optical sensor 1515. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1505 is increased; when the ambient light intensity is low, the display brightness of the display screen 1505 is decreased. In another embodiment, the processor 1501 can also dynamically adjust the shooting parameters of the camera assembly 1506 based on the ambient light intensity collected by the optical sensor 1515.

[0151] The proximity sensor 1516, also known as a distance sensor, is typically installed on the front panel of the terminal. The proximity sensor 1516 is used to detect the distance between the user and the front of the terminal. In one embodiment, when the proximity sensor 1516 detects that the distance between the user and the front of the terminal is gradually decreasing, the processor 1501 controls the display screen 1505 to switch from a screen-on state to a screen-off state; when the proximity sensor 1516 detects that the distance between the user and the front of the terminal is gradually increasing, the processor 1501 controls the display screen 1505 to switch from a screen-off state to a screen-on state.

[0152] Those skilled in the art will understand that Figure 12 The structure shown does not constitute a limitation on the terminal and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0153] In an exemplary embodiment, a computer device is also provided, comprising a processor and a memory, wherein at least one computer program is stored in the memory. The at least one computer program is loaded and executed by one or more processors to enable the computer device to implement any of the aforementioned association mining methods. The computer device may refer to a terminal or a server; this application embodiment does not limit the definition.

[0154] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one computer program that is loaded and executed by a processor of a computer device to enable the computer to implement any of the above-described association mining methods.

[0155] In one possible implementation, the aforementioned computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0156] In an exemplary embodiment, a computer program product is also provided, which includes a computer program or computer instructions that are loaded and executed by a processor to enable a computer to implement any of the association mining methods described above.

[0157] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the above exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0158] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0159] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An association mining method characterized by comprising: The method comprises: obtaining importance degrees of respective reference gene fragments in implementing a processing task corresponding to a disease-related target, the importance degrees being obtained based on a target processing model, the target processing model being used to implement the processing task based on the respective reference gene fragments, the respective reference gene fragments being factors to be mined for association between the disease-related target and the respective reference gene fragments, the processing task being used to predict a processing result corresponding to the disease-related target according to the respective reference gene fragments, the respective reference gene fragments being gene fragments in lymphocyte receptor sequences; based on the importance degrees, obtaining association mining results between the respective reference gene fragments and the disease-related target, the association mining results being used to indicate the association between the respective reference gene fragments and the disease-related target.

2. The method of claim 1, wherein, The obtaining of the importance degrees of the respective reference gene fragments in implementing the processing task corresponding to the disease-related target comprises: obtaining the target processing model, the target processing model being obtained based on a representation feature corresponding to a sample object and an object label, the representation feature corresponding to the sample object being composed of respective factor features of the sample object under the respective reference gene fragments; based on the target processing model, obtaining the importance degrees of the respective reference gene fragments in implementing the processing task corresponding to the disease-related target.

3. The method of claim 1, wherein, The obtaining of the association mining results between the respective reference gene fragments and the disease-related target based on the importance degrees comprises: taking a reference gene fragment with an importance degree satisfying a reference condition as an association factor; based on the association factor, obtaining the association mining results between the respective reference gene fragments and the disease-related target.

4. The method according to any of claims 1 to 3, characterized in that, The method further comprises: obtaining a performance measurement index of the target processing model; based on the performance measurement index, determining a confidence degree of the association mining results.

5. The method of claim 2, wherein, The factor feature is a usage frequency, and the method further comprises: obtaining respective usage frequencies of the sample object under the respective reference gene fragments; based on the respective usage frequencies of the sample object under the respective reference gene fragments, composing the representation feature corresponding to the sample object.

6. The method of claim 5, wherein, The obtaining of the respective usage frequencies of the sample object under the respective reference gene fragments comprises: obtaining a biological tissue sample of the sample object; sequencing lymphocyte receptors in the biological tissue sample to obtain gene fragment usage information, and calculating the respective usage frequencies of the sample object under the respective reference gene fragments based on the gene fragment usage information.

7. The method according to claim 5 or 6, characterized in that, The respective usage frequencies of the sample object under the respective reference gene fragments are represented by structured data.

8. The method of any one of claims 1-3, wherein, The processing task is a classification task according to the respective reference gene fragments; or the processing task is a regression task according to the respective reference gene fragments.

9. The method according to any one of claims 1 to 3, characterized in that, The target processing model is a decision tree model.

10. An association mining device characterized by comprising: The device comprises: The first obtaining unit is configured to obtain importance degrees of respective reference gene segments in a process of implementing a processing task corresponding to a disease-related target, the importance degrees being obtained based on a target processing model, the target processing model being used to implement the processing task based on the respective reference gene segments, the respective reference gene segments being factors to be mined for association between the disease-related target and the respective reference gene segments, the processing task being used to predict a processing result corresponding to the disease-related target according to the respective reference gene segments, and the respective reference gene segments being gene segments in a lymphocyte receptor sequence. The second obtaining unit is configured to obtain association mining results between the respective reference gene segments and the disease-related target based on the importance degrees, the association mining results being used to indicate the association between the respective reference gene segments and the disease-related target.

11. The apparatus of claim 10, wherein, The first obtaining unit is configured to obtain the target processing model, the target processing model being obtained based on a representation feature of a sample object and an object label, the representation feature of the sample object being composed of respective factor features of the sample object corresponding to the respective reference gene segments, and the importance degrees of the respective reference gene segments in the process of implementing the processing task corresponding to the disease-related target being obtained based on the target processing model.

12. The apparatus of claim 10, wherein, The second obtaining unit is configured to take a reference gene segment with an importance degree satisfying a reference condition as an association factor, and obtain the association mining results between the respective reference gene segments and the disease-related target based on the association factor.

13. The apparatus of any of claims 10-12, wherein, The second obtaining unit is further configured to obtain a performance measurement index of the target processing model, and determine a confidence degree of the association mining results based on the performance measurement index.

14. The apparatus of claim 11, wherein, The factor feature is a usage frequency, and the first obtaining unit is further configured to obtain respective usage frequencies of the sample object corresponding to the respective reference gene segments, and compose the representation feature of the sample object based on the respective usage frequencies of the sample object corresponding to the respective reference gene segments.

15. The apparatus of claim 14, wherein, The first obtaining unit is configured to obtain a biological tissue sample of the sample object, sequence lymphocyte receptors in the biological tissue sample to obtain gene segment usage information, and calculate the respective usage frequencies of the sample object corresponding to the respective reference gene segments based on the gene segment usage information.

16. The apparatus of claim 14 or 15, wherein, The respective usage frequencies of the sample object corresponding to the respective reference gene segments are represented by structured data.

17. The apparatus of any of claims 10-12, wherein, The processing task is a classification task based on the respective reference gene segments, or the processing task is a regression task based on the respective reference gene segments.

18. The apparatus of any of claims 10-12, wherein, The target processing model is a decision tree model.

19. A computer device, comprising: The computer device includes a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor, so that the computer device implements the association mining method according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that, The computer readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor to enable the computer to implement the association mining method according to any one of claims 1 to 9.

21. A computer program product, characterised in that, The computer program product comprises computer programs or computer instructions, and the computer programs or the computer instructions are loaded and executed by the processor to enable the computer to implement the association mining method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method for rapidly discovering phenotype related gene based on probabilistic framework and resequencing technology

    CN105404793A

  • Marker and method for quantitative detection of lymphocytes

    CN107955831A