Artificial intelligence feature extraction method and system based on lymphoid tumor micm phenotyping

By using the quantum-encoded Transformer algorithm and the Monte Carlo-optimized Conditional Random Field algorithm, the problems of insufficient data processing efficiency and accuracy in lymphoma diagnosis are solved, enabling more efficient data feature capture and intelligent application of diagnostic models.

CN120015345BActive Publication Date: 2026-02-06CHONGQING HUAXIN YINGFEI INTELLIGENT TECHNOLOGY RESEARCH INSTITUTE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510094561.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2026-02-06
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

Existing medical big data systems lack effective data processing and feature extraction methods in the diagnosis of lymphoma, resulting in insufficient accuracy and efficiency of diagnostic models, making it difficult to achieve standardized and intelligent clinical diagnosis and treatment.

Method used

We employ the Transformer algorithm based on quantum coding for feature extraction, and combine it with a dynamic quantum topology optimization strategy and a conditional random field algorithm optimized by Monte Carlo method to optimize the connection topology between qubits, thereby enhancing the model's data feature capture capability and processing efficiency.

Benefits of technology

It improves the data processing capabilities for lymphoma diagnosis, enhances the model's adaptability to different data characteristics, and improves the classification accuracy and prediction reliability in diagnostic tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015345B_ABST
    Figure CN120015345B_ABST
Patent Text Reader

Abstract

The application belongs to the field of intelligent medical treatment, and particularly relates to an artificial intelligence feature extraction method and system based on lymphoma MICM phenotyping. The method comprises the following steps: S1, obtaining a diagnosis report of a to-be-tested person; S2, obtaining a structured diagnosis report based on the diagnosis report; S3, vectorizing the structured diagnosis report to obtain a vectorized diagnosis report; S4, inputting the vectorized diagnosis report into a quantum coding-based Transformer for feature extraction to obtain MICM phenotyping features; the quantum coding-based Transformer converts input data into a quantum state and extracts quantum state feature information, and converts the quantum state feature information into non-quantum state data and outputs the MICM phenotyping features. The application extracts the required MICM features by using the quantum coding-based Transformer, and enhances the data feature capturing capacity of the model by using the quantum coding-based Transformer algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent medical treatment, and more particularly, to an artificial intelligence feature extraction method, system, device, medium and program product based on lymphoma MICM phenotype typing. BACKGROUND

[0002] Structured report specifications are imperative in the era of medical big data. Not only can they establish a good data foundation in the clinical diagnosis and treatment of hematological tumors and scientific research, but they can also play a better bridging role in communication with the clinic and patients, promoting hematological tumor treatment centers to gradually move towards standardization, dataization and intelligentization, and carrying out in-depth clinical research and innovative treatment methods in the field of hematological tumor treatment. SUMMARY

[0003] In view of the above problems, the present application provides an artificial intelligence feature extraction method based on lymphoma MICM phenotype typing, which captures information in time series data using data processing and feature extraction, improves the parameters of the disease prediction model, and effectively optimizes it, thereby constructing a disease early warning model suitable for clinical use.

[0004] The present application (first aspect) discloses an artificial intelligence feature extraction method based on lymphoma MICM phenotype typing, comprising:

[0005] S101: obtaining a diagnosis report of a subject to be tested;

[0006] S102: obtaining a structured diagnosis report based on the diagnosis report;

[0007] S103: vectorizing the structured diagnosis report to obtain a vectorized diagnosis report;

[0008] S104: inputting the vectorized diagnosis report into a quantum coding-based Transformer for feature extraction to obtain MICM phenotype typing features; the quantum coding-based Transformer converts input data into quantum states and extracts quantum state feature information, and converts the quantum state feature information into non-quantum state data to output MICM phenotype typing features.

[0009] Further, the quantum coding-based Transformer includes an input layer, a quantum coding layer, a quantum gate layer and an output layer; the processing steps of the quantum coding-based Transformer include:

[0010] Step 1: the input layer configures an initial quantum state for each vector of the vectorized diagnosis report to obtain a vector of initial quantum states;

[0011] Step 2: the quantum encoding layer performs quantum encoding on the vector of the initial quantum state to obtain a quantum-encoded vector;

[0012] Step 3: the quantum-encoded vector is forward-propagated through the quantum gate layer to obtain a forward-propagated extracted quantum state feature;

[0013] Step 4: the output layer maps the extracted quantum state feature back to a non-quantum state to obtain the MICM phenotype typing feature;

[0014] Optionally, the steps further comprise a Step 3' between Step 2 and Step 3: optimizing the correlation distance between quantum bits of the quantum-encoded vector through a dynamic quantum topology optimization strategy.

[0015] Further, the manner in which each vector configures the initial quantum state is represented as:

[0016]

[0017] In the formula, represents the initial quantum state, and i is a complex amplitude, |i> is the ground state of a quantum bit, and n is the number of quantum bits used to represent the vector;

[0018] Optionally, the conversion of the quantum encoding layer can be represented as:

[0019] ψ enc = U(θ) | ψ >

[0020] In the formula, ψ enc represents the encoded quantum state, U(θ) is a quantum gate adjusted according to the characteristics of the input data, θ represents the parameters of the quantum gate, and | ψ > is the initial quantum state;

[0021] Optionally, the quantum gate layer is represented as:

[0022] ψ pro = U pro (γ) ψ enc

[0023] In the formula, ψ pro represents the quantum state after forward propagation, U pro (γ) is a quantum gate used in the forward propagation process, and γ is a quantum gate parameter;

[0024] Optionally, the output layer is represented as:

[0025]

[0026] In the formula, ψ dec represents the MICM phenotype typing feature obtained after decoding, and ψ prois to represent the quantum state feature information extracted after forward propagation, is the conjugate transpose of the quantum encoding operation, used to map the quantum state feature information into the MICM phenotype typing feature;

[0027] Optionally, the quantum gate U(θ) is represented as:

[0028] U(θ) = e -iθH

[0029] In the formula, H is the Hamiltonian;

[0030] Optionally, the adjusted Hamiltonian is obtained by using a dynamic quantum topology optimization strategy, and the quantum gate U(θ) is obtained by using the adjusted Hamiltonian;

[0031] Optionally, the dynamic adjustment mode of the dynamic quantum topology optimization strategy can be represented as:

[0032]

[0033] λ jk = softmax(-d jk / τ te )

[0034] In the formula, softmax() is a preset Softmax classification function, H DQTO is the dynamically adjusted Hamiltonian; λ jk is the coupling strength between quantum bits j and k, and are Pauli-Z operations acting on quantum bits j and k, d jk represents the feature correlation distance between quantum bits j and k, τ te is a temperature parameter.

[0035] Further, the construction step of the quantum encoding-based Transformer includes:

[0036] The initial quantum encoding-based Transformer includes an initial input layer, an initial quantum encoding layer, an initial quantum gate layer, and an initial output layer, and the parameters of the initial quantum encoding layer and the initial quantum gate layer are randomly initialized parameters;

[0037] After the vectorized diagnostic report of the training set is subjected to the initial quantum encoding-based Transformer to obtain the initial MICM phenotype typing feature, the model prediction output is obtained based on the initial MICM phenotype typing feature, the difference between the actual value of the training set and the model prediction output is compared to a loss function, and the loss function is iteratively trained until a stop condition is reached to obtain the quantum encoding-based Transformer.

[0038] The stop condition comprises reaching a preset maximum number of iterations;

[0039] Optionally, the loss function is calculated based on the initial MICM phenotyping feature, and is represented as:

[0040]

[0041] In the formula, L represents the loss function, i∈[1,m], i represents the i th sample in the training set, m represents the total number of samples in the training set, y i is the target output of the i th sample, ψ dec,i represents the MICM phenotyping feature of the i th sample, f(ψ dec,i ) is the predicted output of the i th sample of the model.

[0042] Optionally, the updating manner of the parameters of the initial quantum encoding layer is represented as:

[0043]

[0044] In the formula, θ new and θ old respectively represent the initial quantum encoding layer parameters after and before updating, η rate is the learning rate, is the gradient of the loss function with respect to the initial quantum encoding layer parameters θ.

[0045] Optionally, the updating manner of the parameters of the initial quantum gate layer is represented as:

[0046]

[0047] In the formula, γ new and γ old respectively represent the initial quantum gate layer parameters after and before updating, η2 is the learning rate, is the gradient of the loss function with respect to the initial quantum gate layer parameters γ.

[0048] Optionally, the stop condition comprises that the quantum entanglement degree is less than a preset threshold value: after the forward propagation of the quantum gate layer obtains the forward propagation extracted quantum state feature, the entanglement degree between different quantum bits is calculated, and when the entanglement degree is less than the preset threshold value, it indicates that the model converges, and the training is stopped.

[0049] Optionally, the entanglement degree is represented as:

[0050] E = Tr(ρ A logρ A )

[0051] In the formula, E represents the quantum entanglement degree, ρ Ais the reduced density matrix of system A, system A represents the quantum state feature information extracted in the last sub-iteration, and Tr() represents a trace operation,

[0052] Further, the reduced density matrix ρ A is obtained by a trace operation from a total density matrix ρ, the total density matrix ρ is a feature matrix of the quantum state feature information extracted in the current iteration, and the reduced density matrix ρ A can be represented as:

[0053] ρ A = Tr B (ρ)

[0054] In the formula, Tr B () represents taking a trace on the system B part and leaving the state of the system A part, and system B is the quantum state feature information extracted in the current iteration.

[0055] Further,

[0056] The diagnostic report includes morphological information, immunological information, cytogenetic information, and molecular information.

[0057] Optionally, the method for obtaining a structured diagnostic report based on the diagnostic report includes using a machine vision technology to extract the structured diagnostic report according to a standard report model.

[0058] Optionally, the method for obtaining a vectorized diagnostic report based on the structured diagnostic report includes one or more of the following methods: Word2Vec, Doc2Vec, and BERT.

[0059] Further, the artificial intelligence feature extraction method based on lymphoma MICM phenotyping based on quantum coding further includes S105: obtaining patient diagnosis and treatment information based on the MICM phenotyping feature.

[0060] Optionally, it further includes S105': inputting the MICM phenotyping feature into a classifier to obtain a classification result of whether the subject has a hematological tumor.

[0061] Optionally, the MICM diagnostic report includes the treatment information of the patient and the classification result.

[0062] Optionally, the classifier includes one or more of the following: conditional random field algorithm, support vector machine, decision tree, random forest, and logistic regression.

[0063] Further, the classifier is a trained conditional random field, and a method for constructing the trained conditional random field includes:

[0064] obtaining a diagnostic report training set, the training set further comprising a label of whether suffering from a blood tumor or not;

[0065] obtaining a MICM phenotyping feature training set based on the diagnostic report training set;

[0066] initializing parameters of a conditional random field to obtain an initialized conditional random field;

[0067] Monte Carlo random sampling is performed on the parameters of the initialized conditional random field to obtain sample parameters;

[0068] obtaining a predicted label of the MICM phenotyping feature training set based on the sample parameters, and calculating a loss function of the conditional random field based on the predicted label of the MICM phenotyping feature training set and an actual label of the diagnostic report, gradually minimizing the loss function of the conditional random field through iteration until a stop condition is reached to obtain the trained conditional random field;

[0069] Optionally, the loss function of the conditional random field is represented as:

[0070]

[0071] wherein, L(q θ ) represents a loss function based on parameters q θ of the conditional random field, is label data of the i-th sample, q θ (t) is the parameters of the conditional random field obtained through Monte Carlo sampling at the t-th iteration in the iteration process, is a label prediction probability under given feature and parameters q θ (t) λ L1 is an L1 regularization coefficient, and ||q θ ||1 is an L1 regularization term of the parameters q θ (t) .

[0072] Optionally, the sample parameters obtained through Monte Carlo random sampling are represented as:

[0073] q θ (t) = q θ (t-1) + ∈ (t)

[0074] wherein, q θ (t) represents sample parameters of the t-th iteration, q θ (t-1) represents sample parameters of the t-1-th iteration, and ∈ (t)a sampling step size;

[0075] Optionally, the sampling step size ∈ (t) is set as follows for dynamic adaptation:

[0076] ∈ (t) = κ cy · exp(- ρ cy t)

[0077] wherein κ cy is an initial step size, and ρ cy is a decay factor;

[0078] Optionally, after the sample parameters are optimized by using the sparse regularization strategy, the loss function of the conditional random field is calculated by using the optimized sample parameters to replace the sample parameters.

[0079] Optionally, the sparse regularization strategy is expressed as follows:

[0080]

[0081] wherein q represents the optimized sample parameters at the tthiteration, Δq θ (t) is an increment of the sample parameters q θ (t) at the tthiteration, α gz is a learning rate, is a gradient of the loss function of the conditional random field with respect to the parameters, and δ gz is an adjustment factor calculated based on a difference between a model output and an actual classification result, and sign(q θ (t) ) represents a sign function of q θ (t) .

[0082] Optionally, the stop condition includes that an information gain of iteration is less than a preset threshold value: after each iteration, an information gain is calculated based on the sample parameters of the current conditional random field, and the iteration is stopped when the information gain is less than the preset threshold value.

[0083] Optionally, represents a label prediction probability under given features and parameters q θ (t) , and all can be calculated when the sample parameters of the current conditional random field are determined.

[0084] wherein c ∈ [1, C], c represents a classification label, C is all possible categories, and the information gain is expressed as follows:

[0085] where Es i represents the information gain of the i-th sample, is the probability of class c given the feature x, p(c) is the prior probability of class c.

[0086] The second aspect of the present application discloses an artificial intelligence feature extraction system based on lymphoma MICM phenotyping, comprising:

[0087] The acquisition module 201 is configured to acquire a diagnosis report of a to-be-tested person.

[0088] The structured processing module 202 is configured to obtain a structured diagnosis report based on the diagnosis report.

[0089] The vectorization processing module 203 is configured to vectorize the structured diagnosis report to obtain a vectorized diagnosis report.

[0090] The feature extraction module 204 is configured to input the vectorized diagnosis report into a quantum coding-based Transformer to perform feature extraction and obtain MICM phenotyping features; the quantum coding-based Transformer converts input data into a quantum state, extracts quantum state feature information, converts the quantum state feature information into non-quantum state data, and outputs MICM phenotyping features.

[0091] The third aspect of the present application discloses a computer device, comprising a memory and a processor; the memory is configured to store program instructions; the processor is configured to invoke the program instructions, and when the program instructions are executed, the steps of the above method are executed.

[0092] The fourth aspect of the present application discloses a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the above method.

[0093] The fifth aspect of the present application discloses a computer program product, comprising a computer program, which is executed by a processor to implement the steps of the above method.

[0094] The present application has the following beneficial effects:

[0095] 1. The present application extracts the required MICM phenotyping features based on the quantum coding-based Transformer, and uses the quantum coding-based Transformer algorithm to enhance the model's ability to capture data features;

[0096] 2. A dynamic quantum topology optimization strategy is used to optimize the connection topology structure between quantum bits to improve the efficiency of quantum coding and the information processing capacity of the model in the feature extraction process.

[0097] 3. The application also utilizes a conditional random field algorithm based on Monte Carlo optimization for model training, and improves processing efficiency and accuracy through a sparse representation strategy.

[0098] 4. Based on the innovative technology of the application, the data processing capability can be effectively improved in the MICM diagnosis task, the adaptability of the model to different data characteristics is enhanced, and the classification accuracy and prediction reliability of the model in the medical diagnosis task are improved. BRIEF DESCRIPTION OF DRAWINGS

[0099] In order to more clearly illustrate the technical solutions in the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0100] Figure 1 is a method flowchart provided by the first aspect of the embodiment of the application;

[0101] Figure 2 is a program product schematic diagram provided by the second aspect of the embodiment of the application;

[0102] Figure 3 is a schematic diagram of a computer device provided by the embodiment of the application;

[0103] Figure 4 is a schematic diagram of the architecture of an exemplary computing device provided by the embodiment of the application;

[0104] Figure 5 is a schematic diagram of a storage medium provided by the embodiment of the application;

[0105] Figure 6 is a flowchart of the MICM diagnosis report generation and release provided by the embodiment of the application;

[0106] Figure 7 is a flowchart of MICM phenotype typing feature extraction based on a quantum encoding Transformer algorithm provided by the embodiment of the application;

[0107] Figure 8 is a training flowchart of a quantum encoding Transformer algorithm provided by the embodiment of the application;

[0108] Figure 9 is a training flowchart of a conditional random field algorithm based on Monte Carlo optimization. DETAILED DESCRIPTION

[0109] In order to better understand the technical scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application.

[0110] In some of the processes described in this specification and in the accompanying drawings, multiple operations are described in a specific order. However, it should be understood that these operations can be performed in a different order or in parallel, and that the sequence of operations is not necessarily the order in which the operations appear in this text. The sequence of operations is indicated by the sequence of numbers such as S101, S102, etc., which are merely used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and the operations can be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this text are used to distinguish different messages, devices, modules, etc., and do not represent the order or limit the types of "first" and "second".

[0111] The technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0112] Figure 1 is a flowchart of an artificial intelligence feature extraction method based on evaluation of lymphoma MICM phenotyping provided by an embodiment of the present application. Specifically, the method comprises the following steps:

[0113] S101: obtaining a diagnosis report of a subject to be tested;

[0114] In some embodiments, the obtained report is a medical report, including a cell morphology examination report, a flow cytometry analysis report, a cytogenetic examination report, and a fluorescence in situ hybridization examination report.

[0115] S102: obtaining a structured diagnosis report based on the diagnosis report;

[0116] In some embodiments, the obtained diagnosis report is a structured report, and the report content is extracted according to a standard report model to obtain a structured diagnosis report. Figure 6 ).

[0117] In some embodiments, the obtained diagnosis report is an unstructured report, and the report content is extracted according to a standard report model using computer vision technology to obtain a structured report Figure 6 ).

[0118] S103: vectorizing the structured diagnosis report to obtain a vectorized diagnosis report;

[0119] The method of vectorizing the structured diagnostic report to obtain a vectorized diagnostic report includes one or more of the following methods: Word2Vec, Doc2Vec, BERT.

[0120] In some embodiments, the structured diagnostic report is vectorized to obtain a vectorized diagnostic report using BERT.

[0121] In some embodiments, the structured diagnostic report is vectorized to obtain a vectorized diagnostic report using Word2Vec.

[0122] S104: The vectorized diagnostic report is input into a quantum coding-based Transformer to perform feature extraction to obtain MICM phenotyping features; the quantum coding-based Transformer converts input data into a quantum state and extracts quantum state feature information, and converts the quantum state feature information into non-quantum state data to output the features of the MICM.

[0123] In one embodiment, a blood tumor MICM comprehensive diagnosis method and system mainly consists of the following modules:

[0124] (1) Machine vision (OCR) enhanced recognition model module: identifies picture format reports and extracts content in files.

[0125] (2) Report content model library: collect structured reports such as PDF and Word file format reports.

[0126] (3) Database operation module: connect to the database using oledb.

[0127] (4) SFTP file upload module: secure file transfer protocol, ensures file safety during upload.

[0128] (5) Clinical release module: release the final report to the clinical department for access in the form of a browser.

[0129] (6) HIS interface module: automatically extract clinical order information into the system.

[0130] (7) CA signature module: tamper-proof and traceable.

[0131] (8) Report automatic fusion module: identifies information in the report file and merges the file with the order information.

[0132] (9) File conversion module: converts jpg, word, and other format files to PDF format.

[0133] (10) File encryption module: PDF file password encryption.

[0134] The data model involved in the method and system:

[0135] 1 Cell morphology examination report, flow cytometry analysis report, cytogenetic examination report, fluorescence in situ hybridization examination report, 1000 points of each type of report randomly selected from 4000 reports as training data set.

[0136] 2 Convert semi-structured data into structured data and store it in the database.

[0137] 3 Unstructured data, mainly in the form of pictures. First, train a machine vision text recognition model library for samples, and then convert all sample data into the database through the machine vision text recognition model library. Analyze the sample data collected in the database.

[0138] 4 Form a report content model library and a machine vision text recognition model library.

[0139] The method and system generate sub-reports:

[0140] Monitor the report folder, upload the report to the server, identify the content in the report, and combine it with the patient's medical order information ( Figure 6 ). Among them,

[0141] ① Determine whether the collected reports in the monitoring folder are complete.

[0142] The system generates an independent "monitoring folder" according to each patient's examination application, and determines whether the "monitoring folder" contains the automatic cell morphology examination report, flow cytometry analysis report, cytogenetic examination report, fluorescence in situ hybridization examination report and other reports listed in the patient's examination application.

[0143] ② Determine the format of the collected reports in the monitoring folder.

[0144] The system automatically determines whether the collected reports in the monitoring folder are structured reports or unstructured reports, and whether the report form is PDF or Jpg.

[0145] ③ Use machine vision technology to extract unstructured report content according to standard report models.

[0146] ④ Structured standard diagnostic report and clinical medical order information are edited by artificial intelligence to generate MICM diagnostic report.

[0147] Among them, the process of generating MICM diagnostic report by artificial intelligence editing of structured standard diagnostic report and clinical medical order information (also known as Figure 6IV) is generated by using a natural language model, which is a generative model obtained by training, and the specific training process is as shown in Figure 8

[0148] In one specific embodiment, the training of the natural language model comprises the following steps:

[0149] Step 1: Data collection and labeling

[0150] The present application is used for training a natural language model as a generative model, and the training data is collected from various medical reports, including cytological examination reports, flow cytometry analysis reports, cytogenetic examination reports, and fluorescence in situ hybridization examination reports. The structured reports are obtained by extracting the report content according to the standard report model or using machine vision technology to extract the report content according to the standard report model. The collected data is labeled, and the labeling method of the present application is manual labeling, which is analyzed by professional medical experts according to medical standards, and a structured diagnostic report is obtained. The structured diagnostic report constitutes the training data in the form of "question and answer pairs" (as shown in Table 1), and the natural language model Transformer algorithm model is supervised trained; further, according to the collected data, the diagnostic results are labeled, i.e., the training data is constituted in the form of "question + type".

[0151] Table 1: Question and answer pair example

[0152]

[0153]

[0154] Step 2: Training of the Transformer algorithm model based on quantum coding

[0155] The training data collected in step 1 is feature extracted to obtain MICM phenotype typing features, and the present application uses a Transformer algorithm based on quantum coding for feature extraction, which converts the input data into a quantum state through a quantum coding layer, and uses the entanglement and superposition characteristics of quantum information to enhance the model's ability to capture data features.

[0156] The training process of the Transformer algorithm based on quantum coding is as shown in Figure 9

[0157] Step 2-1: encode the training data in text format into vector data,

[0158] The Word2Vec algorithm is used to encode the training data in text format into vector data, i.e., a shallow neural network is used to map words to a vector space, so that semantically similar words are close to each other in the vector space.​​

[0159] Further, the initial quantum state is configured for each input of vector data encoded as vector data, and the configuration of the quantum state is represented as:

[0160]

[0161] In the formula, ψ represents a quantum state, α i is a complex amplitude, |i> is the ground state of a quantum bit, and n is the number of quantum bits.

[0162] The initial value of the parameter θ of the initialized quantum gate is set to π / 4.

[0163] Step 2-2: Perform encoding mapping through the quantum encoding layer to map the initial quantum state of the input data to quantum state data;

[0164] Unlike the traditional Transformer algorithm model, the encoding method of the quantum encoding-based Transformer algorithm performs encoding mapping through the quantum encoding layer: the input data is converted into a quantum state through the quantum encoding layer, that is, each feature of the data is encoded onto a quantum bit, which can be represented as:

[0165] ψ enc = U(θ) | ψ >

[0166] In the formula, ψ enc represents the encoded quantum state, U(θ) is a quantum gate operation adjusted according to the characteristics of the input data, θ represents the parameter of the quantum gate, and | ψ > is the original quantum state.

[0167] The calculation method of | ψ > can be represented as:

[0168]

[0169] In the formula, | ψ > represents a quantum state, ns is the total number of features, p i is the probability of the i-th feature being selected; and the feature is output based on the probability of the feature being selected;

[0170] Further, the calculation method of the probability p i of the feature being selected can be represented as:

[0171]

[0172] In the formula, β cy is an adjustment parameter; Es i is an information gain score based on the i-th feature, which is calculated by a conditional random field algorithm optimized based on the Monte Carlo method; μ cy is the average value of the information gain.

[0173] Further, the implementation of the quantum gate operation U(θ) can be represented as:

[0174] U(θ) = e -iθH

[0175] where H is the Hamiltonian.

[0176] In one embodiment, the Hamiltonian H is characterized by Pauli matrices, which can be represented as:

[0177]

[0178] where are Pauli matrices corresponding to X, Y and Z operations of the qubits.

[0179] Step 2-3: Optimize the connection topology between qubits using a dynamic quantum topology optimization strategy to improve the efficiency of quantum coding and the information processing capacity of the model.

[0180] Specifically, at the quantum coding layer, the initial connection of each qubit is automatically configured based on the preliminary analysis of the data. As the training progresses, the connection is dynamically adjusted according to the changes in the data stream. The adjustment method can be represented as:

[0181]

[0182] λ jk = softmax(-d jk / τ te )

[0183] where softmax() is a preset Softmax classification function, H DQTO is the dynamically adjusted Hamiltonian; λ jk is the coupling strength between qubits j and k, and are Pauli-Z operations acting on qubits j and k. d jk represents the feature correlation distance between qubits j and k, τ te is a temperature parameter. Preferably, τ te is set to 3.

[0184] Step 2-4: The quantum state undergoes entanglement and disentanglement operations between quantum states through quantum logic gates, which are controlled by the parameters learned during the training process to maximize the exchange of information between features, which can be represented as:

[0185] ψ pro = U pro (γ)ψ enc

[0186] In the formula, ψ pro represents the quantum state after forward propagation, U pro (γ) is a quantum gate operation used in the forward propagation process, and γ is a quantum gate parameter. Preferably, γ is set to π / 2.

[0187] Further, the quantum gate operation U pro (γ) can be calculated as:

[0188]

[0189] In the formula, each H k represents a Hamiltonian operating on different quantum levels, and γ k is a corresponding parameter, and m represents the total number of operation levels. Preferably, γ k is set to π / 4.

[0190] In steps 2-5, after the quantum gate operation, the entanglement degree between different quantum bits is calculated,

[0191] The calculation method can be represented as:

[0192] E = Tr(ρ A logρ A )

[0193] In the formula, E represents the quantum entanglement degree, ρ A is the reduced density matrix of system A, and Tr() represents the trace operation. Further, the reduced density matrix ρ A is obtained from the total density matrix ρ by trace operation, which can be represented as:

[0194] ρ A = Tr B (ρ)

[0195] In the formula, Tr B () represents taking the trace of system B part, leaving the state of system A part. If system B is the remaining n-1 quantum bits, it means that only the state of a specific quantum bit is considered from the state of the total system.

[0196] In this embodiment, system B and system A respectively represent the transformed features extracted in this iteration and the transformed features extracted in the last iteration, and the transformed features are the quantum state features output by the quantum gate layer.

[0197] Step 2-6: Unlike the traditional Transformer algorithm model, the decoding method of the quantum encoding-based Transformer algorithm maps the entangled quantum state back to classical information through inverse quantum encoding operation,

[0198] which can be represented as:

[0199]

[0200] where ψ dec represents the decoded quantum state, is the conjugate transpose of the quantum encoding operation, used for the inverse operation.

[0201] Steps 2-7: Calculate the loss function and adjust the parameters of the quantum encoding layer and the quantum gate layer through the backpropagation algorithm, which can be represented as:

[0202]

[0203] where L represents the loss function, y i is the target output, f(ψ dec,i ) is the predicted output of the model, is the gradient of the loss function with respect to the quantum gate parameters θ.

[0204] Further, by bringing the loss function into the partial derivative calculation formula, the calculation method of can be obtained as:

[0205]

[0206] Further, by the chain rule, which can be represented as:

[0207]

[0208] Further, involving the differentiation of the inverse quantum encoding operation, the calculation method can be represented as:

[0209]

[0210] where H DQTO is the dynamically adjusted Hamiltonian.

[0211] Further, the update of λ jk is based on the backpropagation algorithm, which can be represented as:

[0212]

[0213] where η cg is the learning rate, is the gradient of the loss function with respect to the coupling strength, calculated by the chain rule. Preferably, η cg is set to 0.03.

[0214] Further, the update method of the parameters of the quantum gate can be represented as:

[0215]

[0216] wherein θ new and θ old represent the quantum gate parameters after and before updating, respectively, and η rate is the learning rate. Preferably, η rate is set to 0.01.

[0217] Steps 2-8: Repeat the above steps iteratively until a preset stopping iteration condition is met, indicating that the model training is completed.

[0218] The preset stopping iteration condition is to reach a preset maximum number of iterations, and in an embodiment, preferably, the preset maximum number of iterations is set to 1000.

[0219] In some embodiments, the preset stopping iteration condition is that when the entanglement degree is less than a preset threshold, it indicates that the model converges, and then the training is stopped

[0220] Step 3, training of the conditional random field algorithm model based on the Monte Carlo method optimization

[0221] The output features of the encoding part of the quantum coding based Transformer algorithm are input into the conditional random field algorithm for training of the auxiliary diagnosis model. The conditional random field is a statistical modeling method, and the conditional random field algorithm based on the Monte Carlo method optimization is adopted in the present application, and a sparse representation strategy is adopted to improve the efficiency and accuracy of the model in processing data. Specifically, the training process of the conditional random field algorithm based on the Monte Carlo method optimization is as shown in Figure 9 .

[0222] Step 3-1: Set the output features of the encoding part of the quantum coding based Transformer algorithm as q x , and initialize the parameters q θ of the conditional random field model, and the initialization method is random initialization, which can be represented as:

[0223] q θ = N(0, σ 2 I)

[0224] wherein q θ represents the parameter vector of the conditional random field model, N(0, σ 2 I) represents that the parameters are initialized as a normal distribution with a mean of 0 and a variance of σ 2 , and I is an identity matrix. Preferably, σ 2 is set to 0.01.

[0225] Step 3-2: In each iteration, the model parameters q θ are randomly sampled using the Monte Carlo method, which can be represented as:

[0226] q θ (t) = q θ (t-1) + ∈ (t)

[0227] where q θ (t) and q θ (t-1) are the parameters after the tth and (t-1)th iteration, ∈ (t) is the sampling step size;

[0228] Further, the sampling step size ∈ (t) is dynamically adaptive, and is set in the following way:

[0229] ∈ (t) = κ cy · exp(-ρ cy t)

[0230] where κ cy is the initial step size, and ρ cy is the decay factor, which adjusts the step size with the iteration number t to ensure the stability and convergence of the sampling process. Preferably, κ cy is set to 1, and ρ cy is set to 0.95.

[0231] Step 3-3: Using the parameters obtained from the Monte Carlo sampling, train the conditional random field model by optimizing the objective function L(q θ ), the model will learn how to effectively classify based on the input feature vector, and the objective function can be calculated in the following way:

[0232]

[0233] where y is the label data, is the label prediction probability under the given feature and parameter q θ , and λ L1 is the L1 regularization coefficient, and ||q θ ||1 is the L1 regularization term. Preferably, λ L1 is set to 0.3.

[0234] Further, the probability can be calculated in the following way:

[0235]

[0236] where y' represents the feature function, which converts the input features and labels into numerical values ​​for model training; y' represents all possible label configurations.

[0237] Steps 3-4: Optimize parameter q using a sparse regularization strategy during model training. θ By calculating the increment Δq θ Update parameter q θ To ensure that the model focuses on the features most critical to classification, the calculation method can be expressed as:

[0238]

[0239] In the formula, α gz It's the learning rate. It is the gradient of the loss function with respect to the parameters, δ gz It is an adjustment factor calculated based on the difference between the model output and the actual classification result, sign(q) θ ) represents q θ The sign function. Preferably, α gz Set to 0.01.

[0240] Steps 3-5: Further, gradient The calculation method can be expressed as:

[0241]

[0242] In the formula, This represents the expected distribution of the prediction for all possible label configurations y′ under the current parameters.

[0243] In one embodiment, the adjustment factor δ gz The calculation method can be expressed as:

[0244]

[0245] In the formula, η gz It is the adjustment factor learning rate, y i These are the actual category labels. This represents the category predicted by the model, and tanh() is the hyperbolic tangent function. Preferably, η gz Set to 0.05.

[0246] Steps 3-5: Calculate the information gain Es after each iteration. i The calculation method can be expressed as:

[0247]

[0248] In the formula, c represents the category label, C is all possible categories, and p(c|x i ) is a given feature xi the probability of class c under the condition of the features

[0249] Further, the performance of the model on the training set is evaluated, and in one embodiment, the performance of the model on the training set is measured by training accuracy, and the prediction label for calculating the training accuracy can be expressed as:

[0250]

[0251] wherein, is an indicator function, is the predicted label under the given features and model parameters q θ is the label prediction probability under the given features and parameters q θ

[0252] The above steps are repeated iteratively until a preset stopping iteration condition is met, i.e., the model training is completed. In one embodiment, the preset stopping iteration condition is that the training accuracy of the model on the training set reaches 95% or more, or reaches a preset maximum number of iterations, preferably, the preset maximum number of iterations is set to 1000.

[0253] Step 4: Generate MICM diagnosis report

[0254] Generate the MICM diagnosis report using the trained model.

[0255] In one embodiment, the MICM diagnosis report is generated for the obtained report content, and the MICM diagnosis report is in a standardized output format, including two parts, one part is the diagnosis process and treatment suggestion output by the quantum coding-based Transformer algorithm model, and the other part is the auxiliary diagnosis of the disease category output by the conditional random field algorithm model optimized based on the Monte Carlo method.

[0256] In one embodiment, for the report content: “the patient has irregular heart rhythm and chest pain symptoms”, the standardized output format is: “the patient may have non-ST segment elevation myocardial infarction (NSTEMI), and it is recommended to immediately perform blood biochemical detection and electrocardiogram examination.” And “disease auxiliary diagnosis category: acute coronary syndrome”.

[0257] In one embodiment, the MICM diagnosis report includes the treatment information of the patient and the classification result.

[0258] ​​The MICM report generation of the method and system: the system automatically completes the standardized sub-diagnosis report after extracting the content of the unstructured report such as the cell morphology examination report, the flow cytometry analysis report, the cytogenetic examination report, and the fluorescence in situ hybridization examination report, gathers the patient medical order information, and automatically generates the MICM standard report, and after the doctor audits, the report is published to the clinic.

[0259] Figure 3 is a schematic diagram of a computer device provided by an embodiment of the application, as shown in Figure 3 The device can include one or more processors and one or more memories, wherein the memory stores computer readable code which, when executed by the one or more processors, can perform the method described above.

[0260] The processor in this embodiment can be an integrated circuit chip with signal processing capability. The processor can be a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The methods, operations and logic block diagrams disclosed in the embodiments of the disclosure can be implemented or executed. The general purpose processor can be a microprocessor or the processor can also be any conventional processor or the like, which can be of X86 architecture or ARM architecture.

[0261] Generally, various example embodiments of the present disclosure can be implemented in hardware or special-purpose circuitry, software, firmware, logic, or any combination thereof. Certain aspects can be implemented in hardware, while other aspects can be implemented in firmware or software which can be executed by a controller, microprocessor or other computing device. When aspects of the present disclosure are illustrated or described as a block diagram, flow chart, or using some other pictorial representation, it will be understood that the blocks, apparatus, systems, techniques or methods described herein can be implemented in hardware, software, firmware, special-purpose circuitry or logic, general purpose hardware or controller or other computing device, or some combination thereof.

[0262] For example, the method or device according to the embodiments of the present disclosure can also be implemented by means of the architecture of the computing device 3000 as shown in Figure 4 Figure 4 ​As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the methods provided in this disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 4 The architecture shown is merely exemplary and can be omitted as needed when implementing different devices. Figure 4 One or more components in the computing device shown.

[0263] This invention also includes a computer-readable storage medium, such as... Figure 5 The diagram illustrates a storage medium provided in an embodiment of the present invention. The computer storage medium 4020 stores computer-readable instructions 4010. When the computer-readable instructions 4010 are executed by a processor, the method described above according to embodiments of the present disclosure can be performed. The computer-readable storage medium in the embodiments of the present disclosure can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Synchronous Link Dynamic Random Access Memory (SLDRAM), and Direct Memory Bus Random Access Memory (DR RAM). It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0264] This disclosure also provides a computer program product or computer program that, when executed by a processor, implements the steps of the above-described method. For example... Figure 2 As shown, the computer program product or computer program includes:

[0265] The acquisition module 201 is configured to acquire a diagnostic report of a subject;

[0266] The structured processing module 202 is configured to obtain a structured diagnostic report based on the diagnostic report;

[0267] The vectorization processing module 203 is configured to vectorize the structured diagnostic report to obtain a vectorized diagnostic report;

[0268] The feature extraction module 204 is configured to input the vectorized diagnostic report into a quantum coding-based Transformer to perform feature extraction and obtain MICM phenotyping features; the quantum coding-based Transformer converts input data into a quantum state to extract quantum state feature information, and converts the quantum state feature information into non-quantum state data to output MICM phenotyping features.

[0269] It should be noted that the flowcharts and block diagrams in the drawings illustrate the possible architectural, functional, and operational scenarios of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment, or a portion of code that includes one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that shown in the drawings. For example, two blocks that are shown in succession can actually be executed substantially in parallel, and sometimes in reverse order, depending on the involved functions. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0270] Generally, various example embodiments of the present disclosure can be implemented in hardware or special-purpose circuits, software, firmware, logic, or any combination thereof. Certain aspects can be implemented in hardware, while other aspects can be implemented in firmware or software which can be executed by a controller, microprocessor, or other computing device. When various aspects of embodiments of the present disclosure are illustrated or described as a flowchart, a flow diagram, or using some other graphical representation, it is to be understood that the described blocks, devices, systems, techniques, or methods can be implemented in hardware, software, firmware, special-purpose circuits, or logic, general-purpose hardware or controllers, or other computing devices, or some combination thereof.

[0271] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0272] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0273] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to the actual needs to achieve the purposes of the embodiments of the present application.

[0274] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit.

[0275] The example embodiments of the present disclosure described in detail above are merely illustrative, and are not limiting. Those skilled in the art should understand that various modifications and combinations of these embodiments or their features can be made without departing from the principles and spirits of the present disclosure, and such modifications should fall within the scope of the present disclosure.

Claims

1. An artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping, characterized in that, The method comprises: S1: obtaining a diagnostic report of a subject to be tested; S2: obtaining a structured diagnostic report based on the diagnostic report; S3: vectorizing the structured diagnostic report to obtain a vectorized diagnostic report; S4: inputting the vectorized diagnostic report into a quantum coding-based Transformer to perform feature extraction and obtain MICM phenotyping features; the quantum coding-based Transformer converts input data into a quantum state and extracts quantum state feature information, and then converts the quantum state feature information into non-quantum state data to output the MICM phenotyping features.

2. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 1, characterized in that, The quantum coding-based Transformer comprises an input layer, a quantum coding layer, a quantum gate layer, and an output layer; and the processing steps of the quantum coding-based Transformer comprise: Step 1: the input layer configures an initial quantum state for each vector of the vectorized diagnostic report to obtain a vector of an initial quantum state; Step 2: the quantum coding layer performs quantum coding on the vector of the initial quantum state to obtain a quantum-coded vector; Step 3: the quantum-coded vector is forwarded through the quantum gate layer to obtain forward-propagated extracted quantum state features; Step 4: the output layer maps the extracted quantum state features back to a non-quantum state to obtain the MICM phenotyping features.

3. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 2, characterized in that, The steps further comprise a step 3' between step 2 and step 3: optimizing the correlation distance between quantum bits of the quantum-coded vector through a dynamic quantum topology optimization strategy.

4. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 2, characterized in that, The manner of configuring the initial quantum state for each vector is represented as: wherein denotes the initial quantum state, is a complex amplitude, is the ground state of the qubit, is the number of qubits used to represent the vector.

5. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 4, characterized in that, The conversion of the quantum coding layer is represented as: wherein represents an encoded quantum state, is a quantum gate adjusted according to characteristics of input data, represents a parameter of the quantum gate, is an initial quantum state.

6. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 2, characterized in that, The quantum gate layer is represented as: In the formula, denotes the quantum state after forward propagation, is a quantum gate used in the forward propagation process, is a quantum gate parameter, denotes the encoded quantum state.

7. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 2, characterized in that, The output layer is represented as: wherein, denotes the MICM phenotyping feature obtained after decoding, is the quantum state feature information extracted after forward propagation, is the conjugate transpose of the quantum encoding operation for mapping the quantum state feature information into the MICM phenotyping feature.

8. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 2, characterized in that, Quantum gate is represented as: wherein, is a Hamiltonian; the Hamiltonian is adjusted by using a dynamic quantum topology optimization strategy; a dynamic adjustment manner of the dynamic quantum topology optimization strategy is represented as: , , In the formula, is a preset Softmax classification function, is a dynamically adjusted Hamiltonian; is a quantum bit and is a coupling strength between the quantum bits and is a Pauli-Z operation acting on the quantum bits and , represents a feature correlation distance between the quantum bits and , is a temperature parameter. 9.The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 1, characterized in that, The construction steps of the quantum coding-based Transformer comprise: The initial quantum coding-based Transformer comprises an initial input layer, an initial quantum coding layer, an initial quantum gate layer, and an initial output layer, and the parameters of the initial quantum coding layer and the initial quantum gate layer are randomly initialized parameters; After the vectorized diagnostic reports of the training set are input into the initial quantum coding-based Transformer to obtain output MICM phenotyping features, a model prediction output is obtained based on the output MICM phenotyping features, the difference between the actual value of the training set and the model prediction output is compared to a loss function, and the loss function is iteratively trained until a stop condition is reached to obtain the quantum coding-based Transformer; The stop condition comprises reaching a preset maximum number of iterations.

10. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 9, characterized in that, The loss function is calculated based on the output MICM phenotyping features and is represented as: wherein, represents a loss function, , i represents the i-th sample in the training set, and m represents the total number of samples in the training set, is the target output of the i-th sample, represents the MICM phenotyping feature of the output of the i-th sample, is the predicted output of the i-th sample by the model.

11. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 9, characterized in that, The updating manner of the parameters of the initial quantum coding layer is represented as: wherein, and respectively represent the initial quantum encoding layer parameters after and before the update, is a learning rate, is a gradient of the loss function with respect to the initial quantum encoding layer parameters .

12. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 9, characterized in that, The updating manner of the parameters of the initial quantum gate layer is represented as: wherein, and denote the parameters of the updated and the initial quantum gate layer, respectively, is the learning rate, is the gradient of the loss function with respect to the initial quantum gate layer parameters .

13. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 9, characterized in that, The stop condition comprises that the quantum entanglement degree is less than a preset threshold: after the forward-propagated extracted quantum state features are obtained by forwarding through the quantum gate layer, the entanglement degree between different quantum bits is calculated, and when the entanglement degree is less than the preset threshold, it indicates that the model converges, and the training is stopped; The entanglement degree is represented as: , wherein represents the quantum entanglement degree, is the reduced density matrix of system A, system A representing the quantum state characteristic information extracted in the last iteration, represents the trace operation; the reduced density matrix from the total density matrix is obtained by trace operation, the total density matrix is the feature matrix of the quantum state feature information extracted for this iteration, the reduced density matrix is expressed as: , In the formula, represents the trace taken on the system B part, leaving the state of the system A part, and the system B is the quantum state feature information extracted in this iteration.

14. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 1, characterized in that, The diagnostic report comprises morphological information, immunological information, cytogenetic information, and molecular information.

15. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 1, characterized in that, The method for obtaining a structured diagnostic report based on the diagnostic report comprises using machine vision technology to extract the structured diagnostic report according to a standard report model.

16. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 1, characterized in that, The method for vectorizing the structured diagnostic report to obtain a vectorized diagnostic report comprises one or more of the following methods: Word2Vec, Doc2Vec, and BERT.

17. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 1, characterized in that, The method for extracting artificial intelligence features based on lymphoma MICM phenotyping based on quantum coding further comprises S5: obtaining patient diagnosis and treatment information based on the MICM phenotyping features. The method further comprises S5': inputting the MICM phenotyping features into a classifier to obtain a classification result of whether the subject has a hematological tumor.

18. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 17, characterized in that, The classifier comprises one or more of the following: conditional random field algorithm, support vector machine, decision tree, random forest, and logistic regression.

19. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 17, characterized in that, A MICM diagnostic report is output based on the patient diagnosis and treatment information and the classification result.

20. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 17, characterized in that, The classifier is a trained conditional random field, and a method for constructing the trained conditional random field comprises: A diagnostic report training set is obtained, and the training set further comprises a label of whether the subject has a hematological tumor. MICM phenotyping feature training sets are obtained based on the diagnostic report training sets. Parameters of the initialized conditional random field are obtained. Sample parameters are obtained by Monte Carlo random sampling of the parameters of the initialized conditional random field. Predicted labels of the MICM phenotyping feature training sets are obtained based on the sample parameters, and a loss function of the conditional random field is calculated based on the predicted labels and actual labels of the diagnostic report. The trained conditional random field is obtained by gradually minimizing the loss function of the conditional random field through iteration until a stop condition is reached.

21. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 20, characterized in that, The stop condition comprises that the information gain of each iteration based on the sample parameters of the current conditional random field is less than a preset threshold value, and the iteration is stopped when the information gain is less than the preset threshold value. , In the formula, denotes the loss function based on the conditional random field is the label data of the i-th sample, is the parameter of the conditional random field obtained by Monte Carlo sampling at the t-th iteration in the iteration process, is the label prediction probability under the given feature and the parameter is the L1 regularization coefficient, is the L1 regularization term of the parameter ;​​ The loss function of the conditional random field is represented as: , where denotes the sample parameter of the th iteration, denotes the sample parameter of the th iteration, is the sampling step size; The sampling step For dynamic adaptation, the setting is as follows: , wherein is an initial step size, is a decay factor; denotes the label prediction probability under given feature and parameter The sample parameters of the current conditional random field are determined, all can be calculated, where , denotes the classification label, is all possible categories, and the information gain is expressed as follows: , wherein represents the information gain of the i-th sample, is the probability of the class given the feature , is the prior probability of the class .

22. The artificial intelligence feature extraction method based on lymphoid tumor MICM phenotyping according to claim 20, characterized in that, The Monte Carlo random sampling to obtain sample parameters is represented as: The calculation of the loss function of the conditional random field further comprises: after optimizing the sample parameters by using a sparse regularization strategy, the loss function of the conditional random field is calculated by using the optimized sample parameters to replace the sample parameters. = , denotes the optimized sample parameters at the t-th iteration, is the increment of the sample parameters at the t-th iteration, is the learning rate, is the gradient of the loss function of the conditional random field with respect to the parameters, is an adjustment factor calculated based on the difference between the model output and the actual classification result, denotes the sign function.

23. A computer device, comprising: The sparse regularization strategy is represented as:

24. A computer-readable storage medium, characterized in that, The device comprises a memory and a processor; the memory is used to store a computer program; and the processor executes the computer program to implement the steps of the method according to any one of claims 1-22.

25. A computer program product comprising a computer program, characterized in that, A computer program is stored thereon, and the computer program is executed by a processor to implement the steps of the method according to any one of claims 1-22. The computer program is executed by a processor to implement the steps of the method according to any one of claims 1-22.

Citation Information

Patent Citations

  • Temporomandibular joint disease automatic diagnosis method and system based on artificial intelligence

    CN119314661A

  • Diagnosis and treatment recommendation using quantum computing

    US20230207124A1