Traditional Chinese medicine health management cloud platform based on big data
Through the traditional Chinese medicine health management cloud platform based on big data, similar patients are screened, keywords are extracted, and data change sequences and target representative data are constructed, which solves the problem of low accuracy in detecting abnormal patient self-report information in the existing technology, and achieves higher detection accuracy.
Patent Information
- Application Number
- CN202510473135.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
In the prior art, when performing abnormal detection of patient self-report information, it is difficult to accurately detect errors in describing the disease and symptoms or record typos, resulting in poor detection accuracy.
The Chinese medicine health management cloud platform based on big data is adopted to obtain the initial diagnosis and treatment information of patients and self-report information, similar patients are screened, keywords are extracted and clustered, data change sequences and target representative data are constructed, and abnormal detection is performed.
It improves the accuracy of abnormal detection of patients' self-report information, can more accurately detect errors or typos in the description of the condition and symptoms, and improves the accuracy of subsequent condition analysis.
Smart Images

Figure CN119993402A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical data anomaly detection, and in particular to a traditional Chinese medicine health management cloud platform based on big data. Background Art
[0002] Traditional Chinese medicine health management is a comprehensive service that combines traditional Chinese medicine theory with modern health management methods. Therefore, in the process of traditional Chinese medicine health management, it is often necessary to manage the collected patient diagnosis and treatment information and patient self-reported information. Among them, patient diagnosis and treatment information can be the test result information obtained by the patient through the testing equipment before seeing a traditional Chinese medicine doctor. For example, patient diagnosis and treatment information can include: blood routine, urine routine and gastric ultrasound and other examination results. Patient self-reported information can represent the patient's description of his or her own symptoms. Due to differences in patients' own cognition, some patients may find it difficult to accurately describe their own symptoms, and incorrect descriptions of symptoms often cause trouble for subsequent condition analysis. Therefore, in the process of traditional Chinese medicine health management, it is often necessary to perform abnormal detection on the collected patient self-reported information.
[0003] At present, the method commonly used for data anomaly detection is to use regular expressions to detect anomalies in data. However, when using regular expressions to detect anomalies in patient self-reported information, the following technical problems often occur: Regular expressions can often detect format errors and anomalies in patient self-reported information, such as inconsistent medical record numbers and date formats. However, it is often difficult to accurately detect errors in the description of patients' symptoms or typos in the records, resulting in poor accuracy in anomaly detection of patient self-reported information. Summary of the invention
[0004] In order to solve the technical problem of poor accuracy in abnormality detection of patient self-reported information, the present invention proposes a traditional Chinese medicine health management cloud platform based on big data.
[0005] In a first aspect, the present invention provides a traditional Chinese medicine health management cloud platform based on big data, comprising a processor and a memory, wherein the processor is used to process instructions stored in the memory to implement the following steps: Obtain the initial diagnosis and treatment information and self-reported information of the patient to be tested, as well as the initial diagnosis and treatment information and all historical self-reported information of each historical patient; Based on the similarities between the initial diagnosis and treatment information of the patient to be tested and all historical patients, similar patients are screened out from all historical patients; Extract keywords from the self-description information to be detected and all historical self-description information, and cluster all the extracted keywords to obtain the target cluster; According to the keywords belonging to the same target cluster in all historical self-reported information corresponding to all similar patients, a data change sequence corresponding to the target cluster is constructed; According to the keywords belonging to the same target cluster in the self-reported information of the patients to be tested, target representative data corresponding to the target cluster are constructed; According to the similarity between the elements in the data change sequence corresponding to each target representative data and its corresponding target cluster, anomaly detection is performed on each target representative data, thereby realizing anomaly detection of the self-described information to be detected.
[0006] In combination with the first aspect above, in a possible implementation, screening similar patients from all historical patients according to the similarity between the initial diagnosis and treatment information corresponding to the patient to be detected and all historical patients includes: Based on each initial diagnosis and treatment information, a target representative matrix corresponding to each initial diagnosis and treatment information is constructed; Determine the cosine similarity between the target representative matrix corresponding to the patient to be detected and the target representative matrix corresponding to each historical patient as the target similarity between the patient to be detected and each historical patient; Historical patients whose target similarity with the patient to be detected is greater than a preset similarity threshold are screened out from all historical patients as similar patients.
[0007] In combination with the first aspect above, in a possible implementation manner, clustering all extracted keywords to obtain a target cluster includes: Take the extracted keywords as entities and construct the target knowledge graph; Determine the semantic relevance between each two keywords according to the distance between each two keywords in the target knowledge graph, wherein the distance between different keywords in the target knowledge graph is negatively correlated with the semantic relevance between them; According to the semantic relevance between different keywords, all keywords are clustered, and the clusters obtained by clustering are determined as target clusters.
[0008] In combination with the first aspect above, in a possible implementation, constructing a data change sequence corresponding to a target cluster according to keywords belonging to the same target cluster in all historical self-reported information corresponding to all similar patients includes: Filter out the historical self-reported information of the same treatment stage from all the historical self-reported information corresponding to all similar patients to form a historical self-reported information group under each treatment stage; Determine any target cluster as a marker cluster, and construct reference representative data under the marker cluster for each treatment stage based on all keywords belonging to the marker cluster in the historical self-report information group under each treatment stage; The reference representative data of all treatment stages under the marker cluster constitute a data change sequence corresponding to the marker cluster.
[0009] In combination with the first aspect, in a possible implementation, constructing reference representative data of each treatment stage under the marker cluster according to all keywords in the historical self-report information group under each treatment stage belonging to the marker cluster includes: The word vector corresponding to each keyword is obtained, and the average of the word vectors corresponding to all keywords belonging to the tag cluster in the historical self-report information group under each treatment stage is determined as the reference representative data under the tag cluster for each treatment stage.
[0010] In combination with the first aspect above, in a possible implementation, constructing target representative data corresponding to the target cluster according to keywords belonging to the same target cluster in the self-reported information corresponding to the patient to be tested includes: Determine any target cluster as a labeled cluster and obtain the word vector corresponding to each keyword; The mean of the word vectors corresponding to all the keywords in the tag cluster in the self-description information to be detected is determined as the target representative data corresponding to the tag cluster.
[0011] In combination with the first aspect, in a possible implementation, performing anomaly detection on each target representative data according to similarities between elements in a data change sequence corresponding to each target representative data and its corresponding target cluster includes: Determine the target treatment stage corresponding to each target representative data according to the cosine similarity between each target representative data and the reference representative data in the data change sequence corresponding to the corresponding target cluster; The mean of the start time of the target treatment phase corresponding to all target representative data is determined as the treatment representative time; According to the difference between the treatment representative time and the start time of the target treatment stage corresponding to each target representative data, abnormality detection is performed on each target representative data.
[0012] In combination with the first aspect above, in a possible implementation, determining the target treatment stage corresponding to each target representative data according to the cosine similarity between each target representative data and the reference representative data in the data change sequence corresponding to the target cluster corresponding to the target representative data includes: Determine any target representative data as the marked representative data, determine the target cluster corresponding to the marked representative data as the temporary cluster, and determine each reference representative data in the data change sequence corresponding to the temporary cluster as the temporary representative data; Filter out temporary representative data having the largest cosine similarity with the marked representative data from all temporary representative data as comparison representative data; The treatment stage corresponding to the comparison representative data is set as the target treatment stage corresponding to the marked representative data.
[0013] In combination with the first aspect, in a possible implementation, performing abnormality detection on each target representative data according to the difference between the treatment representative moment and the start moment of the target treatment stage corresponding to each target representative data includes: Normalizing the absolute value of the difference between the treatment representative moment and the start moment of the target treatment phase corresponding to each target representative data to obtain the abnormal deviation degree corresponding to each target representative data; If the degree of abnormal deviation corresponding to the target representative data is greater than a preset abnormal threshold, the target representative data is determined to be abnormal.
[0014] In combination with the first aspect above, in a possible implementation manner, a method for detecting anomalies in self-described information to be detected includes: If there is an abnormality in the target representative data, the self-described information to be detected is determined to be abnormal.
[0015] The present invention has the following beneficial effects: The cloud platform for traditional Chinese medicine health management based on big data of the present invention realizes abnormal detection of self-reported information by processing diagnosis and treatment information and self-reported information, solves the technical problem of poor accuracy of abnormal detection of patient self-reported information, and improves the accuracy of abnormal detection of patient self-reported information. Compared with abnormal detection through regular expressions, when the present invention performs abnormal detection on the self-reported information to be detected corresponding to the patient to be detected, the symptom conditions of patients with similar conditions often have a certain correlation. Therefore, similar patients with similar diagnosis and treatment conditions as the patient to be detected are screened out from all historical patients, and the data change sequence representing the overall conditions of all similar patients and the target representative data representing the overall conditions of the patient to be detected under the same target cluster are quantified. By analyzing the similarity between the elements in the data change sequence corresponding to each target representative data and its corresponding target cluster, abnormal detection of self-reported information to be detected is realized, thereby improving the accuracy of abnormal detection of self-reported information to be detected. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0017] Figure 1 A flowchart of the method for implementing the traditional Chinese medicine health management cloud platform based on big data of the present invention; Figure 2 A flowchart of a method for obtaining a target treatment stage of the present invention; Figure 3 A flow chart representing data anomaly detection for the purposes of the present invention. DETAILED DESCRIPTION
[0018] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation methods, structures, features and effects of the technical solutions proposed by the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.
[0019] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0020] refer to Figure 1 , shows a flowchart of a method for implementing a traditional Chinese medicine health management cloud platform based on big data according to the present invention. Specifically, the traditional Chinese medicine health management cloud platform based on big data includes a processor and a memory, and the processor is used to process instructions stored in the memory to implement the following steps: Step S1, obtaining the initial diagnosis and treatment information and self-reported information of the patient to be tested, as well as the initial diagnosis and treatment information and all historical self-reported information corresponding to each historical patient.
[0021] Among them, the patient to be detected may be a patient to be tested for abnormal self-report information. The initial diagnosis and treatment information may characterize the condition of the patient when he first visits the doctor for TCM treatment. The initial diagnosis and treatment information may include: the examination results that assist TCM judgment during the first visit of the patient for TCM treatment, and the diagnosis results of TCM during the first visit. For example, the initial diagnosis and treatment information may include but is not limited to: blood routine test results, urine routine test results, gastric ultrasound test results and TCM diagnosis results. It should be noted that when performing TCM treatment, TCM often observes the test results obtained by the patient through the corresponding equipment to assist in treatment judgment. The test results obtained by the equipment inspection may be but are not limited to: blood routine test results, urine routine test results and gastric ultrasound test results. The self-report information to be detected may be the information obtained by the patient to be detected describing his or her own symptoms. The historical patient may be a patient who has received TCM treatment in the past. The historical self-report information may be the information obtained by the historical patient describing his or her own symptoms.
[0022] As an example, the blood routine, urine routine and gastric ultrasound and other examination results that assist in the TCM diagnosis of the patient to be tested during the first visit to the TCM treatment process can be collected, and the TCM diagnosis results of the patient to be tested during the first visit to the TCM treatment process can be recorded, and these examination results and TCM diagnosis results can be used to form the initial diagnosis and treatment information corresponding to the patient to be tested. The most recent description of the patient's own symptoms during the TCM treatment process can be collected to form the self-reported information corresponding to the patient to be tested. Similarly, the initial diagnosis and treatment information and all historical self-reported information corresponding to each historical patient can be collected.
[0023] Step S2, based on the similarity between the initial diagnosis and treatment information corresponding to the patient to be detected and all historical patients, similar patients are screened out from all historical patients.
[0024] It should be noted that the more similar patients are screened, the better the effect of subsequent abnormality detection on the self-reported information will be.
[0025] As an example, this step may include the following steps: The first step is to construct a target representative matrix corresponding to each initial diagnosis and treatment information based on each initial diagnosis and treatment information, which may include the following sub-steps: In the first sub-step, for the examination results with numerical data in the initial diagnosis and treatment information, a vector consisting of all numerical data representing different indicators in each examination result can be recorded as an initial examination vector.
[0026] For example, a vector consisting of all numerical data representing different indicators in the urine routine examination results can be used as an initial examination vector.
[0027] In the second sub-step, for text information in the initial diagnosis and treatment information, such as TCM diagnosis results, a keyword library can be manually constructed based on expert suggestions to obtain different unique hot vectors corresponding to different keywords; then the words in the TCM diagnosis results are matched with the keywords in the constructed keyword library to obtain a matching keyword sequence, and the unique hot vectors of all keywords in the matching keyword sequence are superimposed to obtain a feature vector of the TCM diagnosis results.
[0028] In the third sub-step, a matrix consisting of all initial examination vectors and characteristic vectors of TCM diagnosis results corresponding to each initial diagnosis and treatment information is used as a target representative matrix corresponding to each initial diagnosis and treatment information.
[0029] For example, each initial inspection vector can be used as a row of the matrix, and the characteristic vector of the TCM diagnosis result can be used as a row of the matrix. Taking the longest row vector in the matrix as the reference, the other row vectors are padded with 0 to make the length of all row vectors in the matrix the same, and the final matrix is recorded as the target representative matrix. Each row in the matrix can be called a row vector.
[0030] It should be noted that the target representative matrix corresponding to the initial diagnosis and treatment information can represent the patient's disease symptoms corresponding to the initial diagnosis and treatment information.
[0031] In the second step, the cosine similarity between the target representative matrix corresponding to the patient to be detected and the target representative matrix corresponding to each historical patient is determined as the target similarity between the patient to be detected and each historical patient.
[0032] Among them, the target representative matrix corresponding to the patient to be detected is the target representative matrix corresponding to the initial diagnosis and treatment information of the patient to be detected. The target representative matrix corresponding to the historical patient is the target representative matrix corresponding to the initial diagnosis and treatment information of the historical patient.
[0033] The third step is to select historical patients whose target similarity with the above-mentioned patient to be detected is greater than a preset similarity threshold from all historical patients as similar patients.
[0034] The preset similarity threshold may be a preset maximum similarity when two patients are considered to be dissimilar. For example, the preset similarity threshold may be 0.7.
[0035] It should be noted that the symptoms of similar patients during their initial TCM treatment are often similar to those of the patients to be tested during their initial TCM treatment.
[0036] Step S3, extracting keywords from the self-description information to be detected and all historical self-description information, and clustering all extracted keywords to obtain a target cluster.
[0037] As an example, this step may include the following steps: The first step is to extract keywords from the self-description information to be detected and all historical self-description information.
[0038] For example, a keyword library can be manually set, and the words in the self-description information to be detected and all historical self-description information can be matched one by one with the keywords in the manually set keyword library, and the words that can be matched in the self-description information to be detected and all historical self-description information can be used as the extracted keywords.
[0039] The second step is to construct the target knowledge graph using the extracted keywords as entities.
[0040] For example, the keywords extracted from the self-description information to be detected and all historical self-description information can be used as entities to construct a knowledge graph, and the constructed knowledge graph can be used as the target knowledge graph.
[0041] It should be noted that relevant descriptions about the disease in the self-reported information to be tested and all historical self-reported information can be collected to identify the entities described. Specifically, the extracted keywords can be taken as entities, and relationships can be extracted to build a knowledge graph network. The knowledge graph built at this time is the target knowledge graph.
[0042] The third step is to determine the semantic relevance between every two keywords based on the distance between them in the target knowledge graph.
[0043] Among them, the distance between two keywords in the target knowledge graph is the shortest path length between the two keywords in the target knowledge graph. The distance between different keywords in the target knowledge graph can be negatively correlated with the semantic relevance between them.
[0044] It should be noted that in actual situations, the correlation between the unique hot vectors corresponding to different keywords is often 0, but in fact there may be a semantic connection between them. For example, if the two keywords are stomachache and acid reflux, the correlation between the unique hot vectors corresponding to stomachache and acid reflux is often 0, but in fact stomachache and acid reflux both describe stomach conditions and have a semantic connection, and their distance in the knowledge graph is relatively small. Therefore, when the distance between two keywords in the target knowledge graph is smaller, it often means that the semantic correlation between the two keywords is greater.
[0045] For example, the formula for determining the semantic relevance between two keywords may be: ; It is i Keywords and j The semantic relevance between keywords. i andj It is the serial number of different keywords. It is i Keywords and j The distance between keywords in the target knowledge graph.
[0046] The fourth step is to cluster all keywords according to the semantic relevance between different keywords, and determine the clusters obtained by clustering as target clusters.
[0047] For example, all keywords may be hierarchically clustered according to the semantic relevance between different keywords, and the clusters obtained by clustering may be recorded as target clusters.
[0048] It should be noted that, when the semantic relevance between two keywords is greater, the two keywords are often classified into the same cluster.
[0049] Step S4, constructing a data change sequence corresponding to the target cluster based on the keywords belonging to the same target cluster in all historical self-reported information corresponding to all similar patients.
[0050] As an example, this step may include the following steps: In the first step, the historical self-reported information corresponding to the same treatment stage is screened out from all the historical self-reported information corresponding to all similar patients to form a historical self-reported information group under each treatment stage.
[0051] Among them, the treatment stage can be a pre-set treatment time period, and its corresponding duration can be a pre-set duration. For example, if the duration corresponding to the treatment stage is 1 day, then the first day of the patient's TCM treatment can be a treatment stage, the second day of the patient's TCM treatment can be a treatment stage, the third day of the patient's TCM treatment can be a treatment stage, and so on, all treatment stages corresponding to all similar patients can be obtained. The method for obtaining the historical self-reported information group under the treatment stage can be: take any treatment stage as a marked treatment stage, and use the historical self-reported information collected in the marked treatment stage from all the historical self-reported information corresponding to all similar patients to form the historical self-reported information group under the marked treatment stage.
[0052] It should be noted that the more treatment stages all similar patients correspond to, the more accurate the results of subsequent self-reported information anomaly detection based on all treatment stages all similar patients correspond to.
[0053] In the second step, any target cluster is determined as a marker cluster, and reference representative data of each treatment stage under the above marker cluster is constructed based on all the keywords belonging to the above marker cluster in the historical self-report information group under each treatment stage.
[0054] For example, Word2Vec can be used to obtain the word vector corresponding to each keyword, and the mean of the word vectors corresponding to all keywords belonging to the above-mentioned tag cluster in the historical self-report information group at each treatment stage can be determined as the reference representative data under the above-mentioned tag cluster at each treatment stage. Among them, Word2Vec is a word embedding model based on a neural network, which can map vocabulary into real number vectors, recorded as word vectors.
[0055] The third step is to use the reference representative data of all treatment stages under the above-mentioned marker clusters to form a data change sequence corresponding to the above-mentioned marker clusters.
[0056] Step S5, constructing target representative data corresponding to the target cluster according to the keywords belonging to the same target cluster in the self-reported information corresponding to the patient to be tested.
[0057] As an example, this step may include the following steps: In the first step, any target cluster is identified as a labeled cluster, and the word vector corresponding to each keyword is obtained through Word2Vec.
[0058] In the second step, the mean of the word vectors corresponding to all the keywords in the above-mentioned self-described information to be detected and belonging to the above-mentioned tag cluster is determined as the target representative data corresponding to the above-mentioned tag cluster.
[0059] Step S6, performing anomaly detection on each target representative data according to the similarity between each target representative data and the elements in the data change sequence corresponding to its corresponding target cluster, thereby realizing anomaly detection of the self-described information to be detected.
[0060] As an example, this step may include the following steps: In the first step, the target treatment stage corresponding to each target representative data is determined based on the cosine similarity between each target representative data and the reference representative data in the data change sequence corresponding to its corresponding target cluster.
[0061] like Figure 2 As shown, determining the target treatment stage corresponding to each target representative data may include the following steps: Step 201, any target representative data is determined as the marked representative data, and the target cluster corresponding to the marked representative data is determined as the temporary cluster, and each reference representative data in the data change sequence corresponding to the temporary cluster is determined as the temporary representative data.
[0062] Step 202 , screening out temporary representative data having the largest cosine similarity with the marked representative data from all temporary representative data as comparison representative data.
[0063] Step 203, setting the treatment stage corresponding to the comparison representative data as the target treatment stage corresponding to the marked representative data.
[0064] In the second step, the mean of the start time of the target treatment phase corresponding to all target representative data is determined as the treatment representative time.
[0065] In the third step, anomaly detection is performed on each target representative data according to the difference between the treatment representative moment and the start moment of the target treatment phase corresponding to each target representative data.
[0066] like Figure 3 As shown, performing anomaly detection on each target representative data may include the following steps: Step 301, normalize the absolute value of the difference between the treatment representative time and the start time of the target treatment stage corresponding to each target representative data, and obtain the abnormal deviation degree corresponding to each target representative data.
[0067] It should be noted that the descriptions of most of the symptoms of the patient's illness are often relatively accurate, that is, the accurate symptom description data in the self-report information is often more than the abnormal description data. Therefore, the representative treatment time can often represent the start time of the treatment stage that the patient to be tested is currently in. At this time, if the start time of the target treatment stage corresponding to the target representative data deviates from the representative treatment time, it often means that the description of the symptoms represented by the target representative data is more likely to be an abnormal description, and the patient or the Chinese medicine practitioner needs to be reminded to conduct a focused check.
[0068] Step 302: If the abnormal deviation degree corresponding to the target representative data is greater than a preset abnormal threshold, the target representative data is determined to be abnormal.
[0069] The preset abnormal threshold may be a pre-set threshold, which may be 0.6.
[0070] It should be noted that if there is no reference representative data in the data change sequence corresponding to the temporary cluster, the abnormal deviation degree corresponding to the target representative data may be set to 1.
[0071] The fourth step is to determine that the self-reported information to be detected is abnormal if there is an abnormality in the target representative data.
[0072] It should be noted that when the self-reported information to be tested is abnormal, it is often necessary to remind the patient or the Chinese medicine practitioner to focus on checking the abnormal target representative data in the self-reported information to be tested.
[0073] In summary, compared with anomaly detection through regular expressions, when the present invention performs anomaly detection on the self-reported information to be detected corresponding to the patient to be detected, the symptoms of patients with similar conditions often have a certain correlation. Therefore, similar patients with similar diagnosis and treatment conditions to the patient to be detected are screened out from all historical patients, and the data change sequence representing the overall condition of all similar patients and the target representative data representing the overall condition of the patient to be detected under the same target cluster are quantified. By analyzing the similarity between the elements in the data change sequence corresponding to each target representative data and its corresponding target cluster, anomaly detection of the self-reported information to be detected is achieved, thereby improving the accuracy of anomaly detection of the self-reported information to be detected.
[0074] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features can be replaced by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A traditional Chinese medicine health management cloud platform based on big data, characterized in that: The method comprises a processor and a memory, wherein the processor is configured to process instructions stored in the memory to implement the following steps: Obtain the initial diagnosis and treatment information and self-reported information of the patient to be tested, as well as the initial diagnosis and treatment information and all historical self-reported information of each historical patient; Based on the similarities between the initial diagnosis and treatment information of the patient to be tested and all historical patients, similar patients are screened out from all historical patients; Extract keywords from the self-description information to be detected and all historical self-description information, and cluster all the extracted keywords to obtain the target cluster; According to the keywords belonging to the same target cluster in all historical self-reported information corresponding to all similar patients, a data change sequence corresponding to the target cluster is constructed; According to the keywords belonging to the same target cluster in the self-reported information of the patients to be tested, target representative data corresponding to the target cluster are constructed; According to the similarity between the elements in the data change sequence corresponding to each target representative data and its corresponding target cluster, anomaly detection is performed on each target representative data, thereby realizing anomaly detection of the self-described information to be detected.
2. According to a big data-based traditional Chinese medicine health management cloud platform according to claim 1, it is characterized in that: The method of screening similar patients from all historical patients based on the similarities between the initial diagnosis and treatment information corresponding to the patient to be detected and all historical patients includes: Based on each initial diagnosis and treatment information, a target representative matrix corresponding to each initial diagnosis and treatment information is constructed; Determine the cosine similarity between the target representative matrix corresponding to the patient to be detected and the target representative matrix corresponding to each historical patient as the target similarity between the patient to be detected and each historical patient; Historical patients whose target similarity with the patient to be detected is greater than a preset similarity threshold are screened out from all historical patients as similar patients.
3. According to a big data-based traditional Chinese medicine health management cloud platform according to claim 1, it is characterized in that: The method clusters all the extracted keywords to obtain a target cluster, including: Take the extracted keywords as entities and construct the target knowledge graph; Determine the semantic relevance between each two keywords according to the distance between each two keywords in the target knowledge graph, wherein the distance between different keywords in the target knowledge graph is negatively correlated with the semantic relevance between them; According to the semantic relevance between different keywords, all keywords are clustered, and the clusters obtained by clustering are determined as target clusters.
4. According to a big data-based traditional Chinese medicine health management cloud platform according to claim 1, it is characterized in that: The step of constructing a data change sequence corresponding to the target cluster based on keywords belonging to the same target cluster in all historical self-reported information corresponding to all similar patients includes: Filter out the historical self-reported information of the same treatment stage from all the historical self-reported information corresponding to all similar patients to form a historical self-reported information group under each treatment stage; Determine any target cluster as a marker cluster, and construct reference representative data under the marker cluster for each treatment stage based on all keywords belonging to the marker cluster in the historical self-report information group under each treatment stage; The reference representative data of all treatment stages under the marker cluster constitute a data change sequence corresponding to the marker cluster.
5. According to a big data-based traditional Chinese medicine health management cloud platform according to claim 4, it is characterized in that: The step of constructing reference representative data of each treatment stage under the marker cluster according to all keywords in the historical self-report information group under each treatment stage that belong to the marker cluster includes: The word vector corresponding to each keyword is obtained, and the average of the word vectors corresponding to all keywords belonging to the tag cluster in the historical self-report information group under each treatment stage is determined as the reference representative data under the tag cluster for each treatment stage.
6. A traditional Chinese medicine health management cloud platform based on big data according to claim 1, characterized in that: The target representative data corresponding to the target cluster is constructed according to the keywords belonging to the same target cluster in the self-reported information corresponding to the patient to be tested, including: Determine any target cluster as a labeled cluster and obtain the word vector corresponding to each keyword; The mean of the word vectors corresponding to all the keywords in the tag cluster in the self-description information to be detected is determined as the target representative data corresponding to the tag cluster.
7. A traditional Chinese medicine health management cloud platform based on big data according to claim 4, characterized in that: The method of performing anomaly detection on each target representative data according to the similarity between the elements in the data change sequence corresponding to each target representative data and its corresponding target cluster comprises: Determine the target treatment stage corresponding to each target representative data according to the cosine similarity between each target representative data and the reference representative data in the data change sequence corresponding to the corresponding target cluster; The mean of the start time of the target treatment phase corresponding to all target representative data is determined as the treatment representative time; According to the difference between the treatment representative time and the start time of the target treatment stage corresponding to each target representative data, abnormality detection is performed on each target representative data.
8. A traditional Chinese medicine health management cloud platform based on big data according to claim 7, characterized in that: The step of determining the target treatment stage corresponding to each target representative data according to the cosine similarity between each target representative data and the reference representative data in the data change sequence corresponding to the target cluster corresponding to the target representative data, comprises: Determine any target representative data as the marked representative data, determine the target cluster corresponding to the marked representative data as the temporary cluster, and determine each reference representative data in the data change sequence corresponding to the temporary cluster as the temporary representative data; Filter out temporary representative data having the largest cosine similarity with the marked representative data from all temporary representative data as comparison representative data; The treatment stage corresponding to the comparison representative data is set as the target treatment stage corresponding to the marked representative data.
9. A traditional Chinese medicine health management cloud platform based on big data according to claim 7, characterized in that: The performing abnormality detection on each target representative data according to the difference between the treatment representative time and the start time of the target treatment stage corresponding to each target representative data comprises: Normalizing the absolute value of the difference between the treatment representative moment and the start moment of the target treatment phase corresponding to each target representative data to obtain the abnormal deviation degree corresponding to each target representative data; If the degree of abnormal deviation corresponding to the target representative data is greater than a preset abnormal threshold, the target representative data is determined to be abnormal.
10. A traditional Chinese medicine health management cloud platform based on big data according to claim 9, characterized in that: The anomaly detection method of the self-described information to be detected includes: If there is an abnormality in the target representative data, the self-described information to be detected is determined to be abnormal.
Citation Information
Patent Citations
Medical record generation method, terminal device and computer readable storage medium
CN110706774A
Intelligent information pushing method and system based on patient medical records
CN117891899A
Patient holographic health record display method and system based on time axis
CN118155794A
Pediatric registration follow-up data processing method and system based on home page of hospitalization medical record
CN118748068A
Medical image report generation method, device and equipment and computer readable storage medium
CN118841125A
Cited By
Patient input and output monitoring method and system based on flexible electronics and multi-parameter fusion
CN120766989A
Multi-modal medical data intelligent association analysis system based on deep learning
CN121034512A