Data Abnormality Detection Method and Device

By receiving clinical data sets and test knowledge sets on the server, generating clinical test data sets and training a similarity sorting model, the problem of low accuracy of test inspections in the existing technology is solved, and more efficient data abnormality detection is achieved, improving the accuracy and user experience of test inspections.

CN111755086BActive Publication Date: 2025-05-27PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010598054.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-28
Publication Date
2025-05-27
Estimated Expiration
2040-06-28

AI Technical Summary

Technical Problem

The prior art lacks methods for doctors to perform abnormal detection in clinical testing, resulting in low accuracy of testing and low user experience.

Method used

Provide a data abnormality detection method. By receiving the clinical data set and the test knowledge set, a clinical test data set is generated, and it is trained as training data of the preset similarity sorting model to obtain a trained target similarity sorting model. Then, the data to be tested is received, and it is input into the trained model to obtain multiple test knowledge. Based on this knowledge, whether the clinical test conclusion corresponding to the data to be tested is in an abnormal state.

Benefits of technology

Through this method, the accuracy of data abnormality detection can be improved, the situation of missed or multi-checked detection can be reduced, and the accuracy and efficiency of inspection and inspection can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111755086B_ABST
    Figure CN111755086B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to data processing, and provide a data anomaly detection method and device, which are applied to a server. The method includes: receiving a clinical data set and an inspection knowledge set, generating a clinical inspection data set based on the clinical data set and the inspection knowledge set, using the clinical inspection data set as training data for a preset similarity ranking model to perform a training operation, and obtaining a trained target similarity ranking model; receiving data to be inspected, inputting the data to be inspected into the trained target similarity ranking model, obtaining multiple inspection knowledge, and determining whether the clinical inspection conclusion corresponding to the data to be inspected is in an abnormal state based on the multiple inspection knowledge. The clinical data set, inspection knowledge set, and clinical inspection data set in the present application can be stored in a blockchain. Using the embodiments of the present application is beneficial to improving the efficiency of data anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a method and device for detecting data anomalies. Background Art

[0002] Inspection and examination is an important process in clinical practice. It mainly uses various tools in the laboratory to evaluate the health status and physiological functions of patients and assist in diagnosis and treatment in clinical medicine. Usually, inspection and examination are given by doctors based on the patient's chief complaint and their own clinical experience. Therefore, the entire inspection and examination process is highly subjective, and due to differences in doctors' clinical experience, the results obtained from inspection and examination are also different, easily leading to missed or over-detected cases. Missed detection results in the absence of key clinical indicators, and over-detection leads to a long cycle of the inspection and examination process. Therefore, there is still a lack of a method for detecting anomalies in doctors' inspection and examination, resulting in low accuracy of inspection and examination and poor user experience. Summary of the Invention

[0003] The embodiments of this application provide a method and device for detecting data anomalies, which are beneficial to improving the accuracy of data anomaly detection.

[0004] In a first aspect of the embodiments of this application, a method for detecting data anomalies is provided, which is applied to a server and includes:

[0005] Receiving a clinical data set and an inspection knowledge set, and generating a clinical inspection data set based on the clinical data set and the inspection knowledge set;

[0006] Using the clinical inspection data set as training data for a preset similarity ranking model to perform a training operation, and obtaining a trained target similarity ranking model;

[0007] Receiving data to be inspected, inputting the data to be inspected into the trained target similarity ranking model, obtaining multiple inspection knowledges, and judging whether the clinical inspection conclusion corresponding to the data to be inspected is in an abnormal state based on the multiple inspection knowledges.

[0008] In a second aspect of the embodiments of this application, a device for detecting data anomalies is provided, which is applied to a server. The device includes: a receiving unit, a training unit, and a judging unit, where

[0009] The receiving unit is configured to receive a clinical data set and an inspection knowledge set, and generate a clinical inspection data set based on the clinical data set and the inspection knowledge set;

[0010] The training unit is configured to use the clinical inspection data set as training data for a preset similarity ranking model to perform a training operation, and obtain a trained target similarity ranking model;

[0011] The determination unit is configured to receive the data to be inspected, input the data to be inspected into the trained target similarity ranking model to obtain multiple inspection knowledge, and determine whether the clinical inspection conclusion corresponding to the data to be inspected is in an abnormal state based on the multiple inspection knowledge.

[0012] In a third aspect of the embodiments of the present application, a server is provided. The server includes a processor, an input device, an output device, and a memory. The processor, the input device, the output device, and the memory are interconnected. Among them, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method described in the first aspect of the embodiments of the present application.

[0013] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program for electronic data exchange. The computer program causes a computer to execute some or all of the steps described in the first aspect of the embodiments of the present application.

[0014] In a fifth aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to cause a computer to execute some or all of the steps described in the first aspect of the embodiments of the present application. The computer program product can be a software installation package.

[0015] Implementing the embodiments of the present application has at least the following beneficial effects:

[0016] Through the embodiments of the present application, applied to a server, the method includes: receiving a clinical data set and an inspection knowledge set, generating a clinical inspection data set based on the clinical data set and the inspection knowledge set, using the clinical inspection data set as the training data of a preset similarity ranking model to perform a training operation to obtain a trained target similarity ranking model; receiving the data to be inspected, inputting the data to be inspected into the trained target similarity ranking model to obtain multiple inspection knowledge, and determining whether the clinical inspection conclusion corresponding to the data to be inspected is in an abnormal state based on the multiple inspection knowledge; in this way, multiple inspection knowledge corresponding to the data to be inspected can be obtained based on the trained target similarity ranking model, and the multiple inspection knowledge is compared with the clinical inspection conclusion corresponding to the data to be inspected to determine whether the clinical inspection conclusion is in an abnormal state, which is beneficial to improving the efficiency of data anomaly detection. Description of the Drawings

[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0018] Figure 1A This is a schematic diagram of the architecture of a data anomaly detection system provided by an embodiment of the present application;

[0019] Figure 1B This is a schematic diagram of the network structure of a preset similarity ranking model provided by an embodiment of the present application;

[0020] Figure 1C This is a schematic diagram of the network process of a data anomaly detection method provided by an embodiment of the present application;

[0021] Figure 2 This is a schematic diagram of the process of a data anomaly detection method provided by an embodiment of the present application;

[0022] Figure 3 This is a schematic diagram of the process of a data anomaly detection method provided by an embodiment of the present application;

[0023] Figure 4 This is a schematic diagram of the structure of a server provided by an embodiment of the present application;

[0024] Figure 5 This is a schematic diagram of the structure of a data anomaly detection device provided by an embodiment of the present application. Detailed implementation manners

[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.

[0026] The terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0027] References to "embodiments" in this application mean that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments.

[0028] To better understand the embodiments of this application, the methods applying the embodiments of this application will be introduced below.

[0029] The server mentioned in the embodiments of this application may include, but is not limited to, a background server, a component server, a cloud server, a data distribution system server, or a data distribution software server, etc. The above are only examples and not an exhaustive list, including but not limited to the above devices.

[0030] Please refer to Figure 1A , Figure 1A which is a schematic flowchart of a data anomaly detection method provided by the embodiments of this application and is applied to a server. The above method includes the following steps:

[0031] 101. Receive a clinical data set and a test knowledge set, and generate a clinical test data set based on the clinical data set and the test knowledge set.

[0032] Among them, clinically, testing is an important process, which mainly uses various tools in the laboratory to evaluate the health status and physiological functions of patients and assist in diagnosis and treatment in clinical medicine. In the embodiments of this application, the above clinical data set may include multiple groups of clinical data, the above test knowledge set may include multiple groups of test knowledge, the above clinical test data set may include multiple clinical test data, and any one of the multiple clinical test data may include a clinical data and a test knowledge. In practical applications, the above clinical data set can be obtained based on historical clinical data and may include the clinical symptoms of patients and the actual test results prescribed by doctors, etc. For example, the diseases of patients, blood routine data, urine routine data, electrocardiogram, etc., which are not limited herein; the above test knowledge may include various symptoms and their corresponding common tests, such as the methods and guiding suggestions for diagnosing acute respiratory tract infection viruses and bacteria, the preparation and staining of blood smears in blood tests, or the extraction volume and other test standards, etc.

[0033] Optionally, to ensure the security of users' medical data, the above-mentioned clinical data set, test knowledge set, and clinical test data set can be stored in the nodes of the blockchain. It should be noted that the blockchain referred to in the embodiments of this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, essentially a decentralized database, is a series of data blocks generated by using cryptographic methods. Each data block contains information on a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0034] In a possible example, the above step 101 of generating a clinical test data set based on the clinical data set and test knowledge set includes:

[0035] 11. Obtain multiple groups of clinical data from the clinical data set, where any one group of the multiple groups of clinical data includes: at least one clinical symptom and a clinical test conclusion corresponding to the at least one clinical symptom;

[0036] 12. Obtain multiple groups of test knowledge from the test knowledge set, where any one group of the multiple groups of test knowledge includes: a disease and at least one symptom test knowledge corresponding to the disease;

[0037] 13. Generate the clinical test data set based on the multiple groups of clinical data and the multiple groups of test knowledge.

[0038] Among them, the above-mentioned diseases may include at least one of the following: fever, chills, cough, runny nose, swollen and red throat, etc., which are not limited here.

[0039] Among them, the clinical test data set generated from the clinical data set and test knowledge set can be used as the training set for subsequent model training operations, which is beneficial to the promotion of subsequent judgment operations on clinical test conclusions.

[0040] In a possible example, the above step 13 of generating the clinical test data set based on the multiple groups of clinical data and the multiple groups of test knowledge may include the following steps:

[0041] 131. Obtain any one group of clinical data from the multiple groups of clinical data as the target clinical data, and calculate the Jaccard similarities corresponding to the multiple groups of test knowledge for the target clinical data;

[0042] 132. Obtain the test knowledge corresponding to the maximum value among the multiple Jaccard similarities as the target test knowledge, generate the mapping relationship between the target clinical data and the target test knowledge as the clinical test data, repeat the above steps to obtain multiple clinical test data corresponding to the multiple groups of clinical data, and combine the multiple clinical test data to obtain the clinical test data set.

[0043] Among them, the above Jaccard similarity is an index used to evaluate the similarity between two data sets of test knowledge and clinical data.

[0044] Among them, since the data volumes of the above multiple groups of clinical data and multiple groups of test knowledge are large and diverse, in order to reduce the workload of subsequent model training and improve the accuracy of subsequent clinical test conclusions, the above multiple groups of clinical data and multiple clinical test data can be evaluated to calculate multiple Jaccard similarities, and the above multiple groups of test knowledge and multiple groups of clinical data can be screened through the Jaccard similarities to obtain the clinical test data set for subsequent model training.

[0045] In specific implementation, the Jaccard similarity between any group of clinical data in the multiple groups of clinical data and each group of test knowledge can be calculated to obtain multiple Jaccard similarities, and the test knowledge corresponding to the maximum value among the multiple Jaccard similarities is selected as the target test knowledge; repeating the above steps, multiple clinical test data corresponding to the multiple groups of clinical data can be obtained, and the above multiple clinical test data form the above clinical test data set. In this way, the clinical test data set for subsequent model training can be obtained.

[0046] In a possible example, step 131 above, calculating the multiple Jaccard similarities corresponding to the target clinical data and the multiple groups of test knowledge, may include the following steps:

[0047] Obtain a preset Jaccard calculation formula, obtain any group of test knowledge in the multiple groups of test knowledge, use the target clinical data and the any group of test knowledge as the input of the preset Jaccard calculation formula to obtain the Jaccard similarity between the target clinical data and the any group of test knowledge, and repeat the above steps to obtain multiple Jaccard similarities.

[0048] Among them, the above preset Jaccard calculation formula can be set by the user or defaulted by the system, which is not limited here.

[0049] For example, the above-mentioned preset Jaccard calculation formula can be set as: J(A,B) = (|A∩B|) / (|A∪B|), where A represents the target clinical data, B represents any set of test knowledge, and J(A,B) represents the Jaccard similarity between the target clinical data A and any set of test knowledge B; that is to say, the quotient of the number of identical tests and examinations of two sets of data (target clinical data A and test knowledge B) and the number of non-repeating elements of the two sets of data (target clinical data A and test knowledge B) can be obtained; thus, multiple Jaccard similarities between the target clinical data and multiple sets of test knowledge can be obtained, and by repeating the above steps, multiple sets of test knowledge corresponding to multiple sets of clinical data can be obtained to obtain a clinical test data set.

[0050] 102. Use the clinical test data set as the training data of the preset similarity ranking model to perform a training operation to obtain a trained target similarity ranking model.

[0051] Among them, the above-mentioned preset similarity ranking model can be set by the user or default by the system, which is not limited here; for example, the above-mentioned preset similarity ranking model can be a Deep Neural Network (DNN).

[0052] Among them, in the embodiment of the present application, the above-mentioned clinical test data set can be used as training data for training the preset similarity ranking model, so that the trained target similarity ranking model may optimize the item most similar to the input in the general direction, which is beneficial to the need to optimize multiple sets of test knowledge corresponding to multiple sets of the most similar clinical data in the embodiment of the present application.

[0053] In a possible example, step 102 above, using the clinical test data set as the training data of the preset similarity ranking model to perform a training operation to obtain a trained target similarity ranking model, may include the following steps:

[0054] 21. Use the multiple sets of clinical data as the input of the first attention layer of the preset similarity ranking model;

[0055] 22. Use the multiple sets of test knowledge as the input of the second attention layer of the preset similarity ranking model;

[0056] 23. Obtain the clinical feature proportion vector of the multiple sets of clinical data from the first attention layer, and obtain the knowledge feature proportion vector of the multiple sets of test knowledge from the second attention layer;

[0057] 24. Input the clinical feature proportion vector and the knowledge feature proportion vector into the preset similarity ranking model to obtain the embedding vectors corresponding to the multiple groups of clinical data and the multiple groups of test knowledge, where the embedding vector represents the similarity between the multiple groups of clinical data and the multiple groups of test knowledge;

[0058] 25. Train the preset similarity ranking model based on the embedding vector to obtain the trained target similarity ranking model.

[0059] Among them, as Figure 1B shown, Figure 1B is a schematic diagram of the network structure of a preset similarity ranking model provided by an embodiment of the present application. The above preset similarity ranking model may include a DNN network model, and the above preset similarity ranking model may further include a first attention layer and a second attention layer.

[0060] In a specific implementation, the server may use multiple groups of clinical data (for example, X1 = (A1i, B1i, E1i, G1i, F1i)) as the input of the first attention layer of the above preset similarity ranking model. The first attention layer includes a first vector with the same dimension as the multiple groups of clinical data, and the first attention layer adjusts the values of the vector elements in the first vector through model training. The value of any vector element in the first vector represents the importance degree of the DNN-based similarity ranking model for any vector element. The higher the value of any vector element, the greater the importance degree of the DNN-based similarity ranking model for any vector element, and vice versa.

[0061] Further, multiple groups of test knowledge (for example, X2 = (A2i, B2i, E2i, G2i, F2i)) may be used as the input of the second attention layer of the preset similarity ranking model. The second attention layer includes a second vector with the same dimension as the multiple groups of test knowledge, and the second attention layer adjusts the values of the vector elements in the second vector through model training. Obtain the clinical feature proportion vector of multiple groups of clinical data from the first attention layer, obtain the knowledge feature proportion vector of multiple groups of test knowledge from the second attention layer, input the clinical feature proportion vector and the knowledge feature proportion vector into the above DNN network respectively, and perform model training operations to obtain the embedding vectors of multiple groups of clinical data and multiple groups of test knowledge, where the embedding vector represents the similarity between multiple groups of clinical data and multiple groups of test knowledge.

[0062] Furthermore, the preset similarity ranking model can be updated based on the embedding vectors to obtain the finally trained target similarity ranking model. Specifically, the cosine loss function of the embedding vectors can be obtained, and the distribution of the embedding vectors can be constrained by this cosine loss function of the embedding vectors, so that the cosine similarity of the corresponding embedding vectors of the matching clinical data and test knowledge in multiple groups of clinical data and multiple groups of test knowledge is larger, and the cosine similarity of the corresponding embedding vectors of the mismatching clinical data and test knowledge is smaller; the preset TopkRankLoss loss function can be obtained, and the preset TopkRankLoss loss function can be embedded to update the preset similarity ranking model to obtain the updated similarity ranking model. Finally, the embedding vectors can be input to perform the model training operation to obtain the trained target similarity ranking model.

[0063] Among them, the above cosine loss function of the embedding vectors can be set by the user or defaulted by the system, which is not limited here; the above preset TopkRankLoss loss function can be set by the user or defaulted by the system, which is not limited here. For example, the preset TopkRankLoss loss function can include:

[0064]

[0065] where x i is any item in the similarity list calculated by the above trained target similarity ranking model for any clinical test data in the clinical test dataset. is the position of x i in the similarity list, maxK is the total length of the similarity list, and positive, simi_positive, and negative indicate that the clinical data in the clinical test data matches, approximately matches, and does not match the test knowledge.

[0066] That is to say, when the clinical data matches the test knowledge, the test knowledge is in the first position in the similarity list, and the generated loss is 0 and no correction is required; when the clinical data and the test knowledge are in an approximate matching relationship or a mismatching relationship, the test knowledge is in the kth position in the similarity list, where k is an integer greater than 1, and the greater the value of k when the similarity between the clinical data and the test knowledge data is lower, the greater the generated loss.

[0067] It can be seen that in the embodiments of the present application, for the requirement of several test items most similar to a patient in the quality control scenario of clinical test examinations, the above algorithm introduces a preset TopkRankLoss loss function; during the training process, a vector cosine loss function is also introduced, and a mechanism of alternating training using both the embedded vector cosine loss function and the preset TopkRankLoss loss function is adopted, so that the trained target similarity ranking model can not only optimize the items most similar to the input in the general direction, but also meet the requirement of several test items most similar in the scenarios of the embodiments of the present application, solving the limitation problem that the existing methods can only optimize the most similar items.

[0068] In addition, the embodiments of the present application introduce an attention mechanism into the preset similarity ranking model, enabling the model to learn the importance of each feature in clinical data and test knowledge during the training process, thereby providing interpretability for the final clinical test results and solving the problem of poor interpretability of neural network models.

[0069] 103. Receive the data to be tested, input the data to be tested into the trained target similarity ranking model to obtain multiple pieces of test knowledge, and determine whether the clinical test conclusion corresponding to the data to be tested is in an abnormal state based on the multiple pieces of test knowledge.

[0070] Among them, the data to be tested may include clinical data to be tested, and the clinical data may include diseases, blood routine data, urine routine data, electrocardiograms, etc., which are not limited herein.

[0071] In a specific implementation, the data to be tested can be input into the trained target similarity ranking model to generate a similarity list and obtain multiple pieces of test knowledge. The server can further determine its position in the similarity list according to the multiple pieces of test knowledge, and determine whether the clinical test conclusion corresponding to the data to be tested is in an abnormal state according to this position. That is to say, the quality of the clinical test conclusion can be inspected to judge the accuracy of the clinical test conclusion.

[0072] It can be seen that in the embodiments of the present application, when doctors face complex conditions and it is difficult to give complete and accurate test results, the embodiments of the present application can provide quality protection for them, thereby reducing their workload and improving work efficiency.

[0073] In a possible example, step 103 above, inputting the data to be tested into the trained target similarity ranking model to obtain multiple pieces of test knowledge may include the following steps:

[0074] 311. Calculate a similarity matrix for the data to be tested based on the trained target similarity ranking model;

[0075] 312. Sort the similarities included in the similarity matrix according to the rule from large to small to obtain the sorted target similarity matrix;

[0076] 313. Determine the first k similarities of the target similarity matrix to obtain k test knowledge corresponding to the k similarities. The multiple test knowledge includes the k test knowledge. Among them, the k test knowledge corresponds to the data to be tested, and k is an integer greater than 1.

[0077] In a specific implementation, the server inputs the above data to be tested into the above target similarity sorting model to obtain a similarity matrix corresponding to the data to be tested. The similarity matrix includes the Jaccard similarity between the data to be tested and multiple test data in the target similarity sorting model.

[0078] Further, in order to improve the efficiency of data anomaly detection, the similarities included in the similarity matrix can be sorted according to a certain rule. For example, they can be sorted from large to small to generate a sorted target similarity matrix; since the greater the similarity, the higher the matching degree. Generally, in order to improve the detection efficiency, the first k similarities in the target similarity matrix can be determined, and the k test knowledge corresponding to the k similarities are the above multiple test knowledge. For example, the test knowledge corresponding to the first 5 similarities can be selected as the criterion for judging whether the following clinical test conclusion is in an abnormal state, where k is an integer greater than 1.

[0079] In a possible example, in step 103 above, judging whether the clinical test conclusion corresponding to the data to be tested is in an abnormal state according to the multiple test knowledge may include the following steps:

[0080] 321. Obtain the clinical test conclusion corresponding to the data to be tested, and judge whether the clinical test conclusion includes the k test knowledge;

[0081] 323. If the clinical test conclusion includes the k test knowledge, judge whether the clinical test conclusion is consistent with the k test knowledge; if it is consistent, determine that the clinical test conclusion is in a non-abnormal state; if it is not consistent, determine that the clinical test conclusion is in an abnormal state;

[0082] 323. If the clinical test conclusion does not include any one of the k test knowledge, determine that the clinical test conclusion is in an abnormal state.

[0083] Among them, the server can determine whether the clinical test conclusion is in an abnormal state based on multiple test knowledge and the clinical test conclusion corresponding to the data to be tested above. The abnormal state can include at least one of the following: multiple test state, missed test state, etc. Specifically, when the above clinical test conclusion includes k pieces of test knowledge and the above clinical test conclusion is inconsistent with the k pieces of test knowledge, it can be determined that the clinical test conclusion is in the multiple test state of the abnormal state; if the above clinical test conclusion does not include any one of the k pieces of test knowledge, it is determined that the clinical test conclusion is in the missed test state of the abnormal state; if the above clinical conclusion includes k pieces of test knowledge and is consistent with the k pieces of test knowledge, it is determined that the clinical test conclusion is in a non-abnormal state.

[0084] As Figure 1C shown, it is a network flow diagram of a data anomaly detection method; the server can input multiple groups of clinical data and multiple groups of test data into the first attention layer and the second attention layer of a preset similarity ranking model respectively, and output the clinical feature proportion vectors of multiple groups of clinical data and the knowledge feature proportion vectors of multiple groups of test knowledge respectively, and input the two into the DNN network model to obtain the embedding vectors corresponding to multiple groups of clinical data and multiple groups of test knowledge. The embedding vector represents the similarity between the multiple groups of clinical data and the multiple groups of test knowledge; furthermore, the embedding vector cosine loss function can be obtained, and the distribution of the above embedding vector can be constrained through the embedding vector cosine loss function, and the preset TopkRankLoss loss function can be obtained; furthermore, the preset similarity ranking model can be updated based on the preset TopkRankLoss loss function, and the updated model can be trained based on the embedding vector to obtain the trained target similarity ranking model; in this way, the attention mechanism is introduced into the preset similarity ranking model, so that the model can learn the importance of each feature in the clinical data and test knowledge during the training process, thereby providing interpretability for the final clinical test result and solving the problem of poor interpretability of the neural network model.

[0085] Finally, the server can input the data to be tested into the target similarity ranking model, obtain the similarity list, and based on the similarity list, determine multiple pieces of test knowledge corresponding to the data to be tested. Finally, by comparing the multiple pieces of test knowledge with the clinical test result corresponding to the data to be tested, it is judged whether the clinical test result is in an abnormal state; in this way, it is beneficial to improve the efficiency of clinical data testing.

[0086] It can be seen that the data anomaly detection method described in the embodiments of the present application is applied to a server, which can receive a clinical data set and an inspection knowledge set, generate a clinical inspection data set based on the clinical data set and the inspection knowledge set, use the clinical inspection data set as the training data of a preset similarity ranking model to perform a training operation, and obtain a trained target similarity ranking model; receive the data to be inspected, input the data to be inspected into the trained target similarity ranking model, obtain multiple inspection knowledge, and judge whether the clinical inspection conclusion corresponding to the data to be inspected is in an abnormal state based on the multiple inspection knowledge; in this way, multiple inspection knowledge corresponding to the data to be inspected can be obtained based on the trained target similarity ranking model, and the clinical inspection conclusion corresponding to the data to be inspected can be compared with the multiple inspection knowledge to judge whether the clinical inspection conclusion is in an abnormal state, which is beneficial to improving the efficiency of data anomaly detection.

[0087] Consistently with the above, please refer to Figure 2 , Figure 2 is a flowchart example of a data anomaly detection method disclosed in the embodiments of the present application, which is applied to a server. The data anomaly detection method may include the following steps:

[0088] 201. Receive a clinical data set and an inspection knowledge set, and obtain multiple groups of clinical data from the clinical data set. Among them, any group of clinical data in the multiple groups of clinical data includes: at least one clinical symptom and the clinical inspection conclusion corresponding to the at least one clinical symptom.

[0089] 202. Obtain multiple groups of inspection knowledge from the inspection knowledge set. Among them, any group of inspection knowledge in the multiple groups of inspection knowledge includes: a disease and at least one symptom inspection knowledge corresponding to the disease.

[0090] 203. Generate the clinical inspection data set based on the multiple groups of clinical data and the multiple groups of inspection knowledge.

[0091] 204. Use the multiple groups of clinical data as the input of the first attention layer of the preset similarity ranking model.

[0092] 205. Use the multiple groups of inspection knowledge as the input of the second attention layer of the preset similarity ranking model.

[0093] 206. Obtain the clinical feature proportion vector of the multiple groups of clinical data from the first attention layer, and obtain the knowledge feature proportion vector of the multiple groups of inspection knowledge from the second attention layer.

[0094] 207. Input the clinical feature weight vector and the knowledge feature weight vector into the preset similarity ranking model to obtain the embedding vectors corresponding to the multiple groups of clinical data and the multiple groups of inspection knowledge, where the embedding vector represents the similarity between the multiple groups of clinical data and the multiple groups of inspection knowledge.

[0095] 208. Train the preset similarity ranking model based on the embedding vector to obtain the trained target similarity ranking model.

[0096] 209. Receive the data to be inspected, input the data to be inspected into the trained target similarity ranking model to obtain multiple inspection knowledge, and judge whether the clinical inspection conclusion corresponding to the data to be inspected is in an abnormal state according to the multiple inspection knowledge.

[0097] Among them, the data anomaly detection method described in the above steps 201 - 209 can refer to Figure 1A the corresponding steps of the described data anomaly detection method.

[0098] It can be seen that the data anomaly detection method described in the embodiments of this application is applied to a server, which can receive a clinical data set and a test knowledge set, and obtain multiple groups of clinical data from the clinical data set. Among them, any group of clinical data in the multiple groups of clinical data includes: at least one clinical symptom and at least one clinical test conclusion corresponding to the at least one clinical symptom; obtain multiple groups of test knowledge from the test knowledge set. Among them, any group of test knowledge in the multiple groups of test knowledge includes: a disease and at least one symptom test knowledge corresponding to the disease; generate a clinical test data set based on the multiple groups of clinical data and the multiple groups of test knowledge; use the multiple groups of clinical data as the input of the first attention layer of the preset similarity ranking model; use the multiple groups of test knowledge as the input of the second attention layer of the preset similarity ranking model; obtain the clinical feature proportion vectors of the multiple groups of clinical data from the first attention layer, and obtain the knowledge feature proportion vectors of the multiple groups of test knowledge from the second attention layer; input the clinical feature proportion vectors and the knowledge feature proportion vectors into the preset similarity ranking model to obtain the embedding vectors corresponding to the multiple groups of clinical data and the multiple groups of test knowledge, where the embedding vectors represent the similarity between the multiple groups of clinical data and the multiple groups of test knowledge; train the preset similarity ranking model based on the embedding vectors to obtain a trained target similarity ranking model; receive the data to be tested, input the data to be tested into the trained target similarity ranking model to obtain multiple pieces of test knowledge, and determine whether the clinical test conclusion corresponding to the data to be tested is in an abnormal state based on the multiple pieces of test knowledge; in this way, the server can train the above-mentioned preset similarity ranking model based on the clinical data set and the test knowledge, that is to say, it makes full use of the requirement for several test items most similar to the patient in the clinical test quality control scenario, which is beneficial to improving the accuracy of the model in calculating similarity, and introduces an attention mechanism in the process of model training, so that the model can learn the importance of each feature in the clinical data and the test knowledge during the training process, thereby providing interpretability for the final clinical test results, and solving the problem of poor interpretability of the neural network model; finally, the above-mentioned clinical test conclusion can be judged and evaluated according to the target similarity ranking model to determine whether the clinical test conclusion is abnormal, which is beneficial to improving the efficiency of data anomaly detection.

[0099] Consistently with the above, please refer to Figure 3 , Figure 3 is a flow example diagram of a data anomaly detection method disclosed in the embodiments of this application, which is applied to a server. The data anomaly detection method may include the following steps:

[0100] 301. Receive a clinical data set and a test knowledge set, and generate a clinical test data set based on the clinical data set and the test knowledge set.

[0101] 302. Use the clinical test data set as the training data of the preset similarity ranking model to perform a training operation to obtain a trained target similarity ranking model.

[0102] 303. Receive the data to be inspected, input the data to be inspected into the trained target similarity ranking model, and obtain multiple inspection knowledge, where the multiple inspection knowledge includes the k inspection knowledge, and the k inspection knowledge corresponds to the data to be inspected, and k is an integer greater than 1.

[0103] 304. Obtain the clinical inspection conclusion corresponding to the data to be inspected, and determine whether the clinical inspection conclusion includes the k inspection knowledge.

[0104] 305. If the clinical inspection conclusion includes the k inspection knowledge, determine whether the clinical inspection conclusion is consistent with the k inspection knowledge; if it is consistent, determine that the clinical inspection conclusion is in a non-abnormal state; if it is inconsistent, determine that the clinical inspection conclusion is in an abnormal state.

[0105] 306. If the clinical inspection conclusion does not include any one of the k inspection knowledge, determine that the clinical inspection conclusion is in an abnormal state.

[0106] Among them, the data anomaly detection method described in the above steps 301-306 can refer to Figure 1A the corresponding steps of the data anomaly detection method described.

[0107] It can be seen that the data anomaly detection method described in the embodiments of the present application is applied to a server, which can receive a clinical data set and an inspection knowledge set, generate a clinical inspection data set based on the clinical data set and the inspection knowledge set; use the clinical inspection data set as the training data of a preset similarity ranking model to perform a training operation to obtain a trained target similarity ranking model; receive the data to be inspected, input the data to be inspected into the trained target similarity ranking model, and obtain multiple inspection knowledge, where the multiple inspection knowledge includes the k inspection knowledge, and the k inspection knowledge corresponds to the data to be inspected, and k is an integer greater than 1; obtain the clinical inspection conclusion corresponding to the data to be inspected, and determine whether the clinical inspection conclusion includes the k inspection knowledge; if the clinical inspection conclusion includes the k inspection knowledge, determine whether the clinical inspection conclusion is consistent with the k inspection knowledge; if it is consistent, determine that the clinical inspection conclusion is in a non-abnormal state; if it is inconsistent, determine that the clinical inspection conclusion is in an abnormal state; if the clinical inspection conclusion does not include any one of the k inspection knowledge, determine that the clinical inspection conclusion is in an abnormal state; thus, the k inspection knowledge with the highest similarity corresponding to the data to be inspected can be determined based on the trained target similarity ranking model, and the clinical inspection conclusion corresponding to the above data to be inspected is compared with the multiple inspection knowledge to determine whether the clinical inspection conclusion is in an abnormal state, which is beneficial to improving the efficiency of data anomaly detection.

[0108] Consistently with the above, please refer to Figure 4 ,Figure 4 A schematic structural diagram of a server provided by an embodiment of the present application is shown as Figure 4 shown, including a processor, a communication interface, a memory, and one or more programs. The processor, the communication interface, and the memory are interconnected. Among them, the memory is used to store computer programs, and the computer programs include program instructions. The processor is configured to call the program instructions. The above one or more programs include instructions for performing the following steps:

[0109] Receive a clinical data set and a test knowledge set, and generate a clinical test data set based on the clinical data set and the test knowledge set;

[0110] Use the clinical test data set as training data for a preset similarity ranking model to perform a training operation, and obtain a trained target similarity ranking model;

[0111] Receive the data to be tested, input the data to be tested into the trained target similarity ranking model, obtain multiple test knowledge, and determine whether the clinical test conclusion corresponding to the data to be tested is in an abnormal state based on the multiple test knowledge.

[0112] It can be seen that the server described in the embodiment of the present application can receive a clinical data set and a test knowledge set, generate a clinical test data set based on the clinical data set and the test knowledge set, use the clinical test data set as training data for a preset similarity ranking model to perform a training operation, and obtain a trained target similarity ranking model; receive the data to be tested, input the data to be tested into the trained target similarity ranking model, obtain multiple test knowledge, and determine whether the clinical test conclusion corresponding to the data to be tested is in an abnormal state based on the multiple test knowledge. In this way, multiple test knowledge corresponding to the data to be tested can be obtained based on the trained target similarity ranking model, and the multiple test knowledge can be compared with the clinical test conclusion corresponding to the above data to be tested to determine whether the clinical test conclusion is in an abnormal state, which is beneficial to improving the efficiency of data anomaly detection.

[0113] In a possible example, in terms of generating the clinical test data set based on the clinical data set and the test knowledge set, the program is used to execute the instructions for the following steps:

[0114] Obtain multiple groups of clinical data from the clinical data set. Among them, any group of clinical data in the multiple groups of clinical data includes: at least one clinical symptom and the clinical test conclusion corresponding to the at least one clinical symptom;

[0115] Obtain multiple groups of test knowledge from the test knowledge set. Among them, any group of test knowledge in the multiple groups of test knowledge includes: a disease and at least one symptom test knowledge corresponding to the disease;

[0116] Generate the clinical test data set based on the multiple groups of clinical data and the multiple groups of test knowledge.

[0117] In a possible example, in terms of generating the clinical test data set based on the multiple groups of clinical data and the multiple groups of test knowledge, the program is used to execute instructions for the following steps:

[0118] Obtain any group of clinical data from the multiple groups of clinical data as the target clinical data, and calculate multiple Jaccard similarities corresponding to the multiple groups of test knowledge;

[0119] Obtain the test knowledge corresponding to the maximum value among the multiple Jaccard similarities as the target test knowledge, generate the mapping relationship between the target clinical data and the target test knowledge as the clinical test data, repeat the above steps, obtain multiple clinical test data corresponding to the multiple groups of clinical data, and combine the multiple clinical test data to obtain the clinical test data set.

[0120] In a possible example, in terms of calculating the multiple Jaccard similarities corresponding to the target clinical data and the multiple groups of test knowledge, the program is used to execute instructions for the following steps:

[0121] Obtain a preset Jaccard calculation formula, obtain any group of test knowledge from the multiple groups of test knowledge, use the target clinical data and the any group of test knowledge as the input of the preset Jaccard calculation formula, obtain the Jaccard similarity between the target clinical data and the any group of test knowledge, repeat the above steps, and obtain multiple Jaccard similarities.

[0122] In a possible example, in terms of using the clinical test data set as the training data of a preset similarity ranking model to perform a training operation to obtain a trained target similarity ranking model, the program is used to execute instructions for the following steps:

[0123] Use the multiple groups of clinical data as the input of the first attention layer of the preset similarity ranking model;

[0124] Use the multiple groups of test knowledge as the input of the second attention layer of the preset similarity ranking model;

[0125] Obtain the clinical feature proportion vector of the multiple groups of clinical data from the first attention layer, and obtain the knowledge feature proportion vector of the multiple groups of test knowledge from the second attention layer;

[0126] Input the clinical feature proportion vector and the knowledge feature proportion vector into the preset similarity ranking model to obtain the embedding vectors corresponding to the multiple groups of clinical data and the multiple groups of test knowledge, where the embedding vectors represent the similarity between the multiple groups of clinical data and the multiple groups of test knowledge;

[0127] Train the preset similarity ranking model based on the embedding vectors to obtain the trained target similarity ranking model.

[0128] In a possible example, in the aspect of inputting the data to be tested into the trained target similarity ranking model to obtain multiple pieces of test knowledge, the program is used to execute the instructions for the following steps:

[0129] Calculate a similarity matrix for the data to be tested based on the trained target similarity ranking model;

[0130] Sort the similarities included in the similarity matrix in descending order to obtain the sorted target similarity matrix;

[0131] Determine the top k similarities of the target similarity matrix to obtain the k pieces of test knowledge corresponding to the k similarities. The multiple pieces of test knowledge include the k pieces of test knowledge, where the k pieces of test knowledge correspond to the data to be tested, and k is an integer greater than 1.

[0132] In a possible example, in the aspect of judging whether the clinical test conclusion corresponding to the data to be tested is in an abnormal state based on the multiple pieces of test knowledge, the program is used to execute the instructions for the following steps:

[0133] Obtain the clinical test conclusion corresponding to the data to be tested, and judge whether the clinical test conclusion includes the k pieces of test knowledge;

[0134] If the clinical test conclusion includes the k pieces of test knowledge, judge whether the clinical test conclusion is consistent with the k pieces of test knowledge; if consistent, determine that the clinical test conclusion is in a non-abnormal state; if inconsistent, determine that the clinical test conclusion is in an abnormal state;

[0135] If the clinical test conclusion does not include any one of the k pieces of test knowledge, determine that the clinical test conclusion is in an abnormal state.

[0136] The above mainly introduces the solution of the embodiment of the present application from the perspective of the execution process on the method side. It can be understood that in order for the server to implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments provided in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0137] The embodiments of the present application can divide the server into functional units according to the above method examples. For example, each functional unit can be divided corresponding to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. It should be noted that the division of units in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0138] Consistent with the above, please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a data anomaly detection device disclosed in the embodiment of the present application, applied to a server. The device includes: a receiving unit 501, a training unit 502, and a judging unit 503, where

[0139] The receiving unit 501 is configured to receive a clinical data set and a test knowledge set, and generate a clinical test data set according to the clinical data set and the test knowledge set;

[0140] The training unit 502 is configured to use the clinical test data set as training data of a preset similarity ranking model to perform a training operation, and obtain a trained target similarity ranking model;

[0141] The judging unit 503 is configured to receive the data to be tested, input the data to be tested into the trained target similarity ranking model, obtain a plurality of test knowledge, and judge whether the clinical test conclusion corresponding to the data to be tested is in an abnormal state according to the plurality of test knowledge.

[0142] It can be seen that the data anomaly detection device described in the embodiments of the present application is applied to a server. The device can receive a clinical data set and a test knowledge set, generate a clinical test data set based on the clinical data set and the test knowledge set, use the clinical test data set as the training data of a preset similarity ranking model to perform a training operation, and obtain a trained target similarity ranking model; receive the data to be tested, input the data to be tested into the trained target similarity ranking model, obtain multiple pieces of test knowledge, and determine whether the clinical test conclusion corresponding to the data to be tested is in an abnormal state based on the multiple pieces of test knowledge; in this way, multiple pieces of test knowledge corresponding to the data to be tested can be obtained based on the trained target similarity ranking model, and the clinical test conclusion corresponding to the data to be tested can be compared with the multiple pieces of test knowledge to determine whether the clinical test conclusion is in an abnormal state, which is beneficial to improving the efficiency of data anomaly detection.

[0143] In a possible example, in terms of generating the clinical test data set based on the clinical data set and the test knowledge set, the receiving unit 501 is specifically configured to:

[0144] Obtain multiple groups of clinical data from the clinical data set, where any one group of the multiple groups of clinical data includes: at least one clinical symptom and the clinical test conclusion corresponding to the at least one clinical symptom;

[0145] Obtain multiple groups of test knowledge from the test knowledge set, where any one group of the multiple groups of test knowledge includes: a disease and at least one symptom test knowledge corresponding to the disease;

[0146] Generate the clinical test data set based on the multiple groups of clinical data and the multiple groups of test knowledge.

[0147] In a possible example, in terms of generating the clinical test data set based on the multiple groups of clinical data and the multiple groups of test knowledge, the receiving unit 501 is specifically configured to:

[0148] Obtain any one group of clinical data from the multiple groups of clinical data as the target clinical data, and calculate multiple Jaccard similarities corresponding to the multiple groups of test knowledge for the target clinical data;

[0149] Obtain the test knowledge corresponding to the maximum value among the multiple Jaccard similarities as the target test knowledge, generate the mapping relationship between the target clinical data and the target test knowledge as the clinical test data, repeat the above steps, obtain multiple pieces of clinical test data corresponding to the multiple groups of clinical data, and combine the multiple pieces of clinical test data to obtain the clinical test data set.

[0150] In a possible example, when calculating multiple Jaccard similarities corresponding to the target clinical data and the multiple sets of test knowledge, the receiving unit 501 is specifically further configured to:

[0151] Obtain a preset Jaccard calculation formula, obtain any one set of test knowledge from the multiple sets of test knowledge, use the target clinical data and the any one set of test knowledge as the input of the preset Jaccard calculation formula to obtain the Jaccard similarity between the target clinical data and the any one set of test knowledge, and repeat the above steps to obtain multiple Jaccard similarities.

[0152] In a possible example, when using the clinical test data set as the training data of a preset similarity ranking model to perform a training operation to obtain a trained target similarity ranking model, the training unit 502 is specifically configured to:

[0153] Use the multiple sets of clinical data as the input of the first attention layer of the preset similarity ranking model;

[0154] Use the multiple sets of test knowledge as the input of the second attention layer of the preset similarity ranking model;

[0155] Obtain the clinical feature proportion vector of the multiple sets of clinical data from the first attention layer, and obtain the knowledge feature proportion vector of the multiple sets of test knowledge from the second attention layer;

[0156] Input the clinical feature proportion vector and the knowledge feature proportion vector into the preset similarity ranking model to obtain the embedding vectors corresponding to the multiple sets of clinical data and the multiple sets of test knowledge, where the embedding vectors represent the similarity between the multiple sets of clinical data and the multiple sets of test knowledge;

[0157] Train the preset similarity ranking model based on the embedding vectors to obtain the trained target similarity ranking model.

[0158] In a possible example, when inputting the data to be tested into the trained target similarity ranking model to obtain multiple sets of test knowledge, the judgment unit 503 is specifically configured to:

[0159] Calculate a similarity matrix for the data to be tested based on the trained target similarity ranking model;

[0160] Sort the similarities included in the similarity matrix in descending order to obtain a sorted target similarity matrix;

[0161] Determine the top k similarities of the target similarity matrix to obtain k pieces of verification knowledge corresponding to the k similarities. The multiple pieces of verification knowledge include the k pieces of verification knowledge, where the k pieces of verification knowledge correspond to the data to be inspected, and k is an integer greater than 1.

[0162] In a possible example, in terms of determining whether the clinical test conclusion corresponding to the data to be inspected is in an abnormal state based on the multiple pieces of verification knowledge, the determination unit 503 is specifically further configured to:

[0163] Obtain the clinical test conclusion corresponding to the data to be inspected, and determine whether the clinical test conclusion includes the k pieces of verification knowledge;

[0164] If the clinical test conclusion includes the k pieces of verification knowledge, determine whether the clinical test conclusion is consistent with the k pieces of verification knowledge; if consistent, determine that the clinical test conclusion is in a non-abnormal state; if inconsistent, determine that the clinical test conclusion is in an abnormal state;

[0165] If the clinical test conclusion does not include any one of the k pieces of verification knowledge, determine that the clinical test conclusion is in an abnormal state.

[0166] The embodiments of the present application further provide a computer-readable storage medium, where the computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute some or all of the steps of any one of the data anomaly detection methods described in the foregoing method embodiments.

[0167] The embodiments of the present application further provide a computer program product, where the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute some or all of the steps of any one of the data anomaly detection methods described in the foregoing method embodiments.

[0168] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0169] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0170] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0171] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0172] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software program modules.

[0173] If the above-mentioned integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. And the aforementioned memory includes: USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs and other media that can store program codes.

[0174] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory, and the memory can include: flash drives, ROM, RAM, magnetic disks, or optical discs, etc.

[0175] The above has introduced the embodiments of the present application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for detecting data anomalies, characterized in that, applied to a server, including: Receiving a clinical data set and an inspection knowledge set, Obtaining multiple groups of clinical data from the clinical data set, wherein any one group of the multiple groups of clinical data includes: at least one clinical symptom and a clinical test conclusion corresponding to the at least one clinical symptom; Obtaining multiple groups of inspection knowledge from the inspection knowledge set, wherein any one group of the multiple groups of inspection knowledge includes: a disease and at least one symptom inspection knowledge corresponding to the one disease; Generating a clinical test data set according to the similarity between the multiple groups of clinical data and the multiple groups of inspection knowledge; Taking the multiple groups of clinical data as the input of the first attention layer of a preset similarity ranking model; Taking the multiple groups of inspection knowledge as the input of the second attention layer of the preset similarity ranking model; Obtaining a clinical feature proportion vector of the multiple groups of clinical data from the first attention layer, and obtaining a knowledge feature proportion vector of the multiple groups of inspection knowledge from the second attention layer; Inputting the clinical feature proportion vector and the knowledge feature proportion vector into the preset similarity ranking model to obtain embedding vectors corresponding to the multiple groups of clinical data and the multiple groups of inspection knowledge, wherein the embedding vectors represent the similarity between the multiple groups of clinical data and the multiple groups of inspection knowledge; Training the preset similarity ranking model based on the embedding vectors to obtain a trained target similarity ranking model; Receiving data to be inspected, inputting the data to be inspected into the trained target similarity ranking model to obtain multiple inspection knowledge, and judging whether the clinical test conclusion corresponding to the data to be inspected is in an abnormal state according to the multiple inspection knowledge.

2. The method according to claim 1, characterized in that, The generating the clinical test data set according to the multiple groups of clinical data and the multiple groups of inspection knowledge includes: Obtaining any one group of clinical data from the multiple groups of clinical data as target clinical data, and calculating multiple Jaccard similarities corresponding to the target clinical data and the multiple groups of inspection knowledge; Obtaining the inspection knowledge corresponding to the maximum value among the multiple Jaccard similarities as the target inspection knowledge, generating a mapping relationship between the target clinical data and the target inspection knowledge as clinical test data, repeating the above steps to obtain multiple clinical test data corresponding to the multiple groups of clinical data, and combining the multiple clinical test data to obtain the clinical test data set.

3. The method according to claim 2, characterized in that, The calculating the multiple Jaccard similarities corresponding to the target clinical data and the multiple groups of inspection knowledge includes: Obtaining a preset Jaccard calculation formula, obtaining any one group of inspection knowledge from the multiple groups of inspection knowledge, taking the target clinical data and the any one group of inspection knowledge as the input of the preset Jaccard calculation formula to obtain the Jaccard similarity between the target clinical data and the any one group of inspection knowledge, repeating the above steps to obtain multiple Jaccard similarities.

4. The method according to claim 1, characterized in that, Inputting the data to be inspected into the trained target similarity ranking model to obtain multiple inspection knowledge, including: Calculating a similarity matrix for the data to be inspected based on the trained target similarity ranking model; Sorting the similarities included in the similarity matrix in descending order to obtain a sorted target similarity matrix; Determining the top k similarities of the target similarity matrix to obtain k inspection knowledge corresponding to the k similarities. The multiple inspection knowledge includes the k inspection knowledge, where the k inspection knowledge corresponds to the data to be inspected, and k is an integer greater than 1.

5. The method according to claim 4, wherein, Judging whether the clinical inspection conclusion corresponding to the data to be inspected is in an abnormal state based on the multiple inspection knowledge includes: Obtaining the clinical inspection conclusion corresponding to the data to be inspected, and judging whether the clinical inspection conclusion includes the k inspection knowledge; If the clinical inspection conclusion includes the k inspection knowledge, judging whether the clinical inspection conclusion is consistent with the k inspection knowledge; if consistent, determining that the clinical inspection conclusion is in a non-abnormal state; if inconsistent, determining that the clinical inspection conclusion is in an abnormal state; If the clinical inspection conclusion does not include any one of the k inspection knowledge, determining that the clinical inspection conclusion is in an abnormal state.

6. A data anomaly detection device, wherein, Applied to a server, the device includes: a receiving unit, a training unit, and a judging unit, wherein, The receiving unit is configured to receive a clinical data set and an inspection knowledge set, obtain multiple groups of clinical data from the clinical data set, where any one group of the multiple groups of clinical data includes: at least one clinical symptom and a clinical inspection conclusion corresponding to the at least one clinical symptom; obtain multiple groups of inspection knowledge from the inspection knowledge set, where any one group of the multiple groups of inspection knowledge includes: a disease and at least one symptom inspection knowledge corresponding to the one disease; generating a clinical inspection data set according to the similarity between the multiple groups of clinical data and the multiple groups of inspection knowledge; The training unit is configured to use the multiple groups of clinical data as the input of the first attention layer of a preset similarity ranking model; use the multiple groups of inspection knowledge as the input of the second attention layer of the preset similarity ranking model; obtain a clinical feature proportion vector of the multiple groups of clinical data from the first attention layer, and obtain a knowledge feature proportion vector of the multiple groups of inspection knowledge from the second attention layer; input the clinical feature proportion vector and the knowledge feature proportion vector into the preset similarity ranking model to obtain embedding vectors corresponding to the multiple groups of clinical data and the multiple groups of inspection knowledge, where the embedding vectors represent the similarity between the multiple groups of clinical data and the multiple groups of inspection knowledge; training the preset similarity ranking model based on the embedding vectors to obtain a trained target similarity ranking model; The determination unit is configured to receive the data to be inspected, input the data to be inspected into the trained target similarity ranking model, obtain a plurality of inspection knowledge, and determine whether the clinical inspection conclusion corresponding to the data to be inspected is in an abnormal state according to the plurality of inspection knowledge.

7. A server, characterized in that, it includes a processor, an input device, an output device and a memory, the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Diagnosis result identification method, diagnosis result identification model training method, computer device and storage medium

    CN110808095A