Medical interaction data quality inspection method, related device, equipment and storage medium

By obtaining the feature representation of medical interaction data and comparing the reference feature set, the problem of quality inspection efficiency and accuracy of medical interaction data is solved, and efficient and accurate data quality inspection and error type expansion are achieved.

CN120356703APending Publication Date: 2025-07-22ANHUI IFLYHEALTH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510197969.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

How to achieve quality inspection of medical interaction data and improve the efficiency and accuracy of data quality inspection, especially in scenarios where data is scarce or difficult to obtain.

Method used

By obtaining the reference feature set and obtaining the feature representation of the fragment to be detected, the target feature representation is selected using feature similarity, and quality inspection is carried out based on the accompanying information of the target feature representation, a reference feature set involving various error causes is constructed for feature comparison, similar errors are detected and new error type extensions are supported.

Benefits of technology

It improves the efficiency and accuracy of medical interactive data quality inspection, can systematically scan and detect similar errors, reduce manual intervention, and reduce quality inspection costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356703A_ABST
    Figure CN120356703A_ABST
Patent Text Reader

Abstract

The invention discloses a medical interaction data quality inspection method, a related device, equipment and a storage medium, and the method comprises the steps: obtaining a reference feature set, and obtaining the first feature representation of each to-be-inspected segment in first medical interaction data; based on the similarity between the first feature representation of the to-be-detected fragment and the second feature representation of each reference fragment, selecting the second feature representation as a target feature representation; and obtaining a quality inspection result of the to-be-inspected fragment based on the attached information represented by the target feature. According to the scheme, the medical interaction data can be subjected to quality inspection, and the efficiency and precision of data quality inspection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and in particular, to a method and related devices, equipment, and storage media for quality inspection of medical interaction data. Background Art

[0002] In the medical field, which is a highly complex and data-intensive industry, generative large language models are gradually becoming an important tool for improving medical efficiency and quality. By automatically processing medical instructions, generating medical record summaries, medical reports, interpreting clinical instructions, etc., the automation level of medical services is greatly improved.

[0003] During the medical consultation process, the performance level of the model largely depends on high-quality medical interaction data. In scenarios where data is scarce or difficult to obtain, synthetic medical interaction data becomes an effective alternative. However, the quality of the data, whether it is real-world collected medical interaction data or model-synthesized medical interaction data, is the key. In view of this, how to perform quality inspection on medical interaction data and improve the efficiency and accuracy of data quality inspection has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a method and related devices, equipment, and storage media for quality inspection of medical interaction data, which can perform quality inspection on medical interaction data and improve the efficiency and accuracy of data quality inspection.

[0005] To solve the above technical problem, in a first aspect of this application, a method for quality inspection of medical interaction data is provided, including: obtaining a reference feature set, and obtaining a first feature representation of each to-be-inspected segment in the first medical interaction data; wherein, the reference feature set contains second feature representations of several reference segments, and the second feature representation is at least attached with: the reason for the error of the reference segment; based on the similarity between the first feature representation of the to-be-inspected segment and the second feature representations of each reference segment, selecting the second feature representation as the target feature representation; and obtaining the quality inspection result of the to-be-inspected segment based on the attached information of the target feature representation.

[0006] To solve the above technical problem, in a second aspect of this application, a device for quality inspection of medical interaction data is provided, including: an obtaining module, a selection module, and a quality inspection module. The obtaining module is used to obtain a reference feature set and obtain a first feature representation of each to-be-inspected segment in the first medical interaction data; wherein, the reference feature set contains second feature representations of several reference segments, and the second feature representation is at least attached with: the reason for the error of the reference segment; the selection module is used to select the second feature representation as the target feature representation based on the similarity between the first feature representation of the to-be-inspected segment and the second feature representations of each reference segment; and the quality inspection module is used to obtain the quality inspection result of the to-be-inspected segment based on the attached information of the target feature representation.

[0007] To solve the above technical problems, a third aspect of the present application provides an electronic device, which at least includes a memory and a processor coupled to each other. The memory stores at least program instructions, and the processor is configured to execute the program instructions to implement the medical interaction data quality inspection method in the above first aspect.

[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the medical interaction data quality inspection method in the above first aspect.

[0009] In the above solution, a reference feature set is obtained, and a first feature representation of each to-be-inspected segment in the first medical interaction data is obtained. The reference feature set includes second feature representations of several reference segments, and the second feature representation is at least attached with: the error reason of the reference segment. Then, based on the similarity between the first feature representation of the to-be-inspected segment and the second feature representations of each reference segment, the second feature representation is selected as the target feature representation, and based on the attached information of the target feature representation, the quality inspection result of the to-be-inspected segment is obtained. Therefore, on the one hand, it is possible to perform quality inspection on medical interaction data through feature similarity comparison and improve the efficiency of data quality inspection. On the other hand, by constructing a reference feature set involving various error reasons for feature comparison, it is possible to systematically scan the entire to-be-inspected data set through known error types to detect whether there are similar errors, and it is possible to ensure as much as possible that all medical interaction data with the same error type is accurately detected, thereby improving the accuracy of data quality inspection. In addition, when it is necessary to support the detection of new error types, it is only necessary to add a reference segment involving the new error type and its second feature representation to the reference feature set, which helps to achieve efficient expansion of error types, reduce the need for manual intervention, and thus reduce the cost of data quality inspection. Description of the Drawings

[0010] Figure 1 is a schematic flowchart of an embodiment of the medical interaction data quality inspection method of the present application; Figure 2 is a schematic diagram of the process of an embodiment of the medical interaction data quality inspection method of the present application; Figure 3 is a schematic framework diagram of an embodiment of the medical interaction data quality inspection device of the present application; Figure 4 is a schematic framework diagram of an embodiment of the electronic device of the present application; Figure 5 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed Embodiments

[0011] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.

[0012] In the following description, specific details such as specific system architectures, interfaces, and technologies are set forth for the purpose of illustration and not limitation in order to provide a thorough understanding of the present application.

[0013] The terms "system" and "network" are often used interchangeably herein. The term " / and" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: the situation where A exists alone, the situation where A and B exist simultaneously, and the situation where B exists alone. In addition, the fragment " / " in this article generally represents an "or" relationship between the preceding and following associated objects. In addition, "plurality" in this article means two or more than two.

[0014] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the medical interaction data quality inspection method of the present application. Specifically, it may include the following steps: Step S11: Obtain a reference feature set and obtain the first feature representation of each fragment to be inspected in the first medical interaction data.

[0015] In the embodiments of the present disclosure, the reference feature set may include second feature representations of several reference segments. The reference segment is a text segment with incorrect expression in a medical scenario. As a possible implementation example, the reference segment may be derived from the text segments with incorrect expression in the second medical interaction data. Of course, the reference segment may also be derived from other sources, such as text segments with incorrect expression in a medical interaction scenario generated by a generative large model. The specific source of the reference segment is not limited herein, and no further examples will be given one by one. It should be noted that the "medical interaction data" such as the first medical interaction data and the second medical interaction data in the embodiments of the present disclosure may include, but are not limited to, human-to-human dialogue data, human-machine dialogue data, machine-generated data, etc. in a medical scenario, which is not limited herein. Taking the "medical interaction data" as human-to-human dialogue data in a medical scenario as an example, the "medical interaction data" may include at least one round of dialogue. The inquiry sentence in each round of dialogue may be input by the patient, and the answer sentence in response to the inquiry sentence in each round of dialogue may be input by the doctor; or, taking the "medical interaction data" as human-machine dialogue data in a medical scenario as an example, the "medical interaction data" may include at least one round of dialogue. The inquiry sentence in each round of dialogue may be input by the patient, and the answer sentence in response to the inquiry sentence in each round of dialogue may be input by a machine (such as a large language model, etc.); or, taking the "medical interaction data" as machine-generated data in a medical scenario as an example, the inquiry sentence and the answer sentence in response to the inquiry sentence in each round of dialogue may both be generated by a machine (such as a large language model). Of course, the above examples are only several possible examples of the "medical interaction data", and other possible situations will not be exemplified one by one herein.

[0016] In one implementation scenario, as a possible implementation example, the text segments with incorrect expression in the second medical interaction data may be directly obtained as the reference segments. Exemplarily, newly stored medical interaction data (such as newly stored human-to-human dialogue data, human-machine dialogue data, machine-generated data, etc. in a medical scenario, which is not limited herein) may be obtained as the second medical interaction data, and the second medical interaction data may be quality inspected based on a data quality inspection model to obtain the text segments with incorrect expression in the second medical interaction data as the reference segments. It should be noted that the data quality inspection model may be trained based on the sample medical interaction data marked with incorrect segments. Or, the sample medical interaction data marked with incorrect segments may also be obtained as the second medical interaction data, and the incorrect segments marked in the second medical interaction data may be selected as the reference segments. Of course, the above examples are only several possible examples of obtaining the reference segments in the actual application process, and other possible ways of obtaining the reference segments will not be exemplified one by one herein.

[0017] In another implementation scenario, as another possible implementation example, different from the foregoing implementation manners, it is possible to first obtain the text fragments with incorrect expressions in the second medical interaction data as candidate fragments, and then perform semantic completion on the candidate fragments based on the second medical interaction data to obtain reference fragments. In the above manner, after obtaining the text fragments with incorrect expressions, semantic completion is further performed to obtain reference fragments, which can ensure semantic integrity as much as possible and help improve the accuracy of subsequent similarity measurement.

[0018] In a specific implementation scenario, in order to obtain candidate fragments, it is possible to obtain newly stored medical interaction data as the second medical interaction data, and perform quality inspection on the second medical interaction data based on a data quality inspection model to obtain the text fragments with incorrect expressions in the second medical interaction data as candidate fragments. It should be noted that the data quality inspection model can be trained based on sample medical interaction data marked with incorrect fragments. In the above manner, by performing quality inspection on the newly stored medical interaction data through the data quality inspection model to obtain the text fragments with incorrect expressions therein as candidate fragments, the efficiency of screening incorrect fragments can be improved.

[0019] In a specific implementation scenario, in order to obtain candidate fragments, it is also possible to obtain sample medical interaction data marked with incorrect fragments as the second medical interaction data, and select the incorrectly marked fragments in the second medical interaction data as candidate fragments. It should be noted that the sample medical interaction data marked with incorrect fragments can be used to train the data quality inspection model, and the data quality inspection model can be used to perform quality inspection on the newly stored medical interaction data to detect the text fragments with incorrect expressions therein. In the above manner, by directly selecting the incorrectly marked fragments in the sample medical interaction data as candidate fragments, the accuracy of screening incorrect fragments can be improved.

[0020] In a specific implementation scenario, after obtaining the candidate fragments, it is possible to perform semantic completion on the candidate fragments based on the semantic completion model with reference to the second medical interaction data to obtain reference fragments, and the semantic completion model is trained based on sample fragment pairs. The sample fragment pairs include a first fragment with incomplete semantics and a second fragment obtained after semantic completion of the first fragment, and the sample fragment pairs can be generated by a large language model. Exemplarily, open-source large models, large language models with parameter fine-tuning, or custom large models, etc., can be used to synthesize the sample fragment pairs, and then the sample fragment pairs are used to train a dedicated model for semantic completion of incorrect fragments, that is, the semantic completion model. In the above manner, by generating sample fragment pairs through a large language model and then training a semantic completion model based on the sample fragment pairs, so as to perform semantic completion on the candidate fragments based on the semantic completion model with reference to the second medical interaction data, the efficiency and accuracy of semantic completion of incorrect fragments can be improved.

[0021] In a specific implementation scenario, for the sake of easy understanding, taking the detection of the error fragment "using bread to assist swallowing" as an example, the reason for the error is that "fish bones cannot be swallowed and discharged using bread". Then, due to the semantic loss in the error fragment (specifically the lack of a subject in this example), the relevant operations of semantic completion described above can be used to perform semantic completion on the error fragment "using bread to assist swallowing", resulting in a semantically complete reference fragment "using bread to assist swallowing when swallowing fish bones". Of course, the above example is only one possible example in the actual application process, and other possible situations will not be listed one by one here.

[0022] In another implementation scenario, as another possible implementation example, different from the foregoing implementation manner, it is possible to first obtain the text fragment with an error in the second medical interaction data as a candidate fragment, and obtain the reason for the error of the candidate fragment. Then, based on the second medical interaction data, semantic completion is performed on the candidate fragment to obtain a reference fragment, and semantic completion is performed on the reason for the error of the candidate fragment based on the second medical interaction data to obtain a semantically complete reason for the error. It should be noted that the specific method for performing semantic completion on the reason for the error can refer to the relevant operations for performing semantic completion on the candidate fragment, which will not be elaborated here. In addition, for the sake of easy understanding, still taking the detection of the error fragment "using bread to assist swallowing" as an example, the reason for the error is "wrong method". After semantic completion, the semantically complete reason for the error is "fish bones cannot be swallowed and discharged using bread". Of course, the above example is only one possible example in the actual application process, and other possible situations will not be listed one by one here.

[0023] In the disclosed embodiment, the second feature representation is at least accompanied by: the cause of the error of the reference fragment. Exemplarily, after the reference fragment is extracted, the reference fragment can be vectorized by embedding (such as BGE-M3 embedding model) and other methods to extract the second feature representation of the reference fragment. Of course, the above example is only a possible example of feature extraction of the reference fragment, and other possible methods are not limited here. In addition, the second feature representation can be accompanied by at least the cause of the error of the reference fragment. Outside of this case, the cause of the error of the reference fragment can specifically be the cause of the error with complete semantics after the aforementioned semantic completion. Of course, in addition to this, the second feature representation can also be accompanied by other information such as the second medical interaction data from which the reference fragment is derived. In addition, the second feature representation can also be accompanied by the cause of the error of the reference fragment and the second medical interaction data from which the reference fragment is derived. As a possible example, the second feature representation of the reference fragment can also be accompanied by label information, which characterizes the medical environment to which the reference fragment belongs, such as data such as departments and scenes, and the specific content of the label information is not limited here. In order to store the reference feature set, a vector database such as Milvus can be used, in which the second feature representation of the reference fragment is stored in vector form, and other information such as incidental information and label information of the second feature representation can be stored in scalar form.

[0024] In addition, the first medical interaction data can be generated by an artificial intelligence model (such as the aforementioned "machine-generated data"). Of course, the first medical interaction data can also be collected on a medical consultation platform (such as the aforementioned "human-to-human dialogue data" and "human-to-machine dialogue data"). The specific source of the first medical interaction data is not limited here.

[0025] In an implementation scenario, after pre-building a reference feature set, the first medical interaction data can be obtained, and then the answer sentences in the first medical interaction data are segmented (eg, segmented according to punctuation, segmented according to semantics, etc.) to obtain the fragments to be tested.

[0026] In another implementation scenario, as another possible implementation example, different from the aforementioned implementation, after obtaining the first medical interaction data, it is also possible to first segment based on the first medical interaction data to obtain an interaction text sequence, and then perform sliding windows based on the interaction text sequence to obtain several sub-text sequences, and then splice the keywords extracted from the interaction texts at the first and last positions in the sub-text sequence and the interaction texts at positions other than the first and last positions in the sub-text sequence as the fragment to be inspected. The above method, through a series of operations such as data segmentation, sliding window operation, key extraction, and text splicing, can ensure that the fragment to be inspected has relatively complete semantic information and detail information as much as possible, which helps to improve the accuracy of subsequent similarity metrics.

[0027] In a specific implementation scenario, for the first medical interaction data, specifically, it can be split by rows based on the first medical interaction data to obtain an interaction text sequence. For ease of description, the following first medical interaction data is used as an example: The following methods can be used to swallow fish bones: Swallow forcefully.

[0028] Use bread to assist swallowing.

[0029] Drink vinegar to soften it and help it be discharged smoothly.

[0030] Use tweezers to pick it out.

[0031] Splitting the above first medical interaction data by rows can obtain an interaction text sequence {"The following methods can be used to swallow fish bones", "Swallow forcefully", "Use bread to assist swallowing", "Drink vinegar to soften it and help it be discharged smoothly", "Use tweezers to pick it out"}. Of course, the above example is only a possible example of splitting the first medical interaction data by rows to obtain an interaction text sequence in the actual application process, and other possible situations are not exemplified one by one here.

[0032] In a specific implementation scenario, after obtaining the interaction text sequence, a sliding window operation can be performed on it. Exemplarily, the interaction text sequence can be slided with a preset window length and a preset window shift to obtain several sub-text sequences. Taking the preset window length as three interaction texts and the preset window shift as one interaction text as an example, for the aforementioned interaction text sequence, sub-text sequences {"The following methods can be used to swallow fish bones", "Swallow forcefully", "Use bread to assist swallowing"}, sub-text sequences {"Swallow forcefully", "Use bread to assist swallowing", "Drink vinegar to soften it and help it be discharged smoothly"}, and sub-text sequences {"Use bread to assist swallowing", "Drink vinegar to soften it and help it be discharged smoothly", "Use tweezers to pick it out"} can be obtained. Of course, the above example is only a possible example of the sliding window operation, and the window length and window shift of the sliding window operation are not limited here, and other possible situations are not exemplified one by one here.

[0033] In a specific implementation scenario, after obtaining a number of sub-text sequences, key extraction and text splicing can be performed on each sub-text sequence respectively to obtain the corresponding fragment to be inspected. For the sake of convenience, taking the aforementioned sub-text sequence as an example, for the sub-text sequence {"The following methods can be used to swallow fish bones", "Swallow forcefully", "Use bread to assist swallowing"}, the keyword "swallow", "fish bones", "adopt", "method" and "bread", "swallow" of the interactive text located at the beginning and end positions can be extracted. Then, through text splicing, the corresponding fragment to be inspected can be obtained:

swallow|fish bones|adopt|method

bread|swallow

swallow

drink vinegar|soften|drain

bread|swallow

tweezers|pick out

[0034] In a specific implementation scenario, after extracting the fragment to be inspected from the first medical interaction data, feature extraction can be performed on the fragment to be inspected to obtain the first feature representation of the fragment to be inspected. It should be noted that the specific method of extracting the first feature representation can be the same as the specific method of extracting the second feature representation. For example, both can use embedding for vectorization, which is not limited here. In addition, the first feature representation can also be attached with label information, which characterizes the medical environment to which the fragment to be inspected belongs, such as data on departments, scenarios, etc. The specific content of the label information is not limited here.

[0035] Step S12: Based on the similarity between the first feature representation of the fragment to be inspected and the second feature representation of each reference fragment, select the second feature representation as the target feature representation.

[0036] In an implementation scenario, as a possible implementation example, the similarity between the first feature representation of the fragment to be inspected and the second feature representations of each reference fragment can be obtained. Then, the second feature representations can be sorted in descending order of similarity, and the second feature representations before a preset order (e.g., the first three, etc.) can be selected as the target feature representation. Of course, in order to further improve the accuracy of feature screening, the second feature representations with similarity higher than a preset threshold and before the preset order can also be selected as the target feature representation. If no second feature representation with similarity higher than the preset threshold is screened out, it can be considered that there may be no misrepresentation in the fragment to be inspected or that the fragment to be inspected has a new error type (i.e., different from all the error types that have appeared in the reference feature set). In this case, manual rechecking can be performed. When it is found through manual rechecking that there is a new error type, the error fragment in the first medical interaction data to which the fragment to be inspected belongs can be extracted as a new reference fragment according to the acquisition method of the aforementioned reference feature set, and the second feature representation of the new reference fragment can be extracted and incorporated into the reference feature set, so as to support data quality inspection for the new error type based on the updated reference feature set in the subsequent process. Of course, the above examples are only several possible examples of selecting the second feature representation after sorting, and other possible methods of selecting the second feature representation after sorting will not be exemplified one by one here.

[0037] In another implementation scenario, as another possible implementation example, different from the foregoing implementation manner, as mentioned above, both the first feature representation and the second feature representation can also be attached with label information. For the convenience of distinction, the label information attached to the first feature representation can be called the first label, and the label information attached to the second feature representation can be called the second label. Then, after sorting the second feature representations in descending order of similarity and selecting the first candidate representation from the sorted second feature representations, the first candidate representation with the second label being the same as the first label can be selected as the second candidate representation. Then, re-ranking (rerank) can be performed based on the second candidate representation, and the target feature representation can be selected from the sorted second candidate representations. In the above manner, after selecting the first candidate representation according to the similarity, the second candidate representation is further selected from among them according to whether the label information is the same, and the second candidate representation is re-ranked to select the target feature representation from among them. This can improve the accuracy of feature screening through two stages of rough screening and fine screening.

[0038] In a specific implementation scenario, after sorting each second feature representation according to similarity, specifically, among the sorted second feature representations, the second feature representations before the first rank and with similarity higher than a preset threshold can be selected as the first candidate representations. Similarly to the foregoing implementation manner, if no second feature representation with similarity higher than the preset threshold is screened out, it can be considered that the fragment to be inspected may have no misrepresentation or there may be a new error type in the fragment to be inspected (that is, different from all the error types that have appeared in the reference feature set). In this case, manual re-inspection can be performed, and when the manual re-inspection shows that there is a new error type, the error fragment in the first medical interaction data to which the fragment to be inspected belongs, which has been manually re-inspected, can be extracted as a new reference fragment by referring to the acquisition method of the foregoing reference feature set, and the second feature representation of the new reference fragment can be extracted and incorporated into the reference feature set, so as to support subsequent data quality inspection of the new error type based on the updated reference feature set.

[0039] In a specific implementation scenario, in order to re-rank each second candidate representation, a re-rank model fine-tuned for the medical direction with respect to BGE-reranker can be used to complete the re-ranking of each second candidate representation. For the technical details of BGE-reranker, please refer to the relevant content, which will not be elaborated here.

[0040] In a specific implementation scenario, after re-ranking each second candidate representation, specifically, among the re-ranked second candidate representations, the second candidate representations before the second rank (such as the first three, etc.) can be selected as the target feature representations.

[0041] Step S13: Obtain the quality inspection result of the fragment to be inspected based on the attached information of the target feature representation.

[0042] Specifically, in the case where a target feature representation is selected, the quality inspection result of the to-be-inspected segment can be obtained based on the accompanying information of the target feature representation. It should be noted that in the case where the accompanying information includes the reason for the error, the quality inspection result of the to-be-inspected segment can specifically include: there is an expression error in the to-be-inspected segment, and the reason for the error of the to-be-inspected segment is the reason for the error attached to the target feature representation; or, in the case where the accompanying information further includes the second medical interaction data from which the reference segment is derived, the quality inspection result of the to-be-inspected segment can, in addition to including the foregoing content, further include similar medical interaction data with the same error type as the to-be-inspected segment, that is, the second medical interaction data from which the reference segment to which the target feature representation belongs is derived. Of course, as described above, if no target feature representation is selected, the quality inspection result of the to-be-inspected segment can include: the to-be-inspected segment may not have an expression error or the to-be-inspected segment has a new error type (that is, different from all the error types that have appeared in the reference feature set). In this case, the to-be-inspected segment can be further manually reinspected. For specific details, reference can be made to the foregoing relevant description, which will not be elaborated here. In addition, as a possible example, after each to-be-inspected segment in the first medical interaction data has passed the quality inspection, the first medical interaction data can be used to train a medical interaction model. That is to say, for any first medical interaction data, if there are no error expressions after each to-be-inspected segment in it has passed the quality inspection respectively, this first medical interaction data can be used to train a medical interaction model.

[0043] To facilitate a comprehensive understanding of a possible implementation example of the embodiments of the present disclosure, please refer to Figure 2 , Figure 2 which is a schematic diagram of the process of an embodiment of the quality inspection method for medical interaction data of this application. As Figure 2As shown in the figure, on the one hand, new incoming medical interaction data can be obtained as the second medical interaction data, and the second medical interaction data can be quality-checked based on a data quality-check model to obtain text fragments with incorrect expressions in the second medical interaction data as candidate fragments. The data quality-check model is trained based on sample medical interaction data marked with incorrect fragments. On the other hand, sample medical interaction data marked with incorrect fragments can be obtained as the second medical interaction data, and the incorrect fragments marked in the second medical interaction data can be selected as candidate fragments. Based on this, the candidate fragments can be semantically completed with reference to the second medical interaction data based on a semantic completion model to obtain reference fragments. By vectorizing the reference fragments, a second feature representation of the reference fragments can be obtained, and a reference feature set can be constructed accordingly. On this basis, for the first medical interaction data to be quality-checked, it can be first split by row to obtain an interaction text sequence, then a sliding window can be performed on the interaction text sequence to obtain several sub-text sequences, and finally, the keyword and phrase extracted from the interaction text at the beginning and end positions in the sub-text sequence and the interaction text other than the beginning and end positions in the sub-text sequence are spliced as the fragment to be inspected. By vectorizing the fragment to be inspected, a first feature representation of the fragment to be inspected can be obtained. Then, the similarity between the first feature representation of the fragment to be inspected and the second feature representation of each reference fragment can be obtained, and the second feature representations can be sorted in descending order of similarity. The first candidate representation can be selected from the sorted second feature representations, and then the first candidate representation with the second label consistent with the first label can be selected as the second candidate representation. It should be noted that the first feature representation is also attached with a first label characterizing the medical environment to which the fragment to be inspected belongs, and the second feature representation is also attached with a second label characterizing the medical environment to which the reference fragment belongs. Finally, re-sorting can be performed based on the second candidate representation, and the target feature representation can be selected from the re-sorted second candidate representations. Furthermore, based on the attached information of the target feature representation, the quality-check result of the fragment to be inspected can be obtained.

[0044] In the above solution, a reference feature set is obtained, and a first feature representation of each to-be-inspected segment in the first medical interaction data is obtained. The reference feature set includes second feature representations of several reference segments, and the second feature representation is at least attached with: the error reason of the reference segment. Then, based on the similarity between the first feature representation of the to-be-inspected segment and the second feature representations of each reference segment, the second feature representation is selected as the target feature representation, so as to obtain the quality inspection result of the to-be-inspected segment based on the attached information of the target feature representation. Therefore, on the one hand, it can realize the quality inspection of medical interaction data through feature similarity comparison and improve the efficiency of data quality inspection. On the other hand, by constructing a reference feature set involving various error reasons for feature comparison, it can systematically scan the entire to-be-inspected data set through known error types to detect whether there are similar errors, and can ensure that all medical interaction data with the same error type are accurately detected as much as possible, improving the accuracy of data quality inspection. In addition, when it is necessary to support the detection of new error types, only the reference segment involving the new error type and its second feature representation need to be added to the reference feature set, which helps to achieve efficient expansion of error types, reduce the need for manual intervention, and thus reduce the data quality inspection cost.

[0045] Please refer to Figure 3 , Figure 3 FIG. is a schematic framework diagram of an embodiment of the medical interaction data quality inspection device of the present application. The medical interaction data quality inspection device 30 includes: an acquisition module 31, a selection module 32, and a quality inspection module 33. The acquisition module 31 is configured to obtain a reference feature set and obtain a first feature representation of each to-be-inspected segment in the first medical interaction data; wherein, the reference feature set includes second feature representations of several reference segments, and the second feature representation is at least attached with: the error reason of the reference segment; the selection module 32 is configured to select the second feature representation as the target feature representation based on the similarity between the first feature representation of the to-be-inspected segment and the second feature representations of each reference segment; the quality inspection module 33 is configured to obtain the quality inspection result of the to-be-inspected segment based on the attached information of the target feature representation.

[0046] In the above solution, the medical interaction data quality inspection device 30 obtains a reference feature set and obtains the first feature representation of each fragment to be inspected in the first medical interaction data. The reference feature set includes the second feature representations of several reference fragments, and the second feature representation is at least attached with: the error reason of the reference fragment. Then, based on the similarity between the first feature representation of the fragment to be inspected and the second feature representations of each reference fragment, the second feature representation is selected as the target feature representation, and based on the attached information of the target feature representation, the quality inspection result of the fragment to be inspected is obtained. Therefore, on the one hand, it can realize the quality inspection of medical interaction data through feature similarity comparison and improve the efficiency of data quality inspection. On the other hand, by constructing a reference feature set involving various error reasons for feature comparison, it can systematically scan the entire dataset to be inspected through known error types to detect whether there are similar errors, and can ensure that all medical interaction data with the same error type are accurately detected as much as possible, and can improve the accuracy of data quality inspection. In addition, when it is necessary to support the detection of new error types, it only needs to add the reference fragment with the new error type and its second feature representation to the reference feature set, which helps to achieve the efficient expansion of error types, reduce the need for manual intervention, and thus can reduce the data quality inspection cost.

[0047] In some disclosed embodiments, the first feature representation is also attached with a first label characterizing the medical environment to which the fragment to be inspected belongs, and the second feature representation is also attached with a second label characterizing the medical environment to which the reference fragment belongs. The selection module 32 includes a sorting and selection sub-module, which is used to sort the second feature representations in descending order of similarity and select the first candidate representation from the sorted second feature representations; the selection module 32 includes a feature selection sub-module, which is used to select the first candidate representation with the second label consistent with the first label as the second candidate representation; the selection module 32 includes a re-sorting and selection sub-module, which is used to re-sort based on the second candidate representation and select the target feature representation from the re-sorted second candidate representations.

[0048] In some disclosed embodiments, the sorting and selection sub-module is specifically used to select, from the sorted second feature representations, the second feature representations before the first order and with a similarity higher than a preset threshold as the first candidate representation.

[0049] In some disclosed embodiments, the re-sorting and selection sub-module is specifically used to select, from the re-sorted second candidate representations, the second candidate representations before the second order as the target feature representations.

[0050] In some disclosed embodiments, the obtaining module 31 includes a data splitting sub-module, configured to split based on the first medical interaction data to obtain an interaction text sequence; the obtaining module 31 includes a sequence sliding window sub-module, configured to perform a sliding window on the interaction text sequence to obtain a plurality of sub-text sequences; the obtaining module 31 includes an extraction and splicing sub-module, configured to splice based on the keyword phrases extracted from the interaction texts at the head and tail positions in the sub-text sequences and the interaction texts other than those at the head and tail positions in the sub-text sequences as the to-be-inspected segment.

[0051] In some disclosed embodiments, the data splitting sub-module is specifically configured to perform line-by-line splitting based on the first medical interaction data to obtain an interaction text sequence.

[0052] In some disclosed embodiments, the sequence sliding window sub-module is specifically configured to perform a sliding window on the interaction text sequence with a preset window length and a preset window shift to obtain a plurality of sub-text sequences; wherein, the preset window length is three interaction texts, and the preset window shift is one interaction text.

[0053] In some disclosed embodiments, the obtaining module 31 includes a segment obtaining sub-module, configured to obtain the text segments with expression errors in the second medical interaction data as candidate segments; the obtaining module 31 includes a semantic completion sub-module, configured to perform semantic completion on the candidate segments based on the second medical interaction data to obtain reference segments.

[0054] In some disclosed embodiments, the segment obtaining sub-module is specifically configured to perform at least one of the following: obtain newly warehoused medical interaction data as the second medical interaction data, and perform quality inspection on the second medical interaction data based on a data quality inspection model to obtain the text segments with expression errors in the second medical interaction data as candidate segments; obtain the sample medical interaction data marked with error segments as the second medical interaction data, and select the marked error segments in the second medical interaction data as candidate segments; wherein, the data quality inspection model is trained based on the sample medical interaction data marked with error segments.

[0055] In some disclosed embodiments, the semantic completion sub-module is specifically configured to perform semantic completion on the candidate segments based on a semantic completion model with reference to the second medical interaction data to obtain reference segments; wherein, the semantic completion model is trained based on sample segment pairs, the sample segment pairs include a first segment with incomplete semantics and a second segment after semantic completion of the first segment, and the sample segment pairs are generated by a large language model.

[0056] In some disclosed embodiments, the first medical interaction data is generated by an artificial intelligence model; and / or, the first medical interaction data is collected on a medical consultation platform; and / or, after each to-be-tested segment in the first medical interaction data has been quality-checked, the first medical interaction data is used to train a medical interaction model; and / or, the reference segment is derived from an erroneous text segment in the second medical interaction data, and the second feature representation is also accompanied by: the second medical interaction data from which the reference segment is derived.

[0057] See also Figure 4 , Figure 4 : is a schematic diagram of the framework of an embodiment of an electronic device of the present application. The electronic device 40 includes at least a memory 41 and a processor 42 coupled to each other, the memory 41 stores at least program instructions, and the processor 42 is used to execute the program instructions to implement the steps in any of the above-mentioned medical interaction data quality inspection method embodiments. For details, please refer to the aforementioned disclosed embodiments, which will not be repeated here. As a possible example, the electronic device 40 may include but is not limited to a smart phone, a tablet computer, a server, etc., and the specific type of the electronic device 40 is not limited here.

[0058] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned medical interaction data quality inspection method embodiments. The processor 42 can also be called a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 42 can be implemented by an integrated circuit chip.

[0059] In the above solution, the electronic device 40 obtains a reference feature set and obtains the first feature representation of each to-be-inspected segment in the first medical interaction data. The reference feature set contains the second feature representations of several reference segments, and the second feature representation is at least attached with: the error reason of the reference segment. Then, based on the similarity between the first feature representation of the to-be-inspected segment and the second feature representations of each reference segment, the second feature representation is selected as the target feature representation, so as to obtain the quality inspection result of the to-be-inspected segment based on the attached information of the target feature representation. Therefore, on the one hand, it can realize the quality inspection of medical interaction data through feature similarity comparison and improve the efficiency of data quality inspection. On the other hand, by constructing a reference feature set involving various error reasons for feature comparison, it can systematically scan the entire to-be-inspected data set through known error types to detect whether there are similar errors, and can ensure that all medical interaction data with the same error type are accurately detected as much as possible, thereby improving the accuracy of data quality inspection. In addition, when it is necessary to support the detection of new error types, only the reference segments involving the new error type and their second feature representations need to be added to the reference feature set, which helps to achieve efficient expansion of error types, reduce the need for manual intervention, and thus can reduce the data quality inspection cost.

[0060] Please refer to Figure 5 , Figure 5 FIG. is a schematic framework diagram of an embodiment of the computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be run by a processor, and the program instructions 51 are used to implement the steps in any of the above embodiments of the medical interaction data quality inspection method.

[0061] In the above solution, the computer-readable storage medium 50 obtains a reference feature set and obtains the first feature representation of each fragment to be inspected in the first medical interaction data. The reference feature set contains the second feature representations of several reference fragments, and the second feature representation is at least attached with: the error reason of the reference fragment. Then, based on the similarity between the first feature representation of the fragment to be inspected and the second feature representations of each reference fragment, the second feature representation is selected as the target feature representation, and based on the attached information of the target feature representation, the quality inspection result of the fragment to be inspected is obtained. Therefore, on the one hand, it can perform quality inspection on medical interaction data through feature similarity comparison and improve the efficiency of data quality inspection. On the other hand, by constructing a reference feature set involving various error reasons for feature comparison, it can systematically scan the entire dataset to be inspected through known error types to detect whether there are similar errors, and can ensure that all medical interaction data with the same error type is accurately detected as much as possible, improving the accuracy of data quality inspection. In addition, when it is necessary to support the detection of new error types, only the reference fragments involving the new error type and their second feature representations need to be added to the reference feature set, which helps to achieve efficient expansion of error types, reduce the need for manual intervention, and thus reduce the cost of data quality inspection.

[0062] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0063] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0064] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0065] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0066] In addition, in each embodiment of the present application, each functional unit may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0067] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0068] If the technical solution of the present application involves personal information, before the product applying the technical solution of the present application processes personal information, it has clearly informed the personal information processing rules and obtained the individual's independent consent. If the technical solution of the present application involves sensitive personal information, before the product applying the technical solution of the present application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirement of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent logo is set to inform that the personal information collection range has been entered and personal information will be collected. If an individual voluntarily enters the collection range, it is considered consent to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are informed by obvious logos / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information themselves; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

Claims

1. A medical interaction data quality inspection method, characterized in that, include: Acquire a reference feature set, and acquire a first feature representation of each to-be-tested segment in the first medical interaction data; wherein the reference feature set includes a second feature representation of a plurality of reference segments, and the second feature representation is accompanied by at least: an error cause of the reference segment; Based on the similarity between the first feature representation of the to-be-detected segment and the second feature representation of each of the reference segments, selecting the second feature representation as the target feature representation; Based on the accompanying information represented by the target feature, a quality inspection result of the fragment to be inspected is obtained.

2. The method according to claim 1, wherein The first feature representation is also accompanied by a first label representing the medical environment to which the to-be-tested segment belongs, the second feature representation is also accompanied by a second label representing the medical environment to which the reference segment belongs, and the second feature representation is selected as the target feature representation based on the similarity between the first feature representation of the to-be-tested segment and the second feature representations of each of the reference segments, including: Sorting the second feature representations in descending order of the similarity, and selecting a first candidate representation from the sorted second feature representations; Selecting a first candidate representation whose second label is consistent with the first label as a second candidate representation; Reordering is performed based on the second candidate representations, and the target feature representation is selected from each of the second candidate representations after the reordering.

3. The method according to claim 2, characterized in that The selecting the first candidate representation from each of the second feature representations after sorting includes: selecting, from each of the second feature representations after sorting, a second feature representation before the first rank and with a similarity higher than a preset threshold as the first candidate representation; And / or, selecting the target feature representation from each of the second candidate representations after reordering includes: selecting the second candidate representation before the second rank as the target feature representation from each of the second candidate representations after reordering.

4. The method according to claim 1, wherein The step of obtaining the fragment to be inspected comprises: Segmenting based on the first medical interaction data to obtain an interaction text sequence; Perform sliding window based on the interactive text sequence to obtain a plurality of sub-text sequences; Keywords extracted from the interactive texts at the first and last positions in the sub-text sequence and the interactive texts other than the first and last positions in the sub-text sequence are spliced to serve as the segment to be inspected.

5. The method according to claim 3, characterized in that, The segmenting based on the first medical interaction data to obtain the interactive text sequence includes: segmenting by line based on the first medical interaction data to obtain the interactive text sequence; And / or, sliding the window based on the interactive text sequence to obtain a plurality of sub-text sequences includes: sliding the window of the interactive text sequence based on a preset window length and a preset window shift to obtain a plurality of sub-text sequences; wherein the preset window length is three of the interactive texts, and the preset window shift is one of the interactive texts.

6. The method according to claim 1, wherein The step of obtaining the reference fragment comprises: Acquire a text segment with an incorrect expression in the second medical interaction data as a candidate segment; The candidate segment is semantically completed based on the second medical interaction data to obtain the reference segment.

7. The method according to claim 6, wherein The step of obtaining a text segment with incorrect expressions in the second medical interaction data as a candidate segment includes at least one of the following: Acquire newly stored medical interaction data as the second medical interaction data, and perform quality inspection on the second medical interaction data based on a data quality inspection model to obtain text segments with incorrect expressions in the second medical interaction data as the candidate segments; Acquire sample medical interaction data marked with error segments as the second medical interaction data, and select the error segments marked in the second medical interaction data as the candidate segments; The data quality inspection model is trained based on sample medical interaction data that are annotated with erroneous fragments.

8. The method according to claim 6, wherein The semantically completing the candidate segment based on the second medical interaction data to obtain the reference segment includes: Based on the semantic completion model, the candidate segment is semantically completed with reference to the second medical interaction data to obtain the reference segment; wherein the semantic completion model is trained based on a sample segment pair, the sample segment pair includes a semantically incomplete first segment and a second segment after the first segment is semantically completed, and the sample segment pair is generated by a large language model.

9. The method according to any one of claims 1 to 8, characterized in that, The first medical interaction data is generated by an artificial intelligence model; And / or, the first medical interaction data is collected on a medical consultation platform; and / or, after each of the to-be-checked segments in the first medical interaction data is quality-checked and found to be correct, the first medical interaction data is used to train a medical interaction model; And / or, the reference segment originates from a text segment with an error in the second medical interaction data, and the second feature representation is further accompanied by: the second medical interaction data from which the reference segment originates.

10. A medical interaction data quality inspection device, characterized in that, include: An acquisition module is used to acquire a reference feature set and acquire a first feature representation of each to-be-tested segment in the first medical interaction data; wherein the reference feature set includes a second feature representation of a plurality of reference segments, and the second feature representation is accompanied by at least: an error cause of the reference segment; A selection module, configured to select the second feature representation as a target feature representation based on the similarity between the first feature representation of the segment to be detected and the second feature representation of each reference segment; A quality inspection module is used to obtain a quality inspection result of the to-be-inspected segment based on the accompanying information represented by the target feature.

11. An electronic device, characterized in that, The medical device comprises at least a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the medical interaction data quality inspection method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the medical interaction data quality inspection method according to any one of claims 1 to 9.