Associated data analysis method for ear-nose-throat symptoms, terminal and storage medium
By determining the analysis method by the number and proportion of anchor words and combining data conflict degree and vector similarity, the problem of low data analysis efficiency in the existing technology is solved, and more efficient and accurate data acquisition is achieved.
Patent Information
- Application Number
- CN202510840655.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In the prior art, data acquisition is performed only based on the similarity of vectorized data, resulting in poor efficiency in data analysis and processing.
By determining the number and proportion of anchor words, relevant data is obtained through data conflict analysis or vector similarity analysis. The target patient text is converted into a high-dimensional vector in combination with the pre-trained model, and relevant data is obtained based on vector similarity.
It improves the efficiency and accuracy of data processing, avoids the shortcomings of a single analysis method, meets actual analysis needs, and enhances the pertinence and accuracy of data acquisition.
Smart Images

Figure CN120724973A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis, and in particular to a method, terminal and storage medium for analyzing associated data of ear, nose and throat symptoms. Background Art
[0002] With the rapid development of medical informatization, artificial intelligence and natural language processing technologies are increasingly being applied in the medical field. Intelligent medical systems can automatically retrieve relevant data within data sharing platforms based on a patient's symptom description, thereby providing effective diagnostic and treatment assistance. However, faced with the massive amount of data within these platforms, search efficiency is often hampered. Therefore, how to quickly obtain relevant data on symptoms is a pressing issue for those skilled in the art.
[0003] Chinese patent publication number CN119446395A discloses a patient diagnosis and treatment method and medium based on a vector database, which relates to the field of intelligent medical technology, including: collecting patient medical data, and vectorizing the medical data including digital information, text information and image information, then performing feature fusion on the vectorized data to obtain a first fusion vector, calculating the similarity between the first fusion vector and the second fusion vector corresponding to each historical patient, and obtaining a preset number of second fusion vectors with the highest similarity; based on the preset number of second fusion vectors with the highest similarity, obtaining the best treatment plan. It can be seen that the above technical solution has the following problems: data acquisition is performed only based on the similarity of the vectorized data, and the data acquisition method is relatively simple, which can easily lead to poor data analysis and processing efficiency. Summary of the Invention
[0004] To this end, the present invention provides a method, terminal and storage medium for analyzing associated data of ENT symptoms, so as to overcome the problem in the prior art that data is acquired only based on the similarity of vectorized data, the data acquisition method is relatively simple, and thus easily leads to poor efficiency in data analysis and processing.
[0005] To achieve the above object, the present invention provides a method for analyzing association data of ENT symptoms, comprising: Determine the anchor word status of the target patient text based on the number of anchor words in a category and the proportion of anchor words in a category; The target analysis method is determined based on the anchor word status of the target patient text, which is to analyze the data conflict degree of a class of anchor words or perform vector similarity analysis; Under the first information acquisition condition, determining the data acquisition method to acquire relevant data based on text similarity or anchor word combination according to the comparison result of the data conflict degree of a type of anchor word with the preset data conflict degree; Under the second information acquisition condition, the target patient text is converted into a high-dimensional vector through the pre-training model, and relevant data is obtained based on the vector similarity; Send relevant data to the client.
[0006] Furthermore, the target analysis method is determined based on the anchor word status of the target patient text; If the anchor word status is that the number of anchor words in one category is greater than the preset number of anchor words in one category and the proportion of anchor words in one category is greater than the preset proportion of anchor words in one category, the target analysis method is to analyze the data conflict degree of anchor words in one category; If the anchor word status is that the number of anchor words in one category is less than or equal to the preset number of anchor words in one category or the proportion of anchor words in one category is less than or equal to the preset proportion of anchor words in one category, the target analysis method is vector similarity analysis.
[0007] Furthermore, the anchor word categories include first-class anchor words and second-class anchor words; One type of anchor words is anchor words whose combined usage richness is greater than the preset combined usage richness and whose data difference is less than or equal to the preset data difference; The second type of anchor words are anchor words whose combined usage richness is less than or equal to the preset combined usage richness or whose data difference is greater than the preset data difference.
[0008] Furthermore, under the first information acquisition condition, the data conflict degree of a type of anchor words corresponding to the target patient text is detected, and the data acquisition method is determined according to the data conflict degree of the type of anchor words; If the data conflict degree of a type of anchor word is less than the preset data conflict degree, the data acquisition method is to obtain relevant data based on text similarity; If the data conflict degree of a type of anchor words is greater than or equal to the preset data conflict degree, the data acquisition method is to perform anchor word combination for the type of anchor words and acquire relevant data based on the anchor word combination; The text similarity is determined based on the number of similarities between a class of anchor words; The first information acquisition condition is to determine that the target analysis method is to analyze the data conflict degree of a type of anchor words.
[0009] Furthermore, the data conflict degree of a type of anchor words corresponding to the target patient text is confirmed as follows: Detect the feature data label values corresponding to all anchor words in the target patient text; Calculate the difference between the feature data label value and the preset feature data label value difference as the data conflict degree; The difference between the feature data label values is denoted as S; Among them, Hi is the feature data label value corresponding to the i-th first-class anchor word in the target patient text, H0 is the average value of the feature data label values corresponding to all first-class anchor words, and n is the total number of first-class anchor words in the target patient text.
[0010] Furthermore, the process of obtaining relevant data based on the anchor word combination includes: Acquire preliminary relevant data based on each anchor word combination, detect the data volume of the preliminary relevant data, and if the data volume of the preliminary relevant data is in a low-order data volume range, increase processing is performed on the anchor word combination; If the amount of the prepared relevant data is within the high-order data amount range, the prepared relevant data is reduced according to the anchor word priority coefficient; If the amount of the prepared relevant data is within the required data amount range, the prepared relevant data is recorded as relevant data and the relevant data corresponding to each anchor word combination is transmitted to the client respectively.
[0011] Furthermore, under the second information acquisition condition, the target patient text is converted into a high-dimensional vector through a pre-trained model, and the first several historical patient texts in descending order of vector similarity with the target patient text, the historical patient image data corresponding to the historical patient texts, and the diagnosis and treatment records corresponding to the historical patient texts are recorded as relevant data and transmitted to the client; The second information acquisition condition is to determine that the target analysis method is vector similarity analysis.
[0012] The present invention also provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor, characterized in that the processor implements the associated data analysis method of the ENT symptoms when executing the computer program.
[0013] The present invention also provides a computer-readable storage medium, which stores a computer program, and is characterized in that when the computer program is executed by a processor, it implements the associated data analysis method of the ear, nose and throat symptoms.
[0014] Compared with the prior art, the beneficial effect of the present invention lies in that, in the technical solution of the present invention, the anchor word category is determined according to the corresponding common use richness and data difference of the anchor word, the common use richness is used to express the degree of common use of the anchor word with other anchor words, and the criticality of the anchor word to the patient image data is determined by the data difference, so that the subsequent analysis of the anchor word is more targeted, avoiding the problem of poor analysis effect caused by the unified vectorized text method in the prior art, thereby improving the subsequent data processing efficiency and accuracy of the present invention.
[0015] Furthermore, in the technical solution of the present invention, the anchor word status of the target patient text is determined based on the number of a type of anchor words and the proportion of a type of anchor words. The anchor word status reflects whether a type of anchor words in the target patient text can meet subsequent analysis requirements, and the target analysis method is determined according to the anchor word status of the target patient text to analyze the data conflict degree of a type of anchor words or vector similarity analysis, so that the target analysis method is more in line with the actual scenario, avoiding the problem that a single analysis method is difficult to meet the actual analysis needs, thereby improving the data analysis accuracy of the present invention.
[0016] Furthermore, under the first information acquisition condition in the technical solution of the present invention, the data acquisition method is determined to be acquiring relevant data based on text similarity or anchor word combination according to the comparison result of the data conflict degree of a type of anchor words and the preset data conflict degree, and the data conflict degree is used to reflect the degree of conflict of the corresponding historical data between the anchor words, and then reflect the accuracy of unified data acquisition using a type of anchor words, and correspondingly select different data acquisition methods, so that the selection of data acquisition method is more in line with actual needs, and at the same time avoids the problem that the acquired data is difficult to meet usage needs due to the single data acquisition method in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram of the method for analyzing association data of ear, nose and throat symptoms of the present invention; Figure 2 This is a flow chart of the present invention for determining a target analysis method based on the anchor word status of a target patient text; Figure 3 The present invention is a flow chart of determining a data acquisition method based on a comparison result of the data conflict degree of a type of anchor word with a preset data conflict degree. DETAILED DESCRIPTION
[0018] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0019] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0020] It should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0021] See also Figures 1 to 3 As shown, the present invention provides a method for analyzing association data of ear, nose and throat symptoms, comprising: Determine the anchor word status of the target patient text based on the number of anchor words in a category and the proportion of anchor words in a category; The target analysis method is determined based on the anchor word status of the target patient text, which is to analyze the data conflict degree of a class of anchor words or perform vector similarity analysis; Under the first information acquisition condition, determining the data acquisition method to acquire relevant data based on text similarity or anchor word combination according to the comparison result of the data conflict degree of a type of anchor word with the preset data conflict degree; Under the second information acquisition condition, the target patient text is converted into a high-dimensional vector through the pre-training model, and relevant data is obtained based on the vector similarity; Send relevant data to the client.
[0022] The present invention is applied to a patient text analysis platform for ENT symptoms, corresponding to the acquisition of associated relevant data to provide medical assistance to patients or doctors. The target patient text is the patient text of the patient with ENT symptoms that currently needs to be analyzed, uploaded by the platform's client. The target patient text is a text describing the patient's symptoms, and the target patient text corresponds to target patient image data. The target patient image data is at least one of the patient's ear endoscope image, nasal endoscope image, and laryngeal endoscope image. The present invention is applied with a database for storing a number of anchor words, historical patient texts, historical patient image data corresponding to historical patient texts, and medical records corresponding to historical patient texts. The specific content of the medical records is written by actual doctors, which is easily understood by those skilled in the art and will not be elaborated on here. The anchor words are keywords in the target patient text that are statistically set by the user.
[0023] The present invention also applies several historical records, and any historical record corresponds to the number of a type of anchor words, the proportion of a type of anchor words, and usage richness, data difference, feature data label value difference and data conflict, and the historical record is correspondingly set with a qualified mark, which records whether the historical record meets the user's needs. Among them, whether a historical record meets the user's needs can be determined based on, but not limited to, user ratings. It can be understood that using a self-established rating mechanism as an indicator of whether the obtained relevant data meets the user's needs to determine whether the historical record meets the needs is content that has been mastered by those skilled in the art and is not limited here.
[0024] Specifically, the target analysis method is determined based on the anchor word status of the target patient text; If the anchor word status is that the number of anchor words in one category is greater than the preset number of anchor words in one category and the proportion of anchor words in one category is greater than the preset proportion of anchor words in one category, the target analysis method is to analyze the data conflict degree of anchor words in one category; If the anchor word status is that the number of anchor words in one category is less than or equal to the preset number of anchor words in one category or the proportion of anchor words in one category is less than or equal to the preset proportion of anchor words in one category, the target analysis method is vector similarity analysis.
[0025] The number of anchor words of a category is the number of different anchor words of a category in the target patient text. It should be noted that the number is not the number of occurrences. If there is an anchor word A, and the anchor word A appears 3 times in the target patient text, the number is also 1.
[0026] The proportion of first-category anchor words = the number of first-category anchor words in the target patient text / the number of second-category anchor words; The values of the preset number of anchor words in a class and the preset proportion of anchor words in a class can be set by the user according to the actual application scenario. It can be understood that the larger the number of anchor words in a class and the proportion of anchor words in a class are, the greater the contribution of anchor words in a class to subsequent analysis. The historical records corresponding to the number of anchor words in a class and the proportion of anchor words in a class that meet the needs can be extracted, and the outliers in the number of anchor words in a class and the proportion of anchor words in a class can be removed respectively. The average values corresponding to the number of anchor words in a class and the proportion of anchor words in a class after removing the outliers are recorded as the preset number of anchor words in a class and the preset proportion of anchor words in a class respectively. The method of removing outliers includes but is not limited to the 3σ criterion method or the IQR method.
[0027] Specifically, the anchor word categories include first-class anchor words and second-class anchor words; One type of anchor words is anchor words whose combined usage richness is greater than the preset combined usage richness and whose data difference is less than or equal to the preset data difference; The second type of anchor words are anchor words whose combined usage richness is less than or equal to the preset combined usage richness or whose data difference is greater than the preset data difference.
[0028] For a single anchor word, the anchor word is recorded as the target anchor word. The method for confirming the richness of the target anchor word is to randomly extract a number of historical patient texts corresponding to the target anchor word. The historical patient texts corresponding to the target anchor word are the historical patient texts in which the target anchor word exists. The specific number of historical patient texts corresponding to the target anchor word is set by the user. The greater the user's demand for the analysis accuracy of the anchor word category, the larger the number of extractions. The number M of other anchor words other than the target anchor word in each historical patient text is detected, and the average value of M corresponding to each historical patient text is calculated. is recorded as the target anchor word's usage richness, where M is not the number of occurrences, but the number of different anchor words in the historical patient text except the target anchor word. Similarly, if there is an anchor word B, and the anchor word B appears 3 times in a historical patient text, the contribution number of M to the historical patient text is still 1. The data difference of the target anchor word is confirmed by detecting the historical patient image data corresponding to each historical patient text corresponding to the target anchor word, and calculating the ear endoscope image difference, nasal endoscope image difference and laryngeal endoscope image difference corresponding to each historical patient text. The average value is recorded as the data difference. The principles of the confirmation methods for the ear endoscope image difference, nasal endoscope image difference and laryngeal endoscope image difference corresponding to several historical patient texts are the same. For example, for the ear endoscope image difference corresponding to several historical patient texts, the area of the marked area corresponding to each ear endoscope image and the number of marked areas corresponding to each ear endoscope image are detected. Ear endoscope image difference = (maximum marked area - minimum marked area) / preset marked area + maximum number of marked areas - minimum number of marked areas) / preset mark The number of regions, the area of the preset marked regions is the average of the areas of the marked regions corresponding to the ear endoscope images, and the number of preset marked regions is the average of the number of marked regions corresponding to the ear endoscope images. It can be understood that the number of marked regions of the ear endoscope images corresponding to different diseases may be different. The area of the marked regions corresponding to a single ear endoscope image is the total area of the marked regions in the ear endoscope image. The calculation method of the difference degree of other nasal endoscope images and the difference degree of laryngeal endoscope images is the same as that of the ear endoscope image, which is easy to understand for those skilled in the art and will not be elaborated here.
[0029] The values of the preset joint use richness and the preset data difference can be set by the user according to the actual application scenario. It can be understood that the greater the user's demand for the accuracy of the subsequent process of using a class of anchor words to determine the text similarity and obtain relevant data based on the text similarity, the greater the value of the preset joint use richness and the smaller the value of the preset data difference. A value selection method is provided to extract the joint use richness and data difference corresponding to the historical records that meet the requirements, remove the outliers in the joint use richness and data difference respectively, and record the corresponding average values of the joint use richness and data difference after removing the outliers as the preset joint use richness and preset data difference respectively.
[0030] Specifically, under the first information acquisition condition, the data conflict degree of a type of anchor words corresponding to the target patient text is detected, and the data acquisition method is determined according to the data conflict degree of the type of anchor words; If the data conflict degree of a type of anchor word is less than the preset data conflict degree, the data acquisition method is to obtain relevant data based on text similarity; If the data conflict degree of a type of anchor words is greater than or equal to the preset data conflict degree, the data acquisition method is to perform anchor word combination for the type of anchor words and acquire relevant data based on the anchor word combination; The text similarity is determined based on the number of similarities between a class of anchor words; The first information acquisition condition is to determine that the target analysis method is to analyze the data conflict degree of a type of anchor words.
[0031] The data acquisition method is to obtain relevant data based on text similarity, detect historical patient texts within a preset time period before the current moment, detect historical patient texts whose text similarity with the target patient text is greater than the preset text similarity, record the detected historical patient texts and the historical patient image data corresponding to the historical patient texts and the medical records corresponding to the historical patient texts as relevant data, and transmit the relevant data to the client separately; the text similarity between the target patient text and any historical patient text = the number of the same type of anchor words between the target patient text and the historical patient text / the number of the same type of anchor words in the target patient text, for example, the target patient text There is a class of anchor words A, B, C, and D in the memory. There is a class of anchor words A, B, and C in the historical patient text. The text similarity between the target patient text and the historical patient text is 3 / 4. It is worth noting that whether the number of times the anchor words A, B, C, and D appear is greater than 1 will not affect the text similarity. The length of the preset time period and the value of the preset text similarity are set by the user. The greater the user's demand for the amount of relevant data obtained, the longer the length of the preset time period. The greater the user's demand for the accuracy of the relevant data, the greater the value of the preset text similarity. A value of preset text similarity is provided, and the preset text similarity is 6 / 10.
[0032] Specifically, the data conflict degree of a type of anchor words corresponding to the target patient text is confirmed as follows: Detect the feature data label values corresponding to all anchor words in the target patient text; Calculate the difference between the feature data label value and the preset feature data label value difference as the data conflict degree; The difference between the feature data label values is denoted as S; Among them, Hi is the feature data label value corresponding to the i-th first-class anchor word in the target patient text, H0 is the average value of the feature data label values corresponding to all first-class anchor words, and n is the total number of first-class anchor words in the target patient text. It can be understood that the order of the values of i corresponding to the first-class anchor words will not affect the calculation results.
[0033] The difference between the characteristic data label value difference and the preset characteristic data label value difference is the value obtained by subtracting the preset characteristic data label value difference from the characteristic data label value difference; The method for confirming the feature data label value corresponding to a single type of anchor word is to detect the historical patient texts containing the type of anchor word within a preset time period before the current moment, extract the ear endoscope images, nasal endoscope images, and laryngeal endoscope images corresponding to each historical patient text, detect the feature data label value corresponding to the endoscope image with the largest number, and record it as the feature data label value corresponding to the type of anchor word. For example, the numbers of ear endoscope images, nasal endoscope images, and laryngeal endoscope images are 10, 8, and 5, respectively, then the feature data label value corresponding to the ear endoscope image is recorded as the feature data label value corresponding to the type of anchor word. If there are at least two types of endoscope images whose corresponding numbers are the maximum, then the length of the preset time period is increased by 15 days until there is only one type of endoscope image whose corresponding number is the maximum. The feature data label value corresponding to any type of endoscopic image = the average value of the aggregation area ratios corresponding to each endoscopic image of that type. The aggregation area ratio corresponding to the endoscopic image = the aggregation area of the corresponding marked area / the area of the endoscopic image. The aggregation area of the marked area is confirmed by constructing a minimum circle that can include the marked area corresponding to the endoscopic image, and recording the area of the circle as the corresponding aggregation area of the marked area. For example, the feature data label value corresponding to the ear endoscope image = the average value of the aggregation area ratios corresponding to all ear endoscope images. For a single ear endoscope image, its corresponding aggregation area ratio = the aggregation area of the marked area corresponding to the ear endoscope image / the area of the ear endoscope image; The values of the preset feature data label value difference and the preset data conflict degree can be set by the user according to actual application requirements. It can be understood that the greater the user's demand for the accuracy of obtaining relevant data based on text similarity, the smaller the preset data conflict degree and the preset feature data label value difference. A value selection method is provided to extract the feature data label value difference and data conflict degree corresponding to the historical records that meet the requirements, remove the outliers in the feature data label value difference and data conflict respectively, and record the corresponding average values of the feature data label value difference and data conflict after removing the outliers as the set feature data label value difference and preset data conflict degree respectively.
[0034] Specifically, the process of obtaining relevant data based on anchor word combinations includes: Acquire preliminary relevant data based on each anchor word combination, detect the data volume of the preliminary relevant data, and if the data volume of the preliminary relevant data is in a low-order data volume range, increase processing is performed on the anchor word combination; If the amount of the prepared relevant data is within the high-order data amount range, the prepared relevant data is reduced according to the anchor word priority coefficient; If the amount of the prepared relevant data is within the required data amount range, the prepared relevant data is recorded as relevant data and the relevant data corresponding to each anchor word combination is transmitted to the client respectively.
[0035] The anchor word combination is confirmed by dividing the first-class anchor words of the target patient text into an initial number of anchor word combinations, wherein the number of anchor word combinations is less than 50% of the total number of the first-class anchor words of the target patient text, a single anchor word combination includes several first-class anchor words, no anchor word of the same class exists between the anchor word combinations, and the difference in the number of the first-class anchor words between the anchor word combinations is less than or equal to 1. The initial number is set by the user. It is understood that the larger the number of anchor word combinations, the larger the amount of preliminary related data obtained; Based on a single anchor word combination, historical patient texts within a preset time period before the current moment are detected, and historical patient texts containing all anchor words of a category within the anchor word combination are detected. The detected historical patient texts, historical patient image data corresponding to the historical patient texts, and medical records corresponding to the historical patient texts are recorded as preliminary related data corresponding to the anchor word combination; The amount of prepared relevant data, that is, the total number of historical patient texts in all prepared relevant data corresponding to each anchor word combination, The values within the low-order data volume range are all smaller than the first preset data volume, the values within the required data volume range are all greater than or equal to the first preset data volume and smaller than the second preset data volume, and the values within the high-order data volume range are all greater than or equal to the second preset data volume; When increasing the anchor word combination, the absolute value L of the difference between the data volume of the preliminary related data and the first preset data volume is detected, the anchor word combination is reconfirmed, and the number of the anchor word combination is increased according to the absolute value L. The number of the anchor word combination = the initial number + the adjustment amount, and the adjustment amount is positively correlated with the absolute value L, the adjustment amount = L×K, and the value of the adjustment amount is rounded upward. It can be understood that the greater the user's allowed adjustment of the number of anchor word combinations, the greater the demand for preliminary related data, and the larger the value of K. A value of K is provided, and the value of K is 0.8; the preliminary related data obtained by the anchor word combination after the increase processing is recorded as related data, and the related data corresponding to each anchor word combination is transmitted to the client respectively.
[0036] When performing the reduction processing of the preliminary related data according to the anchor word priority coefficient, the priority coefficient corresponding to each type of anchor word is detected. The priority coefficient of a type of anchor word is the number of historical patient texts with this type of anchor word in the preset time period before the current moment. The first three types of anchor words in ascending order of the anchor word priority coefficient are deleted and the anchor word combination and the preliminary related data are re-performed; the preliminary related data obtained by the anchor word combination after the preliminary related data reduction processing according to the anchor word priority coefficient is recorded as related data, and the related data corresponding to each anchor word combination is transmitted to the client respectively.
[0037] The values of the first preset data amount and the second preset data amount can be set by the user according to actual application requirements. The minimum amount of relevant data required by the user is recorded as the first preset data amount, and the maximum amount of relevant data required by the user is recorded as the second preset data amount. This is content that is easy for technical personnel in this field to understand and will not be elaborated here.
[0038] Specifically, under the second information acquisition condition, the target patient text is converted into a high-dimensional vector through the pre-training model, and the first several historical patient texts in descending order of vector similarity with the target patient text, the historical patient image data corresponding to the historical patient texts, and the diagnosis and treatment records corresponding to the historical patient texts are recorded as relevant data and transmitted to the client; The second information acquisition condition is to determine that the target analysis method is vector similarity analysis.
[0039] Among them, the pre-trained model is the BERT model. The training of the BERT model is easy for technicians in this field to understand and will not be described in detail here. For the vector similarity between the target patient text and any historical patient text, the BERT model performs vector averaging on the target patient text and the historical patient text respectively, and uses cosine similarity to calculate the similarity of the average values of the two text vectors. Among them, how to calculate cosine similarity is content that technicians in this field have already mastered and will not be described in detail here.
[0040] The present invention also provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor, characterized in that the processor implements the associated data analysis method of the ENT symptoms when executing the computer program.
[0041] The present invention also provides a computer-readable storage medium, which stores a computer program, and is characterized in that when the computer program is executed by a processor, it implements the associated data analysis method of the ear, nose and throat symptoms.
[0042] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
[0043] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for analyzing association data of ear, nose and throat symptoms, characterized in that: include: Determine the anchor word status of the target patient text based on the number of anchor words in a category and the proportion of anchor words in a category; The target analysis method is determined based on the anchor word status of the target patient text, which is to analyze the data conflict degree of a class of anchor words or perform vector similarity analysis; Under the first information acquisition condition, determining the data acquisition method to acquire relevant data based on text similarity or anchor word combination according to the comparison result of the data conflict degree of a type of anchor word with the preset data conflict degree; Under the second information acquisition condition, the target patient text is converted into a high-dimensional vector through the pre-training model, and relevant data is obtained based on the vector similarity; Send relevant data to the client.
2. The method for analyzing association data of ear, nose and throat symptoms according to claim 1, characterized in that: Determine the target analysis method based on the anchor word status of the target patient text; If the anchor word status is that the number of anchor words in one category is greater than the preset number of anchor words in one category and the proportion of anchor words in one category is greater than the preset proportion of anchor words in one category, the target analysis method is to analyze the data conflict degree of anchor words in one category; If the anchor word status is that the number of anchor words in one category is less than or equal to the preset number of anchor words in one category or the proportion of anchor words in one category is less than or equal to the preset proportion of anchor words in one category, the target analysis method is vector similarity analysis.
3. The method for analyzing association data of ear, nose and throat symptoms according to claim 2, characterized in that: The anchor word categories include first-class anchor words and second-class anchor words; One type of anchor words is anchor words whose combined usage richness is greater than the preset combined usage richness and whose data difference is less than or equal to the preset data difference; The second type of anchor words are anchor words whose combined usage richness is less than or equal to the preset combined usage richness or whose data difference is greater than the preset data difference.
4. The method for analyzing association data of ear, nose and throat symptoms according to claim 2, wherein: Under the first information acquisition condition, detecting the data conflict degree of a type of anchor words corresponding to the target patient text, and determining the data acquisition method according to the data conflict degree of the type of anchor words; If the data conflict degree of a type of anchor word is less than the preset data conflict degree, the data acquisition method is to obtain relevant data based on text similarity; If the data conflict degree of a type of anchor words is greater than or equal to the preset data conflict degree, the data acquisition method is to perform anchor word combination for the type of anchor words and acquire relevant data based on the anchor word combination; The text similarity is determined based on the number of similarities between a class of anchor words; The first information acquisition condition is to determine that the target analysis method is to analyze the data conflict degree of a type of anchor words.
5. The method for analyzing association data of ear, nose and throat symptoms according to claim 4, characterized in that: The method for confirming the data conflict degree of a type of anchor words corresponding to the target patient text is as follows: Detect the feature data label values corresponding to all anchor words in the target patient text; Calculate the difference between the feature data label value and the preset feature data label value difference as the data conflict degree; The difference between the feature data label values is denoted as S; Among them, Hi is the feature data label value corresponding to the i-th first-class anchor word in the target patient text, H0 is the average value of the feature data label values corresponding to all first-class anchor words, and n is the total number of first-class anchor words in the target patient text.
6. The method for analyzing association data of ear, nose and throat symptoms according to claim 5, characterized in that: The process of obtaining relevant data based on anchor word combinations includes: Acquire preliminary relevant data based on each anchor word combination, detect the data volume of the preliminary relevant data, and if the data volume of the preliminary relevant data is in a low-order data volume range, increase processing is performed on the anchor word combination; If the amount of the prepared relevant data is within the high-order data amount range, the prepared relevant data is reduced according to the anchor word priority coefficient; If the amount of the prepared relevant data is within the required data amount range, the prepared relevant data is recorded as relevant data and the relevant data corresponding to each anchor word combination is transmitted to the client respectively.
7. The method for analyzing association data of ear, nose and throat symptoms according to claim 2, characterized in that: Under the second information acquisition condition, the target patient text is converted into a high-dimensional vector through the pre-training model, and the first several historical patient texts in descending order of vector similarity with the target patient text, the historical patient image data corresponding to the historical patient texts, and the diagnosis and treatment records corresponding to the historical patient texts are recorded as relevant data and transmitted to the client; The second information acquisition condition is to determine that the target analysis method is vector similarity analysis.
8. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Patient diagnosis and treatment method based on vector database and medium
CN119446395A
Automatic coding method based on clinical data standard terms
CN118886398A
GMP standard dynamic updating and management system based on big data
CN119357397A
Information processing system, information processing method, and program
JP2020154684A
Siamese Neural Networks for Flagging Training Data in Text-Based Machine Learning
US20210232911A1