Anchor point matching-based medical data analysis and processing system

By constructing a big data keyword category system and a sparse inner product approximation function, we achieve structured expression and efficient matching of users' natural language needs, solve the problem of personalized recommendation in existing medical recommendation systems under multimodal information, and improve the intelligence and accuracy of medical services.

CN120656743AActive Publication Date: 2025-09-16BEIJING HENGSHENG YUNTAI NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510745958.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-16
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

Existing medical service recommendation systems find it difficult to implement user-friendly medical self-service systems in situations where multimodal information is mixed and patients have highly heterogeneous medical backgrounds. They are unable to accurately understand the natural language requirements input by users and lack the matching complexity of high-dimensional semantic spaces, resulting in a lack of personalization or serious deviations in recommendation results.

Method used

A big data keyword category system covering the semantics of multiple types of medical services is constructed, and a confidence assessment mechanism is combined to achieve structured expression of natural language requirements. The sparse inner product approximation function is used to map keyword vectors into multi-dimensional anchor indexes, and hospital matching and recommendation are completed through central anchor aggregation and multi-factor scoring strategies.

Benefits of technology

It has significantly improved the automation, intelligence and practicality of medical service acquisition, and has the advantages of accurate semantic recognition, efficient structural modeling, precise matching logic, and strong personalized recommendations, which improves the accuracy and efficiency of users' medical navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656743A_ABST
    Figure CN120656743A_ABST
Patent Text Reader

Abstract

The invention discloses a medical data analysis and processing system based on anchor point matching, and relates to the technical field of big data acquisition, the system comprises a medical database, a patient data entry unit, a big data keyword extraction unit, a retrieval analysis unit and a screening unit; the medical database is established through big data acquisition, and a hospital center anchor point of a hospital corresponding to each piece of data is determined; the patient data entry unit is used for providing a medical demand description text for a patient to be entered; the big data keyword extraction unit is used for extracting medical demand big data keywords from the medical demand description text; the retrieval analysis unit is used for obtaining an anchor point position of the object vector; and the screening unit is used for calculating the hospital center anchor point closest to the center anchor point, and taking the hospital corresponding to the hospital center anchor point as the optimal hospital of the patient. According to the invention, the automation, intelligence and practicability levels of medical service acquisition are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data collection, and in particular to a medical data analysis and processing system based on anchor point matching. Background Art

[0002] In the current healthcare system, intelligent and self-service patient access is becoming a key development direction for medical information technology. With the uneven distribution of medical resources across cities and regions, the diversification of medical needs, and the increasing demand for personalized services, traditional methods such as manual guidance, manual registration, and static information retrieval are no longer able to meet the practical needs of the vast majority of patients for efficient, accurate, and accessible healthcare services. Especially in situations where multimodal information is mixed and patients' medical backgrounds are highly heterogeneous, the key challenge is how to implement a user-friendly self-service healthcare system that accurately understands the natural language input of users and performs intelligent matching and recommendations based on structured medical big data.

[0003] Currently, numerous research and system attempts have been made to explore medical service recommendations, semantic parsing, and hospital navigation. Existing technical solutions can be broadly categorized as follows: The first category involves medical information retrieval systems based on keyword matching. These systems typically employ simple string matching, Boolean logic operations, or TF-IDF models to map patient-entered keywords to static tags or service items in the hospital database to achieve preliminary screening and recommendation. For example, some registration platforms offer the ability to filter by department, disease, hospital level, and other criteria. Representative platforms such as WeDoctor and Good Doctor Online also integrate conditional filtering modules, but these are largely limited to menu-based selection or keyword searches, lacking the ability to understand the implicit semantics in natural language. These systems often fail to establish accurate service matching paths when faced with ambiguous input, poorly expressed, or cross-category requirements, resulting in a lack of personalized or significantly biased recommendation results. The second category involves medical semantic recognition systems that incorporate natural language processing (NLP) models. With the widespread application of deep learning in language modeling, some systems have begun to experiment with using pre-trained language models such as BERT, ERNIE, and BioBERT to semantically represent user-entered medical text, thereby extracting structural representations of elements such as diseases, examinations, and treatments. Some research, such as the "BERT-based Medical Inquiry Dialogue System," has already been able to transform inquiry content into structured knowledge graph entities. However, most of these systems remain at the experimental research level and lack the ability to process the entire process across heterogeneous medical institutions and nationwide medical data. In particular, multi-dimensional matching strategies between entities and hospital resources still rely on manual rules, making it difficult to adapt to the complexity of matching in high-dimensional semantic spaces. The third category is hospital recommendation platforms based on recommender system technology. These systems often draw on algorithms such as collaborative filtering and matrix factorization used in e-commerce and content recommendation. They model patient visit history and hospital service records, inferring the medical paths of similar groups, and implementing recommendations based on the principle of "people who have seen this disease also visited this hospital." Representative research includes medical service recommendations integrated with graph neural networks and collaborative recommendations based on knowledge graphs. However, these systems generally rely on large-scale behavioral data and user preference histories. For first-time users or those seeking medical care, recommendation models face severe cold-start issues. Furthermore, these models suffer from poor interpretability and low structural matching accuracy. In particular, they lack the ability to understand and structure the semantic content of user input in real time, making them unsuitable for high-risk, high-precision scenarios like medical guidance. Summary of the Invention

[0004] To address these technical issues, we propose a medical data analysis and processing system based on anchor matching. This system constructs a big data keyword classification system encompassing the semantics of multiple medical services, combines it with a confidence assessment mechanism to structure the expression of natural language requirements, and uses a sparse inner product approximation function to map keyword vectors into multidimensional anchor indexes. Furthermore, it uses central anchor aggregation and a multi-factor scoring strategy to achieve hospital matching and recommendation. This system boasts accurate semantic recognition, efficient structural modeling, precise matching logic, and strong personalized recommendations, significantly enhancing the automation, intelligence, and practicality of medical service acquisition.

[0005] In order to achieve the above objects, the technical solution adopted by the present invention is:

[0006] A medical data analysis and processing system based on anchor point matching, the system includes: a medical database, a patient data entry unit, a big data keyword extraction unit, a retrieval analysis unit and a screening unit; the medical database is established through big data collection, wherein each piece of data corresponds to a hospital; each piece of data includes a plurality of medical demand big data keywords with different big data keyword categories and stored in the form of a sparse inner product approximation table; according to the sparse inner product approximation table corresponding to all medical demand big data keywords of each piece of data, the hospital center anchor point of the hospital corresponding to each piece of data is determined; the patient data entry unit is used to provide the patient with a medical demand description text; the big data keyword extraction unit is used to extract the medical demand description text from the medical demand description text Medical demand big data keywords are extracted from this book; the retrieval analysis unit is used to group all the medical demand big data keywords into a retrieval entity, and after feature extraction of the retrieval entity, the features are represented in vector form according to the vector space model to obtain an object vector set; the number of sparse inner product approximation tables is determined, and a sparse inner product approximation table is constructed according to a sparse inner product approximation function family; each object vector in the object vector set is mapped through the sparse inner product approximation table to obtain the anchor point position of the object vector; the screening unit is used to calculate the common central anchor point position of all object vectors in the retrieval entity, and then calculate the central anchor point of the hospital closest to the central anchor point position, and the hospital corresponding to the central anchor point of the hospital is used as the best hospital for the patient.

[0007] Furthermore, the big data keyword categories include at least: location, department name, hospital grade, consultation time, examination items, symptom description, treatment method, expert qualification and medical insurance category; the primary key of each data in the medical database is the hospital number of the hospital.

[0008] Furthermore, the big data keyword extraction unit extracts the medical demand big data keywords from the medical demand description text, and then determines whether the medical demand big data keywords can be classified into a specific big data keyword category by calculating the confidence of the medical demand big data keywords.

[0009] Furthermore, the number M of sparse inner product approximation tables satisfies the following constraints:

[0010]

[0011] Among them, δ is the tolerable missed match rate; P is the number of object vectors in the object vector set; P1 is the threshold value of the number of big data keywords.

[0012] Furthermore, the confidence level is:

[0013]

[0014] Among them, Γ c (w i ) is the i-th keyword w i Confidence score of belonging to big data keyword category c; is the jth standard reference keyword in big data keyword category c; To represent the keyword w i and The semantic similarity in the embedding space between them; Standard reference keywords The frequency of documents appearing in the medical database; n is the number of standard reference keywords in the big data keyword category.

[0015] Furthermore, the sparse inner product approximation function family is:

[0016]

[0017] in, For the i-th object vector f i The anchor point in the mth sparse inner product approximation table; T is the number of hash bits of the sparse inner product approximation table; f i,k is the value of the i-th object vector in the k-th dimension; is the random projection weight of the k-th dimension corresponding to the t-th hash function in the m-th sparse inner product approximation table; ρ i,k is the i-th object vector f i Normalized confidence deviation in the kth dimension; Generates hash anchor index values ​​for bitwise concatenation; d is the dimension of the object vector.

[0018] Furthermore, the center anchor position for:

[0019]

[0020] Where z={z1,z2,…,z M} is the center anchor candidate vector, z1 is the first center anchor candidate, z2 is the second center anchor candidate, z M is the Mth center anchor candidate; To represent the anchor point in the mth sparse inner product approximation table Hash density frequency; is the discrete integer vector space composed of all anchor point spaces.

[0021] Furthermore, the best hospitals are:

[0022]

[0023] Among them, H * The number for the best hospital; is the coordinate of the central anchor point in the mth sparse inner product approximation table; is the anchor point position in the mth sparse inner product approximation table of hospital h; δ(·) is the anchor point matching score function, which is the inverse of the Hamming distance; is the query frequency of hospital h in the mth sparse inner product approximation table near the central anchor point; GeoD(h,u) is the geographical distance between hospital h and patient location u; Gather for the hospital.

[0024] Furthermore, each medical demand big data keyword corresponds to a big data keyword category; the number of extracted medical demand big data keywords exceeds the set big data keyword number threshold; the number of medical demand big data keywords under each big data keyword category is at most one.

[0025] Compared with the existing technology, the beneficial effects of the present invention are: it has a highly intelligent, structured and refined ability to understand medical needs and match services, significantly improving the accuracy and efficiency of user medical navigation. By constructing a big data keyword classification system covering multiple semantic dimensions such as location, symptoms, examinations, hospital level, and expert qualifications, the system can automatically parse the natural language medical needs entered by users into structured semantic units, avoiding the ambiguity and ambiguity under traditional manual selection or keyword retrieval methods. The system introduces a confidence assessment mechanism to classify and attribute each keyword to ensure the accuracy and uniqueness of the structural expression. At the structural matching level, the system innovatively adopts sparse anchor point modeling and multi-table mapping strategies to achieve efficient discretization of semantic vectors, ensuring both semantic fidelity and good computational efficiency. By constructing a central pinpoint model, the system realizes the aggregate reasoning of multiple keyword structural information, making the matching process more stable and having overall semantic consistency. In the hospital selection stage, the system integrates multi-factor scoring mechanisms such as structural matching, historical visit popularity and geographical distance to effectively balance medical service capabilities and medical convenience, significantly improving the personalization and practicality of recommendation results. Overall, the system has systematic technical advantages in medical semantic understanding, structural modeling and intelligent recommendation. It can provide high-precision and high-efficiency self-service support for medical services in multi-source heterogeneous medical data scenarios, and has significant practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a schematic diagram of the system structure of the medical data analysis and processing system based on anchor point matching proposed by the present invention;

[0027] Figure 2 Schematic diagram of experimental results comparing the performance of the anchor point matching-based medical data analysis and processing system with existing technologies;

[0028] Figure 3 This is a schematic diagram of the experimental results showing the impact of the number of sparse inner product approximation tables on the accuracy of hospital center anchor point matching;

[0029] Figure 4 Schematic diagram of the experimental results on the impact of geographical distance on the recommendation of the best hospital. DETAILED DESCRIPTION

[0030] The following description is intended to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are merely examples, and those skilled in the art may conceive of other obvious variations.

[0031] Reference Figure 1As shown, a medical data analysis and processing system based on anchor point matching, the system includes: a medical database, a patient data entry unit, a big data keyword extraction unit, a retrieval analysis unit and a screening unit; the medical database is established through big data collection, wherein each data corresponds to a hospital; each data includes a plurality of medical demand big data keywords with different big data keyword categories, and stored in the form of a sparse inner product approximation table; according to the sparse inner product approximation table corresponding to all medical demand big data keywords of each data, the hospital center anchor point of the hospital corresponding to each data is determined; the patient data entry unit is used to provide the patient with a medical demand description text; the big data keyword extraction unit is used to extract the medical demand description text from the medical demand description Medical demand big data keywords are extracted from the text; the retrieval analysis unit is used to form all the medical demand big data keywords into a retrieval entity, and after feature extraction of the retrieval entity, the features are represented in vector form according to the vector space model to obtain an object vector set; the number of sparse inner product approximation tables is determined, and a sparse inner product approximation table is constructed according to a sparse inner product approximation function family; each object vector in the object vector set is mapped through the sparse inner product approximation table to obtain the anchor point position of the object vector; the screening unit is used to calculate the common central anchor point position of all object vectors in the retrieval entity, and then calculate the central anchor point of the hospital closest to the central anchor point position, and the hospital corresponding to the central anchor point of the hospital is used as the best hospital for the patient.

[0032] The medical database, the system's data foundation, is constructed using a real-time big data collection mechanism. Data sources include online platforms of medical institutions at all levels, government registration databases, medical insurance designated hospital directories, medical capacity registration systems, and third-party medical review platforms. The structure of the collected data is designed around multidimensional medical service characteristics. Each hospital is modeled as a data record containing multiple service capacity fields with different semantic categories. For example, information such as a hospital's specialty scope, primary disease categories, diagnostic and treatment technology types, auxiliary medical equipment availability, medical insurance reimbursement ratio, location and transportation accessibility, and medical staff qualifications are uniformly abstracted as medical big data keywords and categorized by semantic categories, forming a multi-level keyword hierarchy. Keywords are stored using a sparse structure and mapped to a sparse inner product approximation table using a family of hash functions, serving as a discrete structural index of hospital characteristics. This design ensures that the system maintains high processing efficiency and good semantic search resolution even when processing large-scale hospital data sets.

[0033] The patient data entry unit is primarily responsible for capturing user input describing medical needs. This input is typically expressed in natural language and includes a description of the patient's condition, desired treatment options, preferred cost, geographic location, and hospital level requirements. This unit provides a human-computer interface and performs language normalization on the input, including noise word removal, synonym normalization, medical terminology standardization, and syntactic analysis. This ensures that subsequent processing modules receive a clearly structured and semantically complete raw data carrier. This processed text is then passed to the keyword extraction module for semantic analysis.

[0034] The task of the keyword extraction module is to identify key medical service intent words in the medical needs description text and to clarify the category to which each keyword belongs. In this system, each keyword must be uniquely attributed to a certain category, such as disease category, treatment method category, diagnostic equipment category, payment method category, medical level category, etc. This classification is a prerequisite for system semantic modeling, which ensures the dimensional orthogonality and independence of anchor point reasoning of subsequent vector modeling. The keyword extraction module relies on the entity vocabulary constructed by the medical ontology, and combines the domain language model to perform contextual semantic analysis to extract semantically directional keyword groups. The system limits each category to retain a maximum of one keyword to avoid information redundancy or conflict within the category. The extraction results will form a keyword set as the "retrieval entity" for subsequent processing by the system.

[0035] The main task of the retrieval and analysis unit is to convert the extracted keyword set into a spatial representation suitable for structured semantic search and to pre-match hospital data through a sparse inner product approximation mechanism. First, the system will vectorize each keyword through a medical semantic embedding model to obtain its expression in a high-dimensional semantic space. This vector not only expresses the literal meaning of the keyword, but also incorporates the contextual relationship between the word and other related words in the medical semantic system. For example, words such as "diabetes" and "endocrinology," "insulin treatment," and "blood glucose meter" will show a similar geometric distribution in this space, while being clearly distinguishable from "orthopedic surgery" and "tumor targeting." The semantic vector model used by the system can be a medical-specific pre-trained model based on BERT, or a semantic vector dictionary constructed using traditional Word2Vec training medical data corpus. The keyword vector set constitutes the feature expression of the user's input requirements in the semantic space.

[0036] The vectorized keyword set also needs to be mapped to the system's structural index system. To this end, the system introduces a sparse inner product approximate table modeling mechanism, which projects each vector to a predefined anchor position in the discrete space by constructing multiple approximate hash tables. These hash tables are generated using a specific family of approximate hash functions, which have a high probability of retaining the proximity relationship between similar vectors in the original semantic space, that is, two semantically similar keywords, their mapping result anchors will fall in adjacent or the same bucket in the hash table. Each keyword will be mapped to multiple hash tables separately to obtain multiple anchor positions. This design provides a redundant fault-tolerant mechanism in the spatial dimension and improves the system's robustness to natural language variants. For example, if a user uses "abnormal liver function" to describe their needs, the system can recognize that it has an intersection with "liver disease" and "cirrhosis" in the anchor space, thereby retrieving hospitals related to liver disease treatment in the hospital data.

[0037] Obtaining the anchor point position is only the first step. The system also needs to integrate the anchor point sets from multiple keywords to solve the central anchor point position of the entire search entity in the sparse space. This process is based on the anchor point aggregation inference algorithm, that is, the system performs weighted clustering calculations on the positions of all keyword anchor points in each hash table to determine the "semantic center point" in the multidimensional anchor point space. This center point represents the position of the user's overall needs in the structural semantic space, and is the benchmark reference for subsequent similarity matching with the hospital feature anchor points. Since different keywords have different semantic weights when expressing user intent, the system will also adjust the weights in the anchor point aggregation process. For example, the anchor point contribution of disease keywords will be higher than that of geographic keywords or cost keywords to ensure that the final match is more focused on the core medical needs.

[0038] After completing the construction of the central anchor point of the retrieval entity, the system enters the screening stage. At this point, the system will traverse the central anchor point position of each hospital in the medical database and perform a sparse spatial distance measurement with the user's retrieval central anchor point. This distance is not a simple numerical difference, but an approximate similarity measure based on the anchor point structure. It may be estimated using algorithms such as Boolean coincidence of anchor point positions, hash collision rate, and mapping sequence edit distance. Ultimately, the system selects the central anchor point of the hospital closest to the retrieval central anchor point as the optimal matching result. The hospital in the matching results is the best medical institution recommended by the system to the patient, which has the best fit between medical capabilities, service characteristics and user expectations. This process does not require manual selection by the user. The system completes it automatically based on structural reasoning, truly realizing an automated service process from natural language input to high semantic matching.

[0039] Throughout the system's operation, the core technical bottlenecks and innovations focus on how to transform unstructured natural language medical needs into structured, indexable information units, and achieve efficient and accurate matching in the vast medical service space. Keyword extraction is the basis of semantic understanding, semantic vectorization is the key to semantic mapping, sparse anchor point modeling is the core of structural indexing, and central anchor point reasoning and approximate distance matching determine the accuracy and rationality of the final recommendation. By organically combining the above-mentioned technical modules, this system breaks through the traditional rough matching mechanism based on keywords or tags, and establishes a medical service matching system with semantic understanding as the core, spatial indexing as the foundation, and structural reasoning as the engine.

[0040] The system is highly versatile and adaptable. In addition to being suitable for self-service recommendations for individual users, it can also be embedded in government medical systems, medical insurance management systems, remote diagnosis and treatment platforms, etc., to achieve intelligent allocation of medical resources in a wider range of scenarios. Because the system is based on a sparse approximate table structure, its data update, expansion and maintenance costs are much lower than traditional dense index systems, and it is suitable for large-scale medical data access and real-time matching needs in dynamically changing scenarios. In addition, the system has good interpretability, and each step of the matching process can trace the keyword semantic path, anchor point space projection and final distance calculation logic, providing data support and technical guarantee for subsequent user behavior analysis, recommendation strategy optimization and medical behavior supervision.

[0041] Furthermore, the big data keyword categories include at least: location, department name, hospital grade, consultation time, examination items, symptom description, treatment method, expert qualification and medical insurance category; the primary key of each data in the medical database is the hospital number of the hospital.

[0042] Furthermore, the big data keyword extraction unit extracts the medical demand big data keywords from the medical demand description text, and then determines whether the medical demand big data keywords can be classified into a specific big data keyword category by calculating the confidence of the medical demand big data keywords.

[0043] The location category identifies the geographic area where the user desires to seek medical treatment, such as city, county, or subway station. The department name category identifies the clinical departments the user desires to visit, such as internal medicine, orthopedics, obstetrics and gynecology, and pediatrics. The hospital grade category corresponds to the user's expectations of the medical institution's grade, such as tertiary hospitals and secondary hospitals. The visit time category includes the user's desired time of day, such as weekdays, holidays, and weekend night clinics. The examination item category identifies the examination requirements the user mentions, such as CT, MRI, blood routine, and gastroscopy. The symptom description category extracts subjective descriptions related to physical conditions or pathological reactions, such as cough, abdominal pain, fever, and lumps. The treatment method category covers the patient's desired intervention method, such as surgery, medication, physical therapy, and traditional Chinese medicine. The expert qualification category identifies whether the user requires high-end medical personnel, such as chief physicians, postdoctoral fellows, and overseas experts. The medical insurance category identifies the patient's payment method, such as urban employee medical insurance, new rural cooperative medical system, self-funded, and public medical insurance. Each category in the system acts as a semantic container, responsible for receiving keywords classified into that category from natural language input.

[0044] The medical database is designed with the hospital as the smallest data unit. Each record corresponds to a specific hospital, with the hospital number as the unique primary key, ensuring the certainty of each entity in the database. The database fields are structured according to the aforementioned keyword categories, with the fields within each category corresponding to the hospital's service features. For example, the "Examination Items" category records all the examination capabilities of the hospital, while the "Expert Qualifications" category records the titles and expertise of the hospital's resident experts. The "Medical Insurance Category" field records the medical insurance settlement methods supported by the hospital. This design ensures that the system can directly compare user requirements with the hospital's service capability fields during the matching process, establishing a precise semantic correspondence. After the user enters a text description of their medical need through the interface, the text is first passed to the big data keyword extraction unit for natural language parsing. Using a language model optimized for the medical context, the system performs word segmentation, part-of-speech tagging, named entity recognition, and syntactic analysis on the input text. It extracts words or phrases that may have medical semantic relevance and initially constructs a set of raw keywords. These original keywords are usually semantically ambiguous, that is, the same word may have multiple categories of candidate affiliation at the same time. For example, "director" may refer to both expert qualifications and department titles. Therefore, after completing the initial keyword extraction, the system needs to further determine whether each keyword can be uniquely classified into a specific category.

[0045] To solve the above classification problem, the system integrates a keyword attribution judgment mechanism based on confidence evaluation in the keyword extraction module. i , calculate its belonging confidence score Γ under each category c in turnc (w i ). The confidence calculation is based on the semantic similarity between the keyword and all reference words in the standard category vocabulary, and the frequency of occurrence of reference words in the medical database is introduced as a weighting factor, thereby forming an information fusion judgment model that takes into account both semantic proximity and common usage in the medical field. Ultimately, each keyword will select a category with the highest confidence value among all categories as its candidate category, and compare the score with the preset confidence threshold. Only when the confidence exceeds the threshold value, the system considers that the keyword can be effectively classified into the category and includes it in the structured keyword set to participate in subsequent vector modeling and sparse anchor mapping. If the confidence is insufficient, the system will remove the keyword or mark it as an unclassified word to prevent it from interfering with the construction of the system structure vector space. By introducing a confidence evaluation mechanism, the system solves the problem of word meaning uncertainty that is prevalent in natural language input. Users may use ambiguous words, abbreviations, colloquial expressions or omitted structures for description, and traditional keyword matching models often cannot accurately identify their semantic attribution. This system quantitatively models the semantic similarity of words within a category context and, combined with word frequency distributions from real medical corpora, constructs an interpretable and adjustable keyword attribution determination mechanism, achieving high-fidelity conversion from natural language input to structured semantic output. This conversion is a prerequisite for the system's subsequent object vector modeling, sparse anchor mapping, central anchor inference, and hospital recommendations. Furthermore, the system supports dynamic expansion of keyword categories and standard reference lexicons. When high-frequency words that cannot be effectively categorized appear in user input, the system manually reviews and semantically expands them through a backend logging module, assigning them to existing categories or creating new ones. The standard reference lexicon and document frequency statistics are also updated. This mechanism ensures that the system's semantic recognition capabilities are continuously optimized and enhanced as user data accumulates, forming a continuously learning and dynamically enhanced semantic perception capability, thereby continuously improving the system's practicality and adaptability in real-world medical service scenarios.

[0046] refer to Figure 2 ,Furthermore, the number M of sparse inner product approximation tables satisfies the following constraints:

[0047]

[0048] Among them, δ is the tolerable missed match rate; P is the number of object vectors in the object vector set; P1 is the threshold value of the number of big data keywords.

[0049] In anchor matching-based medical data analysis and processing systems, the performance of the entire system is highly dependent on its approximate matching capabilities on large-scale medical datasets. The core technical foundation of this capability is the construction of sparse inner product approximation tables to achieve fast and high-fidelity object vector retrieval. The system's retrieval and analysis unit uses sparse inner product approximation tables as its basic structure to map the semantic vectors of patient needs into a discrete anchor point space, enabling efficient comparison of high-dimensional semantic features in sparse space. To ensure the accuracy and stability of the matching, the number of sparse inner product approximation tables, denoted as M, must meet certain constraints. This constraint relationship reflects the mathematical balance mechanism between system fault tolerance and recognition rate in multi-object, multi-category, and multi-anchor retrieval scenarios.

[0050] The design of this formula is based on a comprehensive consideration of the system's query error tolerance, keyword quantity threshold, and object vector size. First, the parameter δ represents the probability of missed matches that the system can accept, also known as the tolerable missed match rate. In the medical self-service recommendation scenario, missed matches mean that the hospital results retrieved by the user fail to cover their actual semantic intent, which will have a direct impact on the effectiveness of medical services. Therefore, the parameter value should be controlled at a very small level, usually around 10 -2 to 10 -6 The specific value can be set according to the application environment. When δ is smaller, it means that the system needs a higher anchor point space density coverage to reduce the matching omissions caused by approximate hash mapping, which requires an increase in the number of sparse tables M.

[0051] The parameter P is the number of object vectors in the object vector set. In this system, this set comes from the semantic vector mapped to each keyword in the keyword set obtained by the system through keyword extraction and semantic classification after the patient enters the medical needs text. Therefore, P is actually equivalent to the number of keywords involved in semantic modeling in user needs. Each keyword occupies a position in the anchor point space as an independent semantic dimension, and the anchor point it maps will participate in the aggregation operation of the final central anchor point. The more objects there are, the wider the anchor point space it covers. To maintain a certain matching accuracy and recall rate, the system needs to build more sparse inner product tables to meet the needs of all objects being hit simultaneously in the hash space.

[0052] Parameter P1 is the system's preset threshold for the number of big data keywords. In the medical self-service system, in order to prevent the semantic space from being overly fragmented, the system limits the extraction of at most one keyword under each category to form a high-quality, low-redundancy semantic input structure. P1 represents the lower limit of the number of keywords allowed by the system in a valid query, that is, if the number of keywords entered by the user does not exceed this threshold, the system will consider that the semantic information is insufficient and will not start the query process. This parameter and the number of object vectors P constrain each other to form the system's basic control over input scale and information density. When P is much higher than P1, it indicates that the user input information is sufficient. In order to cover higher-dimensional anchor spaces, the system will increase the number of sparse tables accordingly to prevent the anchor points from deviating from the aggregation center due to sparse high-dimensional space mapping.

[0053] The fractional term on the right side of the formula is the minimum coverage estimation expression in the sparse inner product space. Its numerator is This represents the minimum spatial coverage required by the system to ensure that the probability of missed matches does not exceed δ. The denominator, log(P) - log(P1), represents the information density ratio between the number of input objects and the system's keyword threshold. When P is much larger than P1, the denominator increases, the entire fraction decreases, and the number of sparse tables required decreases. Conversely, if P is close to P1, the system assumes that the input information is of low dimensionality or has a narrow category distribution. To maintain recall, the number of sparse tables needs to be increased to improve the spatial hit rate.

[0054] The whole formula multiplied by the number of objects P reflects the parallel mapping requirements under the multi-object anchor space aggregation modeling. Since each object vector needs to be mapped in all sparse tables, the actual number of hash tables required by the system is not static, but a function of the number of objects. If each sparse table is regarded as a slice of the semantic space under a certain anchor dimension, the minimum requirement of M represents the system's ability to fully cover the semantic space of patient retrieval entities in a limited anchor structure. When M is not set enough, some keywords entered by the patient may lack corresponding anchors in some or all tables, causing the system to lose the dimensional reference point when aggregating at the central anchor, resulting in semantic ambiguity or recommendation offset.

[0055] The underlying logic of this formula comes from the recall analysis principle of the locality sensitive hashing algorithm in the approximate search of high-dimensional vector space. In traditional LSH theory, in order to effectively identify two approximate object vectors in a high-dimensional space, it is necessary to construct enough hash tables to ensure that the two vectors are mapped to the same or similar positions in at least one table. Corresponding to this system, if each keyword vector in the patient's needs cannot be effectively mapped in the sparse anchor point space, the corresponding medical service characteristics will be weakened in the subsequent hospital comparison, which will ultimately affect the accuracy of the recommendation. Therefore, in order to ensure the overall matching performance of the system, it is necessary to dynamically calculate and set the appropriate number of sparse tables based on the keyword density, the probability of hash table collision and the matching error control range allowed by the system.

[0056] In medical application scenarios, the practical significance of this formula is particularly important. Due to the highly professional and refined nature of medical services, different patients may use different terms to describe the same disease. This linguistic difference will appear as a multidimensional distribution in the vector space. If the system uses a unified table number setting without considering the input scale, matching confidence and keyword distribution differences, it is very easy to have missed matches or wrong matches in actual operation, which will seriously affect the credibility of the system recommendations and user experience. By introducing the above-mentioned sparse table number constraint formula, the system can achieve dynamic tuning of the sparse space structure, so that the system can maintain stable semantic coverage capabilities under different user input conditions.

[0057] In addition, this formula can also guide the system's optimization of resource allocation. When facing large-scale, highly concurrent queries, the system can use this formula to calculate the minimum sparse table requirements for the current request, and load the corresponding number of sparse tables through the asynchronous table cache mechanism, avoiding the waste of computing resources caused by loading all hash tables, thereby achieving a balance between resource utilization and retrieval performance. Combined with the cache heat analysis mechanism, this formula can also be used for dynamic table number trimming. That is, for frequently accessed common keyword combinations, the corresponding table set can be set with higher redundancy to improve query speed, while for low-frequency requests, the number of tables is compressed to save system resources.

[0058] refer to Figure 3 and Figure 4 , further, the confidence is:

[0059]

[0060] Among them, Γ c (w i ) is the i-th keyword w i Confidence score of belonging to big data keyword category c; is the jth standard reference keyword in big data keyword category c; To represent the keyword wi and The semantic similarity in the embedding space between them; Standard reference keywords The frequency of documents appearing in the medical database; n is the number of standard reference keywords in the big data keyword category.

[0061] In this formula, Γ c (w i ) represents the i-th keyword w i The confidence score of the keyword being judged by the system to belong to category c. This score is used to evaluate the semantic consistency between the keyword and the target category. The higher the value, the closer the semantic features of the keyword are to the semantic center of category c, and the higher the confidence of the system in its classification. This indicator is not only used for subsequent keyword screening and vector modeling, but also used to limit irrelevant keywords from entering the semantic index structure to avoid semantic pollution. In the formula, The keyword w i and the jth standard reference keyword in category c The semantic similarity value of . This similarity calculation is performed in the semantic embedding space. Usually, the vocabulary is pre-projected into a high-dimensional semantic space through a medical language model such as BioBERT, Med-BERT or the medical semantic Word2Vec model, and then the value is obtained through cosine similarity calculation. Since each category c usually contains multiple standard reference words, these words together constitute the semantic boundary of the category, the system performs similarity calculation on the keyword w. i The similarity with these words is summarized and statistically analyzed.

[0062] In the confidence formula, the numerator is a weighted sum term, the core of which is to sum each semantic similarity term Multiply by the logarithm of the frequency of the reference keyword in the medical database, that is, This weighting mechanism reflects the system's trust preference for high-frequency medical keywords. In actual medical databases, certain standard keywords such as "diabetes", "coronary heart disease", "CT examination", "hospitalization reimbursement", etc. appear much more frequently than rare terms. By introducing a logarithmic weighting mechanism for document frequency, the system can enhance the confidence scores of words that are close to common semantic centers, and suppress the risk of misclassification of words that have accidental semantic similarity but ambiguous semantic attribution. The denominator is the square root of the similarity value, which plays a normalization role, so that the overall score is not affected by the number of reference words n. The square root method is used to construct a normalization standard consistent with the vector norm, so that the attribution scores are comparable when calculated between multiple categories. Addend 10 -6 It is a numerical stability constant to avoid zero values ​​in the denominator and does not participate in the substantive adjustment of the confidence level.

[0063] In the specific operation of the anchor matching-based medical data analysis and processing system, the confidence calculation process mainly occurs after the user completes the text input and before the keyword extraction module outputs. The system will calculate the confidence score Γ for each category of all possible medical-related words in the user's text. c (w i ), and retain the words with the highest scores in each category as the representative keywords of that category. For example, if a patient's input text contains keywords such as "cirrhosis", "interferon treatment", "medical insurance reimbursement", and "Shenzhen Nanshan", the system may classify "cirrhosis" into the disease category because it has high semantic similarity and high frequency of occurrence with standard reference words under disease categories such as "chronic liver disease", "abnormal liver function", and "viral hepatitis"; "interferon treatment" may be classified into the treatment method category; "medical insurance reimbursement" is classified into the cost model category; and "Shenzhen Nanshan" is classified into the geographical location category. The keywords of each category that are finally retained are the optimal results under the confidence calculation.

[0064] The introduction of this mechanism solves the classification conflict problem that exists in traditional keyword extraction when faced with semantic redundancy and attribution ambiguity. Since users often use free expression sentences when describing medical needs, and synonyms, generalized words, and specific words appear alternately, it is difficult to accurately determine their category attribution by simply relying on keyword surface matching. By introducing the fusion modeling of semantic similarity and document frequency, the system can improve the stability and accuracy of structured input while maintaining semantic understanding capabilities. In addition, the confidence score is not only used for single classification, but also as a regulating factor for subsequent vector weight modeling. During the system vectorization process, each retained keyword will be converted into a semantic vector and participate in the projection of the sparse anchor space. And the confidence score Γ c (w i ) can be used as a vector weighting coefficient to determine the semantic weight of the keyword in the vector combination. A higher score indicates a more important keyword, and its vector has a greater influence in subsequent anchor projection and center anchor aggregation. Conversely, even if a keyword with a lower score is included in the model, its weight is appropriately weakened to reduce its impact on the final result.

[0065] Furthermore, the sparse inner product approximation function family is:

[0066]

[0067] in, For the i-th object vector f i The anchor point in the mth sparse inner product approximation table; T is the number of hash bits of the sparse inner product approximation table; f i,k is the value of the i-th object vector in the k-th dimension; is the random projection weight of the k-th dimension corresponding to the t-th hash function in the m-th sparse inner product approximation table; ρ i,k is the i-th object vector f i Normalized confidence deviation in the kth dimension; Generates hash anchor index values ​​for bitwise concatenation; d is the dimension of the object vector.

[0068] This formula describes how to transform the i-th object vector f i , in the mth sparse inner product approximation table, after T hash mappings, a discrete hash value is generated as the sparse anchor point index of the vector This sparse anchor value is composed of multiple binary hash functions, which represents the position of the object in the hash table and is used for subsequent similarity retrieval and center anchor aggregation operations. In the formula, each object vector f is first defined. i The dimension of is d, which is the representation dimension in the high-dimensional semantic space. This vector comes from the structured semantic expression obtained by the system through keyword extraction, semantic classification and vector embedding after the user inputs the medical needs. Each dimension f in the object vector i,k Corresponding to the projection value of a keyword semantic feature in the k-th dimensional space.

[0069] Each sparse table consists of T independent hash functions, each of which compresses a high-dimensional vector into a single binary bit. These hash functions are parameterized by a set of pre-generated randomly generated projection vectors. Control. For the tth hash function of the mth sparse table, its role is to perform a weighted sum projection on the components of the input vector in all dimensions to form a scalar judgment result. If the result is positive, the output is 1, otherwise it is -1, which is mapped to a binary representation by the sign function sign(·) and used as the hash bit at that position. Unlike traditional hash projection, this system introduces a normalized confidence deviation factor ρ in the hash function i,k , which is used to adjust the contribution weight of different dimensional components in the projection. This deviation value represents the confidence uncertainty of the i-th object in the k-th dimension during the semantic vector construction process, which may come from factors such as semantic ambiguity, context ambiguity, or fuzzy category boundaries when modeling word vectors. The system introduces This mapping function keeps the confidence deviation within a stable range while also implementing a smooth penalty for uncertain components. This function ensures that when the confidence deviation is small, meaning the dimension component is more reliable, it contributes more to the overall projection. Conversely, when the dimension has a large deviation, its impact on the hash projection is suppressed, reducing the impact of semantic perturbations on the consistency of anchor point generation.

[0070] The hash function sums the products of each dimension, converts them into binary symbols through the sign function, and then multiplies them by 2 t-1, forming an integer hash value. All T-bit hash values ​​are concatenated by bitwise concatenation. Combined into the final anchor index value This construction method allows the system to encode the results into low-dimensional sparse representation while maintaining high-dimensional semantic distribution information, thereby facilitating fast query, aggregation and matching in multiple sparse hash tables. During the operation of the system, each keyword vector will be mapped to multiple sparse hash tables, each with a different random projection vector. To achieve multi-angle segmentation of the semantic space. In this way, the same object vector may obtain different anchor indexes in different tables, thereby improving the overall coverage and query robustness. When the mapping results of multiple object vectors are clustered in similar positions, the system can aggregate their anchor positions to form a semantic center anchor for subsequent hospital matching. This design ensures that even if the user inputs keywords in different ways of expressing them, as long as their semantics are similar, their vector mapping results will be close in the anchor space, thereby achieving semantic fault-tolerant matching. In addition, since each hash function only performs weighted summation and symbol judgment, its computational complexity is low and easy to deploy in parallel, making it suitable for fast indexing in massive high-dimensional semantic vectors in medical databases. The introduction of the normalized confidence function further enhances the system's adaptability to input uncertainty and semantic ambiguity, enabling the system to maintain stable matching performance in open semantic scenarios.

[0071] Furthermore, the center anchor position for:

[0072]

[0073] Where z={z1,z2,…,z M} is the center anchor candidate vector, z1 is the first center anchor candidate, z2 is the second center anchor candidate, z M is the Mth center anchor candidate; To represent the anchor point in the mth sparse inner product approximation table Hash density frequency; is the discrete integer vector space composed of all anchor point spaces.

[0074] This formula is essentially a least squares optimization model with a weighted regularization term, which is used to find a central anchor vector z={z1,z2,…,z M}, so that it is mapped to the anchor point set obtained by all object vectors The overall distance is the smallest. The center needle point vector C q Each dimension z mIt represents the candidate position of the central anchor point in the mth sparse approximation table. The entire vector constitutes a discrete point in the sparse space, which is used to represent the structured semantic center of gravity of the entire medical demand entity.

[0075] The outer minimization operation of this formula It means selecting a set of optimal center pinpoint positions from all possible anchor point position combinations. Since the pinpoint space in the sparse table is a discrete integer set, the optimization process is performed in the discrete space Z M Each element of this discrete space is an M-dimensional vector, representing the candidate anchor point position on all M sparse tables.

[0076] The inner layer of the formula is a double summation structure. The first summation index i traverses each object vector in the object vector set, that is, the vector representation mapped to each keyword extracted from the patient's needs; the second summation index m traverses the number of the sparse inner product approximation table. Each object vector has an anchor value in each table. The system tries to find a central pin point z m , so that it is anchored to these objects in all tables The gap is the smallest.

[0077] The calculation of each gap term uses the weighted square distance form, where the distance calculation term is That is, the absolute difference between the candidate center anchor point and the object anchor point, which is used to characterize their relative position difference in the hash space in the sparse table. This difference is weighted by a normalization factor with a logarithmic penalty term, which is in Indicates the anchor point in the mth sparse table The hash density frequency of , that is, the frequency of the anchor point appearing in all hospital data.

[0078] The core purpose of introducing the logarithmic weighting of hash density frequency is to adjust the distance penalty for "hot" anchor positions. If an anchor appears frequently in the table, it indicates that the information it represents has generalization or broad matching capabilities. The system will reduce the penalty for the distance error associated with it to prevent such anchors from overly dominating the matching results. Conversely, if an anchor is relatively rare, the information it contains is often more discriminative, and the system will increase the penalty for deviation from it accordingly to preserve the semantic expressiveness of key semantic anchors.

[0079] The entire objective function accumulates the penalty for each anchor point deviation through a square weighted structure to form an optimized target surface. The system finally selects z={z1,...,z M}As the central needle point position of the entire medical demand Cq , which represents the semantic center of gravity of the requirement in the sparse space. This center of gravity serves as the reference vector for the system to perform spatial matching with hospital anchors in the medical database in subsequent stages. It reflects the joint structural expression of all user-entered keywords in the multi-anchor space.

[0080] At the system application level, the significance of central pinpoint calculation lies not only in aggregating input information, but also in compressing high-dimensional semantic vectors into low-dimensional index structures, and enhancing the robustness of the system in the face of partial input loss, ambiguous expressions, or semantic conflicts. This aggregation method does not rely on vector linear superposition or mean operations, but instead uses the statistical characteristics of pinpoint distribution in sparse space to construct a minimization objective function, thereby possessing stronger discrete matching capabilities and semantic reconciliation capabilities. In particular, when faced with synonymous expressions, pragmatic ambiguity, or word frequency bias in medical semantics, the system can automatically adjust the anchor point aggregation logic to avoid a certain keyword dominating the overall semantic judgment and improve matching fairness.

[0081] Because the center point vector is composed of a set of discrete index values, the result can be directly compared with the center point of each hospital in the hospital database. The system can use simple sparse space distance measures such as Hamming distance, hash collision rate, or discrete point sequence difference to determine the closest hospital anchor point, thereby making the final recommendation. This mechanism frees the system from the computational complexity and instability associated with floating-point vector comparisons in high-dimensional semantic space, significantly improving matching efficiency while maintaining high accuracy.

[0082] Furthermore, the best hospitals are:

[0083]

[0084] Among them, H * The number for the best hospital; is the coordinate of the central anchor point in the mth sparse inner product approximation table; is the anchor point position in the mth sparse inner product approximation table of hospital h; δ(·) is the anchor point matching score function, which is the inverse of the Hamming distance; is the query frequency of hospital h in the mth sparse inner product approximation table near the central anchor point; GeoD(h,u) is the geographical distance between hospital h and patient location u; Gather for the hospital.

[0085] In this formula, the hospital set H contains all hospital records with structured anchor information in the medical database, and each record has an anchor position vector in multiple sparse inner product approximation tables. These anchor positions represent the structural position of the hospital in the semantic space, which is the multi-dimensional coding result of the hospital's service capabilities. The core of the formula is the anchor representation of each candidate hospital h, and the central anchor vector generated by user demand Perform structural matching in sparse space to measure the consistency between the hospital and user needs in multiple semantic dimensions. The basic calculation unit of matching degree is That is, the similarity evaluation function between the position of the central needle point in the mth sparse table and the corresponding anchor point position of the hospital. This function uses the inverse form of the Hamming distance, that is:

[0086]

[0087] The Hamming distance Ham(·) here represents the number of bit differences between two sparse space anchor points in the discrete index space, reflecting the degree of deviation between the two at the structural encoding level. When the two anchor points are exactly the same, the distance is 0 and the matching degree is maximized; if there are multiple bit differences, the matching degree decreases accordingly. The purpose of using its reciprocal form is to construct a mechanism where closer anchor point combinations receive higher scores, ensuring that structural consistency occupies a core position in the recommendation process. Each matching item is also multiplied by a heat factor in Represents the number of times hospital h was found near the center pinpoint in the mth anchor table during historical searches. This value is automatically recorded through log accumulation during system runtime and indicates whether the hospital has historically appeared frequently in matching requests for similar semantic anchors. It serves as an implicit indicator of the hospital's adaptability to this type of demand. The introduction of this term empowers the system with user behavior awareness, enabling the recommendation mechanism to not only rely on static representations of structural matching but also incorporate empirical feedback from dynamic interactions, improving the system's adaptability to real-world user demand scenarios.

[0088] The matching score of each semantic dimension in the formula is also divided by a geographic cost term 1+GeoD(h,u), where GeoD(h,u) represents the geographic distance between hospital h and the current patient location u. This distance calculation is based on the spatial Euclidean distance or route cost estimation between the patient's location information (such as IP resolution, GPS positioning, or manual address input by the user) and the hospital's registered geographic coordinates. By introducing the geographic distance term and using it as a penalty factor in the score, the system can automatically balance the semantic matching of medical services and the convenience of patient treatment. Among hospitals with similar service capabilities, the system will tend to recommend hospitals closer to patients to improve actual service accessibility and satisfaction.

[0089] The entire scoring function is calculated for all M anchor tables separately, and finally the sum of the scores of all dimensions is used as the comprehensive score of the hospital. Finally, the hospital number with the highest score is selected as the recommendation result, that is:

[0090] H * = Hospitals with the best semantic match, high search popularity, and optimal geographical distance;

[0091] This recommendation strategy implements a multi-factor fusion decision-making mechanism with semantic structure priority, behavioral heat weighting, and distance penalty regulation. The system ensures the structural consistency of medical needs through sparse structure hash matching, learns the behavioral preferences of user groups through a frequency adjustment mechanism, and ensures the actual availability of results through geographical factor constraints. Especially when faced with multiple candidate hospitals with similar structural expressions but significant geographical differences, the system can dynamically weigh service capabilities and distance convenience to make recommendations that better meet individual needs. In addition, the function is also highly interpretable. When displaying the recommendation results, the system can simultaneously output dimensional indicators such as matching score, historical heat, and geographical distance to assist users in understanding the reasons for the recommendation, and also provide behavioral feedback data support for the operators of the medical platform, which is conducive to further optimizing service layout and supply and demand scheduling.

[0092] Furthermore, each medical demand big data keyword corresponds to a big data keyword category; the number of extracted medical demand big data keywords exceeds the set big data keyword number threshold; the number of medical demand big data keywords under each big data keyword category is at most one.

[0093] The system stipulates that each medical demand big data keyword must uniquely belong to a big data keyword category. This design originates from the "monosemantizing category constraint" principle in the semantic modeling process, which aims to prevent the ambiguous attribution of keywords from causing confusion in structural expression. In actual medical texts, users may enter words such as "director," "tertiary hospital," "weekend," "medical insurance," "liver function," and "CT." These words have different semantic attributions in different contexts. For example, "director" may be a doctor's title or an administrative position; "weekend" can refer to time or imply a scheduling strategy. To avoid confusion caused by such polysemous words, the system uses a semantic confidence assessment mechanism based on the keyword category context to automatically determine the category to which each keyword is most likely to belong in the current semantic environment, and excludes other possible attributions, so that each keyword is uniquely marked as a sub-item under a category in the extraction results. This mechanism not only improves the category orthogonality of the vectorized representation, but also enhances the dimensional independence during anchor point aggregation, which is beneficial for subsequent sparse space modeling.

[0094] Secondly, the system stipulates that the number of extracted medical demand big data keywords must exceed a set big data keyword threshold. This threshold mechanism primarily ensures sufficient information density in the input semantic content, ensuring sufficient dimensionality when constructing the structured semantic space representation. Extracting too few keywords will result in insufficient vector dimensionality, excessive anchor point sparsity, and distorted semantic center aggregation, thus affecting the stability and interpretability of subsequent hospital matching. To this end, the system sets a lower limit on the number of keywords. Only when the number of extracted valid keywords exceeds the lower limit will the system initiate the anchor point mapping and center point calculation process. Otherwise, the user will be prompted that the input information is incomplete or unclear, and prompted to provide a supplementary description. This mechanism also, to a certain extent, avoids false triggering of invalid matches or erroneous recommendations due to low semantic resolution. Furthermore, the system stipulates that the number of medical demand big data keywords in each big data keyword category must be at most one. This constraint controls the sparsity of the system's structured input dimensions and is the basis for constructing a stable and non-redundant anchor space representation. In this system, the semantic vector space is not a dense continuous space, but a sparse structured space. Each category dimension corresponds to an independently mappable semantic channel in the sparse table. If multiple keywords exist within a category, the spatial representation of that category's dimensions will overlap, interfere, or even cancel each other out, disrupting structural comparability within the anchor space. To address this, after keyword extraction, the system ranks candidate keywords within each category by confidence, retaining only the highest-scoring keywords as representative keywords for that category in subsequent mapping. This mechanism not only enhances the representation certainty of each category anchor, but also reduces subsequent computational complexity, improving the interpretability and efficiency of the overall matching.

[0095] The above three constraint mechanisms jointly construct the semantic control logic of the system input. From "each keyword is uniquely classified" to "extraction quantity exceeds the threshold" to "only one keyword is retained in each category", it constitutes a three-in-one input preprocessing framework from semantic ambiguity elimination, information density guarantee, and dimensional independence maintenance. This framework not only ensures that the system can still stably extract semantic units with structural significance when faced with complex, changeable, and non-standard natural language input, but also provides a highly standardized, sparsity-controllable, and dimensionally separable structural input for a series of core modules such as subsequent sparse inner product mapping, anchor point aggregation, distance evaluation, and hospital recommendation, laying a solid foundation for vectorized modeling and index structure generation of medical needs.

[0096] The following is a specific implementation example of the anchor-matching-based medical data analysis and processing system of the present invention. This example covers the entire process from user input of medical requirement text, keyword extraction and classification, semantic vector construction, anchor point generation, central pinpoint calculation, and matching hospital scoring and recommendation. It also includes clear numerical values ​​and intermediate calculation results to facilitate understanding of the system's operational mechanism in practical engineering applications.

[0097] A patient in a certain district of a certain city enters the following medical need description: "I've been experiencing frequent abdominal bloating lately. I'd like to find a tertiary hospital for a gastroscopy. It would be best if I could schedule an appointment on a weekend, have medical insurance cover it, and have a highly qualified doctor."

[0098] The system extracts the following keywords from this natural language segment based on semantic models and entity recognition tools:

[0099] Symptom description category: abdominal distension;

[0100] Examination items: Gastroscopy;

[0101] Hospital grade: tertiary;

[0102] Clinic time: weekends;

[0103] Medical insurance category: medical insurance;

[0104] Expert Qualifications: Highly qualified doctor (mapped to chief physician)

[0105] The number of keywords is 6, which exceeds the keyword number threshold P1=4 set by the system and meets the conditions.

[0106] Assume that each keyword is converted into a 5-dimensional semantic vector (the actual dimension is more than 128 for the sake of simplicity) through a medical word vector model (such as BioWord2Vec). Construct the following object vector set:

[0107] f1=abdominal distension=[0.4,0.1,0.0,0.2,0.3];

[0108] f2=gastroscopy=[0.3,0.5,0.1,0.0,0.2];

[0109] f3=Triple A=[0.6,0.0,0.2,0.1,0.1];

[0110] f4 = weekend = [0.2, 0.3, 0.4, 0.1, 0.0];

[0111] f5=Medical Insurance=[0.1,0.2,0.0,0.5,0.3];

[0112] f6=chief physician=[0.5,0.2,0.1,0.0,0.4].

[0113] Assume the system uses M = 4 sparse inner product approximation tables, each with T = 3 hash bits, i.e., each hash table has 3 independent hash functions, generating anchor points for each object vector. Taking the first object f1 as an example, the following simulated random weights are used: In the first hash table, the random weight corresponding to the first hash function is:

[0114] Assume that the confidence deviation is: ρ 1,k =0.1 (all dimensions).

[0115] Calculate the normalization correction term: the factor for each dimension is

[0116] Projection results:

[0117]

[0118] s1=0.12-0.02+0+0.08+0.06=0.24·0.645≈0.1548;

[0119]

[0120] Repeat 3 times to get the anchor point of a table as three binary 101 (ie decimal 5). In this way, the anchor point mapping is completed for all M = 4 tables and all P = 6 objects to get the anchor point set

[0121] For each hash table m, the system performs weighted distance optimization on the anchor points of all objects in the table to find the center. Taking the first hash table as an example, the anchor points of the six objects are: 4, 5, 5, 6, 5, 4.

[0122] Set the anchor point hash density frequency:

[0123] HDF1(4)=500;

[0124] HDF1(5)=1200;

[0125] HDF1(6)=800;

[0126] Find the center position z1∈{4,5,6} and calculate the cost:

[0127] Take z1=5 as an example:

[0128]

[0129] Use an approximation:

[0130]

[0131] The cost is approximately:

[0132]

[0133] Repeat the process and select z1 that minimizes the cost as the center needle point coordinate of the first dimension The same is true for the other dimensions, and finally C is obtained q =[5,…].

[0134] Suppose there is hospital A in the hospital database, and its four-dimensional anchor point coordinates are a A =[5,4,6,5], and the other hospital B is [6,5,5,6]. The center anchor point is C q =[5,4,5,5], the patient is located in a certain district, and the geographical distances between the two hospitals are:

[0135] GeoD(A,u)=2km;

[0136] GeoD(B,u)=8km;

[0137] The query frequencies are:

[0138] Calculate the score of Hospital A:

[0139]

[0140] For example, item 1: log(1+300)≈5.71; 1+GeoD(A,u)=3; score approx.

[0141] The total score of the four factors is approximately 6.85. Hospital B, due to its distance, has a significant loss in popularity, but its score is approximately 5.20. The final recommendation is Hospital A, which is the user's top recommendation.

[0142] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A medical data analysis and processing system based on anchor point matching, characterized in that: The system includes: a medical database, a patient data entry unit, a big data keyword extraction unit, a retrieval analysis unit and a screening unit; the medical database is established by big data collection, wherein each piece of data corresponds to a hospital; each piece of data includes a plurality of medical demand big data keywords with different big data keyword categories and stored in the form of a sparse inner product approximation table; according to the sparse inner product approximation table corresponding to all medical demand big data keywords of each piece of data, the hospital center anchor point of the hospital corresponding to each piece of data is determined; the patient data entry unit is used to provide the patient with a medical demand description text; the big data keyword extraction unit is used to extract medical demand big data keywords from the medical demand description text According to keywords; the retrieval analysis unit is used to form all the medical demand big data keywords into a retrieval entity, and after feature extraction of the retrieval entity, the features are expressed in vector form according to the vector space model to obtain an object vector set; the number of sparse inner product approximation tables is determined, and a sparse inner product approximation table is constructed according to a sparse inner product approximation function family; each object vector in the object vector set is mapped through the sparse inner product approximation table to obtain the anchor point position of the object vector; the screening unit is used to calculate the common central anchor point position of all object vectors in the retrieval entity, and then calculate the central anchor point of the hospital closest to the central anchor point position, and the hospital corresponding to the central anchor point of the hospital is used as the best hospital for the patient.

2. The medical data analysis and processing system based on anchor point matching according to claim 1, characterized in that: The big data keyword extraction unit extracts the medical demand big data keywords from the medical demand description text, and then determines whether the medical demand big data keywords can be classified into a specific big data keyword category by calculating the confidence of the medical demand big data keywords.

3. The medical data analysis and processing system based on anchor point matching according to claim 2, characterized in that: The number of sparse inner product approximation tables M satisfies the following constraints: Among them, δ is the tolerable missed match rate; P is the number of object vectors in the object vector set; P1 is the threshold value of the number of big data keywords.

4. The medical data analysis and processing system based on anchor point matching according to claim 3, characterized in that: The confidence level is: Among them, Γ c (w i ) is the i-th keyword w i Confidence score of belonging to big data keyword category c; is the jth standard reference keyword in big data keyword category c; To represent the keyword w i and The semantic similarity in the embedding space between them; Standard reference keywords The frequency of documents appearing in the medical database; n is the number of standard reference keywords in the big data keyword category.

5. The medical data analysis and processing system based on anchor point matching according to claim 4, characterized in that: The sparse inner product approximation function family is: in, For the i-th object vector f i The anchor point in the mth sparse inner product approximation table; T is the number of hash bits of the sparse inner product approximation table; f i,k is the value of the i-th object vector in the k-th dimension; is the random projection weight of the k-th dimension corresponding to the t-th hash function in the m-th sparse inner product approximation table; ρ i,k is the i-th object vector f i Normalized confidence deviation in the kth dimension; ⊕ is the hash anchor index value generated by bitwise concatenation; d is the dimension of the object vector.

6. The medical data analysis and processing system based on anchor point matching according to claim 5, characterized in that: Center anchor point position for: Where z={z1,z2,…,z M } is the center anchor candidate vector, z1 is the first center anchor candidate, z2 is the second center anchor candidate, z M is the Mth center anchor candidate; To represent the anchor point in the mth sparse inner product approximation table Hash density frequency; is the discrete integer vector space composed of all anchor point spaces.

7. The medical data analysis and processing system based on anchor point matching according to claim 6, characterized in that: The best hospitals are: Among them, H * The number for the best hospital; is the coordinate of the central anchor point in the mth sparse inner product approximation table; is the anchor point position in the mth sparse inner product approximation table of hospital h; δ(·) is the anchor point matching score function, which is the inverse of the Hamming distance; is the query frequency of hospital h in the mth sparse inner product approximation table near the central anchor point; GeoD(h,u) is the geographical distance between hospital h and patient location u; Gather for the hospital.

8. The medical data analysis and processing system based on anchor point matching according to claim 7, characterized in that: Each medical demand big data keyword corresponds to a big data keyword category; the number of extracted medical demand big data keywords exceeds the set big data keyword number threshold; the number of medical demand big data keywords under each big data keyword category is at most one.

9. The medical data analysis and processing system based on anchor point matching according to claim 8, characterized in that: The big data keyword categories include at least: location, department name, hospital grade, consultation time, examination items, symptom description, treatment method, expert qualification and medical insurance category; the primary key of each data in the medical database is the hospital number of the hospital.

Citation Information

Patent Citations

  • Method for constructing intelligent matching of chest pain center

    CN118297360A

  • Employee outpatient service ataxia supervision platform and method

    CN118553411A

  • AI-based intrusion protection response data processing method and server

    CN119249411A

  • Ontology mapper

    US10268687B1

  • Ontology mapper

    US8856156B1