Medical Aesthetic Entity Alignment Method, Device, Equipment and Readable Storage Medium

By collecting and analyzing medical beauty project data, extracting and screening entity attributes, and building mapping keys based on similarity, the problem of lack of standard entity definitions in the medical beauty industry is solved, and efficient entity alignment and specification updates are achieved.

CN113887231BActive Publication Date: 2025-06-27CHENGDU MEIERBEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111223916.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-18
Publication Date
2025-06-27
Estimated Expiration
2041-10-18

AI Technical Summary

Technical Problem

The medical beauty industry lacks unified industry norms and standard entity definitions, resulting in irregular project naming and difficult entity alignment.

Method used

By collecting medical beauty project data, extracting entities and their attributes, filtering standard entity sets and non-standard entity sets, building mapping keys based on the similarity of entity attributes to achieve entity alignment.

Benefits of technology

An automated entity alignment method is provided, which can stabilize and timely update the specifications of project entities, reduce manual participation, improve efficiency and iterability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113887231B_ABST
    Figure CN113887231B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of entity alignment, and discloses a medical beauty entity alignment method, device, equipment and readable storage medium. The method includes: collecting medical beauty project data; extracting entities based on the medical beauty project data, and the entity attributes of the entities include at least one of entity semantic vectors, project entity vectors and project structure attributes; screening the entities to obtain a first standard entity set and a non-standard entity set; constructing mapping keys between the non-standard entities in the non-standard entity set and the first standard entities in the first standard entity set based on the similarity of entity attributes. The present invention solves the problem that in the existing medical beauty industry, due to the lack of relatively standardized industry standards and industry general guidelines, the naming of its projects is seriously non-standardized, and the entity alignment of project names is difficult.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of entity alignment, and specifically refers to a medical beauty entity alignment method, device, equipment and readable storage medium. Background Art

[0002] The entity alignment task is an NLP (Natural Language Processing) task after entity recognition. Its main content is to determine whether two or more entities from different information sources refer to the same object in the real world. If multiple entities represent the same object, an alignment relationship is constructed between these entities, and at the same time, the information contained in the entities is fused and aggregated. The existing entity alignment methods are mainly of two types. One is rule-based, such as using dictionaries and edit distances for entity alignment or using attribute similarity matching methods in knowledge graphs for entity alignment; the second is based on deep learning models, mapping entities into a low-dimensional vector space, then calculating the similarity between entities, and in knowledge graphs, methods such as the TransE algorithm are also used for vector representation and then calculating similarity.

[0003] The medical beauty industry has obvious differences from other industries, especially in terms of entity alignment. Entities in other industries, whether traditional industries such as real estate and finance or some emerging industries such as e-commerce and online education, actually have relatively standardized industry standards and general industry guidelines when describing and defining entities. The medical industry, which is closest to the medical beauty industry, has a strict industry standard and general entity specification definition, while the medical beauty industry is developing rapidly and there is no general industry standard entity definition for various entities in medical beauty.

[0004] The medical beauty industry is far from the entity naming norms of other industries, especially traditional industries. Therefore, in the medical beauty industry and industries such as the medical beauty industry that are developing rapidly and have no relevant industry standards, the entity alignment of project names is more difficult. Summary of the Invention

[0005] Based on the above technical problems, the present invention provides a medical beauty entity alignment method, device, equipment and readable storage medium, which solves the problem that due to the lack of relatively standardized industry standards and general industry guidelines in the existing medical beauty industry, the project naming is seriously non-standardized and the entity alignment of project names is difficult.

[0006] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0007] A medical beauty entity alignment method, comprising:

[0008] Collect medical beauty project data;

[0009] Extract entities based on medical beauty project data, and the entity attributes of the entities include at least one of entity semantic vectors, project entity vectors, and project structure attributes;

[0010] Filter the entities to obtain a first standard entity set and a non-standard entity set;

[0011] Construct mapping keys between non-standard entities in the non-standard entity set and first standard entities in the first standard entity set based on the similarity of entity attributes.

[0012] Furthermore, filtering the entities to obtain the first standard entity set includes:

[0013] Preliminarily filter the entities, and add the entities with the data source being medical beauty institutions to the first candidate set;

[0014] Count the frequencies of the entity project names of the first candidate entities in the first candidate set. If the frequency count result is greater than the first preset threshold, add the first candidate entities to the second standard entity set;

[0015] Remove the second standard entity set from the first candidate set to obtain the second candidate set;

[0016] Calculate the weights of the second candidate entities in the second candidate set. If the weight calculation result is greater than the second preset threshold, add the second candidate entities to the third standard entity set;

[0017] Combine the second standard entity set and the third standard entity set to obtain the first standard entity set.

[0018] Furthermore, counting the frequencies of the entity project names of the first candidate entities in the first candidate set includes:

[0019] Determine the major project category to which the entity project name of the first candidate entity belongs;

[0020] Obtain the number of medical beauty institutions with the major project category;

[0021] Obtain the number of medical beauty institutions with the entity project name of the first candidate entity;

[0022] Obtain the ratio of the number of medical beauty institutions with the entity project name of the first candidate entity to the number of medical beauty institutions with the major project category.

[0023] Furthermore, calculating the weights of the second candidate entities in the second candidate set includes:

[0024] Construct a mutual exclusion graph between the second candidate entities in the second candidate set based on the entity recognition model, and perform weight ranking on the mutual exclusion graph to obtain the first weight of the second candidate entity;

[0025] Based on the similarity between the entity semantic vector and the project entity vector in the entity attributes, obtain the second standard entity with the highest similarity in the second standard entity set for the second candidate entity. The similarity score between the second candidate entity and the second standard entity with the highest similarity to the entity semantic vector and the project entity vector is the second weight;

[0026] Based on the similarity of the project structure attributes in the entity attributes, obtain the second standard entity with the highest similarity in the second standard entity set for the second candidate entity. The similarity score between the second candidate entity and the second standard entity with the highest similarity to the project structure attributes is the third weight;

[0027] Subtract the second weight and the third weight of the two candidate entities from their first weight to obtain the weight difference.

[0028] Further, constructing a mapping between the non-standard entities in the non-standard entity set and the first standard entities in the first standard entity set based on the similarity of entity attributes mainly includes:

[0029] Calculate the entity attribute similarity between the non-standard entity and the first standard entity. The entity attribute similarity includes at least one of the entity semantic vector similarity, the project entity vector similarity, and the project entity attribute similarity;

[0030] Based on the entity attribute similarity, select the first standard entity with the highest similarity to the non-standard entity to establish a mapping key.

[0031] Further, the calculation method of the entity attribute similarity includes:

[0032] For the entity semantic vector or the project entity vector in the entity attributes, use the cosine similarity calculation method to calculate the similarity of the entity semantic vector or the project entity vector;

[0033] For the project structure attributes in the entity attributes, use the coincidence degree calculation method to calculate the similarity of the project structure attributes.

[0034] Further, extracting entities based on medical beauty project data includes:

[0035] Attribute the medical beauty project data with the same entity project name to the entity data set of the same entity;

[0036] Clean the data in the entity data set, and classify the data in the entity data set according to the dictionary interpretation category, the project comparison category, the case category, the medical beauty project category, the structured data category of medical beauty project attributes, and the consultation chat category;

[0037] Knowledge representation is performed based on the classification results of data in the entity dataset, including:

[0038] Extract entity semantic vectors based on lexicographical explanation data;

[0039] Construct an entity mutual exclusion table based on item comparison data;

[0040] Perform document vectorization based on case data;

[0041] Obtain standard entity screening data based on medical aesthetic item categories;

[0042] Extract item structure attributes based on data of structured data of medical aesthetic item attributes;

[0043] Extract item entity vectors based on consultation chat data.

[0044] A medical aesthetic entity alignment device, comprising:

[0045] A data acquisition module, which is used to acquire medical aesthetic item data;

[0046] An entity extraction module, which is used to extract entities based on medical aesthetic item data, and the entity attributes of the entities include at least one of entity semantic vectors, item entity vectors, and item structure attributes;

[0047] An entity classification module, which is used to screen entities to obtain a first standard entity set and a non-standard entity set;

[0048] An entity mapping module, which is used to construct a mapping key between non-standard entities in the non-standard entity set and first standard entities in the first standard entity set based on the similarity of entity attributes.

[0049] A computer device, comprising a memory and a processor. When a computer program stored in the memory is executed by the processor, the processor executes the steps of the above-mentioned medical aesthetic entity alignment method.

[0050] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the above-mentioned medical aesthetic entity alignment method.

[0051] Compared with the prior art, the beneficial effects of the present invention are:

[0052] The present invention aims to solve the problem that in the process of the rapid development of the medical beauty industry, there is no unified industry standard for medical beauty projects and no standardized naming of medical beauty projects. The present invention covers multi-party data at the underlying level and data of different types and structures, and the data is also authoritative and reliable enough, providing a reliable data basis for the subsequent entity alignment of medical beauty projects.

[0053] The whole process of the present invention is an automated link, which requires little manual participation, can update the specifications of project entities stably and in a timely manner, and achieves high timeliness and effectiveness. A variety of deep learning algorithms and traditional robotics learning algorithms are adopted in the present invention, and the technical solutions are more advanced and effective. Compared with the traditional method (at present, in most industries, the definition and determination of standard entities are mostly to hire industry experts to manually sort out and confirm the definition of standard entities and construct a standard entity set), this method is more efficient, has strong iterability, and is based on theoretical basis according to the content of data statistics (compared with manual definition).

[0054] In addition, in the present invention, after processing and sorting out the data obtained from different sources, deep learning, machine learning and statistical analysis are used for knowledge representation, so that these multi-source heterogeneous data can be associated with each other and can be effectively calculated and weighted, thereby obtaining a standard project entity set. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. Among them:

[0056] Figure 1 It is a schematic flow diagram of the method for entity alignment of medical beauty.

[0057] Figure 2 It is a schematic flow diagram of the process for screening entities to obtain the first standard entity set.

[0058] Figure 3 It is a schematic flow diagram of the process for statistically analyzing the frequencies of the entity project names of the first candidate entities in the first candidate set.

[0059] Figure 4 It is a schematic flow diagram of the process for calculating the weights of the second candidate entities in the second candidate set.

[0060] Figure 5 It is a schematic flow diagram of the key process for constructing a mapping between the non-standard entities in the non-standard entity set and the first standard entities in the first standard entity set based on the similarity of entity attributes.

[0061] Figure 6 It is a schematic flow diagram of the process for extracting entities based on medical beauty project data.

[0062] Figure 7 It is a schematic block diagram of the structure of a medical beauty entity alignment device.

[0063] Among them, 1 is the data acquisition module, 2 is the entity extraction module, 3 is the entity classification module, and 4 is the entity mapping module. Specific implementation manners

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0065] Refer to Figure 1 , in some embodiments, a medical beauty entity alignment method includes:

[0066] S101, collecting medical beauty project data;

[0067] Among them, the data sources generally come from three aspects:

[0068] Firstly, obtaining network data, mainly obtaining text data from Baidu Encyclopedia, professional medical beauty websites, and content sharing platforms (such as Xiaohongshu, Weibo, etc.);

[0069] Secondly, obtaining medical beauty project data of major medical beauty institutions across the country, and establishing cooperation with medical beauty institutions in major cities across the country to obtain the real medical beauty project data and project case data of major medical beauty institutions;

[0070] Thirdly, obtaining customer consultation data through a large amount of customer consultation data accumulated by a medical beauty consultation platform.

[0071] S102, extracting entities based on the medical beauty project data, and the entity attributes of the entities include at least one of an entity semantic vector, a project entity vector, and a project structure attribute;

[0072] Preferably, semantic extraction of entities is performed based on a model to obtain an entity semantic vector and a project entity vector.

[0073] Among them, for entity semantic vector extraction, a BERT model is used for semantic extraction, and the entire entry explanation is transformed into a 768-dimensional vector for representation. The transformed knowledge data is (medical beauty project - 768-dimensional vector), denoted as the project semantic matrix;

[0074] The BERT model essentially learns a good feature representation for words by running self-supervised learning methods on a vast amount of corpus. In this application, the vector of the last semantic representation layer of the BERT model is taken as the entity semantic vector.

[0075] Among them, for the extraction of project entity vectors, the BERT model + CRF model is used to extract medical beauty project entities in chat data, and the corresponding vectors are extracted and averaged to represent the entity in vector form. The transformed knowledge representation is (medical beauty project - 768-dimensional vector), denoted as the project entity vector.

[0076] The BERT model + CRF model is a secondary development based on open-source models. The network parameters of the trained BERT model are used to initialize our BERT model + CRF model, and this model is used for entity extraction and vector representation.

[0077] Among them, the project structure attributes refer to that general medical beauty projects all contain some common structure attributes, such as surgical methods, treatment courses, suitable populations, treatment times, effects, recovery times, postoperative rest times, whether to be hospitalized, etc. All medical beauty projects can form their own entity attribute graphs, and the transformed knowledge representation is the project structure attributes.

[0078] S103, screen the entities to obtain the first standard entity set and the non-standard entity set;

[0079] S104, construct a mapping key between the non-standard entities in the non-standard entity set and the first standard entities in the first standard entity set based on the similarity of entity attributes.

[0080] Preferably, after completing the entity alignment work, an entity mapping table, a standard entity knowledge representation table, a standard entity attribute graph, etc. can also be generated incidentally.

[0081] Refer to Figure 2 , in some embodiments, screening the entities to obtain the first standard entity set includes:

[0082] S201, conduct a preliminary screening of the entities, and add the entities with the data source being medical beauty institutions to the first candidate set;

[0083] Among them, the data in the first candidate set directly comes from the official medical beauty project data of medical beauty institutions. However, since there is no unified standard in the medical beauty industry, the naming of medical beauty projects by major medical beauty institutions is different. Projects that may be exactly the same may be named completely unrelated names, but these medical beauty projects are the most real medical beauty project entities, and the standard entities required for entity alignment in this application are generated from here.

[0084] S202. Perform frequency statistics on the entity item names of the first candidate entities in the first candidate set. If the frequency statistics result is greater than the first preset threshold, add the first candidate entities to the second standard entity set;

[0085] S203. Obtain the second candidate set by removing the second standard entity set from the first candidate set;

[0086] S204. Calculate the weights of the second candidate entities in the second candidate set. If the weight calculation result is greater than the second preset threshold, add the second candidate entities to the third standard entity set;

[0087] S205. Combine the second standard entity set and the third standard entity set to obtain the first standard entity set.

[0088] See Figure 3 , preferably, performing frequency statistics on the entity item names of the first candidate entities in the first candidate set includes:

[0089] S301. Determine the major project category to which the entity item name of the first candidate entity belongs;

[0090] S302. Obtain the number of medical beauty institutions with the major project category;

[0091] S303. Obtain the number of medical beauty institutions with the entity item name of the first candidate entity;

[0092] S304. Obtain the ratio of the number of medical beauty institutions with the entity item name of the first candidate entity to the number of medical beauty institutions with the major project category.

[0093] Specifically, the first preset threshold is 50%. Being greater than the first preset threshold means that more than half of the medical beauty institutions have adopted this entity item name.

[0094] For example, if the entity item name is double eyelid and its major project category is eye plastic surgery. Suppose it is statistically found that there are 5000 medical beauty institutions offering double eyelid projects and 6000 medical beauty institutions offering eye plastic surgery projects, then the ratio is 83.33%, which is greater than the first preset threshold of 50%, so double eyelid is taken as the standard entity.

[0095] See Figure 4 , preferably, calculating the weights of the second candidate entities in the second candidate set includes:

[0096] S401. Based on the entity recognition model, construct a mutual exclusion graph between the second candidate entities in the second candidate set, and perform weight sorting on the mutual exclusion graph to obtain the first weight of the second candidate entities;

[0097] Among them, for the construction of the mutually exclusive graph, the BERT model + CRF model for entity recognition is used to extract the medical beauty project entities mentioned in the data. These medical beauty projects belong to mutually exclusive medical beauty projects. Each piece of project comparison data can obtain a set of mutually exclusive medical beauty project entities. Then these data form a mutually exclusive table, and then a graph representing the mutual exclusion of medical beauty projects is generated based on this mutually exclusive table. The final knowledge representation of the project comparison data is a mutually exclusive graph.

[0098] Among them, the improved algorithm ProjectRank based on PageRank is used to sort the weights of the above-mentioned mutually exclusive graph to obtain the first weight of each entity project.

[0099] The ProjectRank algorithm is an improved algorithm based on PageRank. First, a mutually exclusive graph is constructed through the mutual exclusion relationship of entities. Then, the same initial value of 1 is set for all entity projects in the graph (by default, the weights of all entities are equal at the beginning). Then, a batch of standard entities is provided according to the second standard entity set, and the initial values of the entities corresponding to this batch of standard entities are all set to 2 (non-standard entities will eventually align with standard entities, so the weight of standard entities is set to twice that of non-standard entities). In the ProjectRank algorithm, it is considered that the result of the influence of a project entity on the entire medical beauty entity system is that the entities mutually exclusive with it also have a certain influence. The more projects mutually exclusive with this entity, the higher the frequency of comparison of this project, and the higher the recognition of people, and it should have a higher weight. Assuming the influence of a project is PR, then the influence of this project should be equal to the sum of the PRs of all projects mutually exclusive with it. The iterative algorithm still uses the method of the adjacency matrix of PageRank for PR value iteration. The difference is that PageRank is a directed graph, while our ProjectRank is an undirected graph. And our goal is to find new standard entities. Therefore, we don't need to iterate until the PR value converges like the PageRank algorithm. Only a limited number of iterations are required. In our algorithm, 3 iterations are selected (because the goal is to select from non-standard entities, so the number of iterations should be odd, and through actual calculation, 3 iterations have the best effect).

[0100] S402, obtaining the second standard entity with the highest similarity in the second standard entity set for the second candidate entity based on the similarity between the entity semantic vector and the project entity vector in the entity attributes. The similarity score between the second candidate entity and the second standard entity with the highest similarity to the entity semantic vector and the project entity vector is the second weight;

[0101] S403. Obtain the second standard entity with the highest similarity in the second standard entity set for the second candidate entity based on the similarity of the project structure attributes in the entity attributes. The similarity score between the second candidate entity and the second standard entity with the highest similarity in project structure attributes is the third weight.

[0102] S404. Subtract the second weight and the third weight of the second candidate entity from its first weight to obtain the weight difference.

[0103] Specifically, the second preset threshold is zero. Being greater than the second preset threshold means that the weight difference obtained by subtracting the second weight and the third weight from the first weight is positive, which satisfies the condition. Then, the second candidate entity can be added to the third standard entity set.

[0104] Refer to Figure 5 , in some embodiments, constructing a mapping between the non-standard entities in the non-standard entity set and the first standard entities in the first standard entity set based on the similarity of entity attributes mainly includes:

[0105] S501. Calculate the entity attribute similarity between the non-standard entity and the first standard entity. The entity attribute similarity includes at least one of the entity semantic vector similarity, the project entity vector similarity, and the project entity attribute similarity.

[0106] S502. Select the first standard entity with the highest similarity to the non-standard entity based on the entity attribute similarity to establish the mapping key.

[0107] In some embodiments, the calculation method of the similarity of entity attributes includes:

[0108] For the entity semantic vector or the project entity vector in the entity attributes, use the cosine similarity calculation method to calculate the similarity of the entity semantic vector or the project entity vector.

[0109] For the project structure attributes in the entity attributes, use the coincidence degree calculation method to calculate the similarity of the project structure attributes.

[0110] Among them, taking the included angle of the vectors as the consideration angle, the inner product of the vectors (the sum of the products of the corresponding elements) divided by the product of the norms of the two vectors is used as the calculation result. Since the cosine similarity represents the difference in direction and is not sensitive to distance, sometimes when the difference in distance is also concerned, a mean value will be subtracted from each value first, which is called the adjusted cosine similarity.

[0111] In addition to the cosine similarity calculation method, the commonly used vector similarity calculation methods after text vectorization also include the Manhattan distance, the Euclidean distance, the Pearson correlation coefficient, etc.

[0112] Among them, the calculation of the overlap degree of project structure attributes, the similarity of project structure attributes means that, for example, the standard entity has 20 structure attributes, the non-standard entity has 30 structure attributes, the set of structure attributes is 40, and 10 of them have the same attribute values, so the similarity is 10 / 40 = 0.25.

[0113] Refer to Figure 6 , in some embodiments, the entities extracted based on medical beauty project data include:

[0114] S601, attributing medical beauty project data with the same entity project name to the entity data set of the same entity;

[0115] S602, cleaning the data in the entity data set and classifying the data in the entity data set according to the types of glossary explanations, project comparisons, cases, medical beauty project classes, structured data classes of medical beauty project attributes, and consultation chat classes;

[0116] S603, performing knowledge representation based on the classification results of the data in the entity data set, including:

[0117] Extracting entity semantic vectors based on glossary explanation data;

[0118] Constructing an entity exclusive table based on project comparison data;

[0119] Performing document vectorization based on case data;

[0120] Obtaining standard entity screening data based on medical beauty project classes;

[0121] Extracting project structure attributes based on structured data of medical beauty project attributes;

[0122] Extracting project entity vectors based on consultation chat data.

[0123] Among them, after the data is obtained, the data is processed, the obtained data is cleaned and classified, and the obtained data is mainly classified into the following six categories:

[0124] 1. Glossary explanation class, mainly from Baidu Encyclopedia, professional medical beauty websites, and medical beauty institution project data; specifically, the BERT model is used to extract entity semantic vectors from glossary explanation data;

[0125] 2. Project comparison class, mainly from medical beauty websites and content platforms; specifically, the BERT model + CRF model is used to construct an entity exclusive table for project comparison data;

[0126] 3. Case class, mainly from medical beauty websites, content platforms, and hospital data; specifically, the TF-IDF model is used to perform document vectorization on case data;

[0127] Among them, the second standard entity with the highest similarity in the second standard entity set is calculated for the second candidate entity through TF-IDF similarity.

[0128] TF-IDF is a statistical method used to evaluate the importance of a word or term for a document set or a single document in a corpus. The importance of a word increases in direct proportion to the number of times it appears in a document, but decreases in inverse proportion to the frequency of its appearance in the corpus. The main idea of TF-IDF is that if a certain word appears frequently (high TF) in an article and rarely appears in other articles, then this word or phrase is considered to have good category discrimination ability and is suitable for classification.

[0129] Among them, the term frequency (TF) represents the frequency of a term (keyword) appearing in the text. This number is usually normalized (generally, the term frequency is divided by the total number of words in the article) to prevent it from biasing towards long documents.

[0130] Among them, the inverse document frequency (IDF): The IDF of a specific term can be obtained by dividing the total number of documents by the number of documents containing that term and then taking the logarithm of the resulting quotient.

[0131] If the fewer documents contain the term TF and the larger the IDF, it indicates that the term has good category discrimination ability. A high term frequency within a specific document and a low document frequency of that term in the entire document set can produce a high-weight TF-IDF. Therefore, TF-IDF tends to filter out common words and retain important words.

[0132] 4. Medical beauty project category, mainly from the project data of medical beauty institutions; it can be used to indicate the data source of the entity. For an entity with medical beauty project category data, its entity data source can be identified as coming from a medical beauty institution for screening standard entities.

[0133] 5. Structured data category of medical beauty project attributes, mainly from the project data of medical beauty institutions and professional medical beauty websites.

[0134] 6. Consultation chat category, mainly from customer consultation data; specifically, the BERT model + CRF model is used to extract project entity vectors for consultation chat category data.

[0135] In this embodiment, the purpose is to initially process the collected medical beauty project data in the early stage to facilitate subsequent entity alignment operations. By classifying and summarizing multiple data belonging to the same entity, it is possible to more easily and conveniently extract the corresponding data from the classified data for subsequent entity alignment operations, thereby accelerating the efficiency of entity alignment operations.

[0136] Of course, for the collected entity data set of entities, it is not necessarily the case that it contains all of the above six types of data. It may only contain one or several of these types. In this case, other types of data can also be used to extract the corresponding data.

[0137] Refer to Figure 7 , in some embodiments, the present application also discloses a medical beauty entity alignment device, including:

[0138] A data acquisition module 1, which is used to acquire medical beauty project data;

[0139] An entity extraction module 2, which is used to extract entities based on the medical beauty project data. The entity attributes of the entity include at least one of an entity semantic vector, a project entity vector, and a project structure attribute;

[0140] An entity classification module 3, which is used to screen the entities to obtain a first standard entity set and a non-standard entity set;

[0141] An entity mapping module 4, which is used to construct a mapping key between the non-standard entities in the non-standard entity set and the first standard entities in the first standard entity set based on the similarity of entity attributes.

[0142] In some embodiments, the present application also discloses a computer device, including a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor is caused to execute the steps of the above-mentioned medical beauty entity alignment method.

[0143] Among them, the computer device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through means such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device.

[0144] The memory at least includes one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or D interface display memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device. Of course, the memory may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is commonly used to store the operating system and various application software installed on the computer device, such as the program code of the medical beauty entity alignment method. In addition, the memory can also be used to temporarily store various data that have been output or will be output.

[0145] In some embodiments, the processor may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data, such as running the program code of the medical beauty entity alignment method.

[0146] In some embodiments, the present application also discloses a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the above-mentioned medical beauty entity alignment method.

[0147] Wherein, the computer-readable storage medium stores an interface display program, and the interface display program can be executed by at least one processor to cause the at least one processor to execute the steps of the program code of the medical beauty entity alignment method as described above.

[0148] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0149] The above is the embodiment of the present invention. The above embodiments and the specific parameters in the embodiments are only for clearly expressing the verification process of the invention and are not used to limit the patent protection scope of the present invention. The patent protection scope of the present invention still takes its claims as the criterion. Any equivalent structural changes made by using the content of the specification and drawings of the present invention should, by the same token, be included in the protection scope of the present invention.

Claims

1. A method for aligning medical beauty entities, characterized in that Including: Collecting medical beauty project data; Extracting entities based on the medical beauty project data, where the entity attributes of the entities include at least one of entity semantic vectors, project entity vectors, and project structure attributes; Screening the entities to obtain a first standard entity set and a non-standard entity set; Constructing mapping keys between the non-standard entities in the non-standard entity set and the first standard entities in the first standard entity set based on the similarity of entity attributes; Screening the entities to obtain the first standard entity set includes: Conducting a preliminary screening of the entities and adding the entities with the data source being medical beauty institutions to the first candidate set; Counting the frequencies of the entity project names of the first candidate entities in the first candidate set. If the frequency count result is greater than a first preset threshold, adding the first candidate entity to the second standard entity set; Removing the second standard entity set from the first candidate set to obtain a second candidate set; Calculating the weights of the second candidate entities in the second candidate set. If the weight calculation result is greater than a second preset threshold, adding the second candidate entity to the third standard entity set; Combining the second standard entity set and the third standard entity set to obtain the first standard entity set; Calculating the weights of the second candidate entities in the second candidate set includes: Constructing a mutual exclusion graph between the second candidate entities in the second candidate set based on an entity recognition model, and performing weight sorting on the mutual exclusion graph to obtain the first weight of the second candidate entities; Obtaining the second standard entity with the highest similarity in the second standard entity set for the second candidate entity based on the similarity of the entity semantic vector and the project entity vector in the entity attributes. The similarity score between the second candidate entity and the second standard entity with the highest similarity of the entity semantic vector and the project entity vector is the second weight; Obtaining the second standard entity with the highest similarity in the second standard entity set for the second candidate entity based on the similarity of the project structure attribute in the entity attributes. The similarity score between the second candidate entity and the second standard entity with the highest similarity of the project structure attribute is the third weight; Subtracting the second weight and the third weight of the second candidate entity from its first weight to obtain a weight difference.

2. The medical aesthetic entity alignment method according to claim 1, wherein Counting the frequencies of the entity project names of the first candidate entities in the first candidate set includes: Determining the major project category to which the entity project name of the first candidate entity belongs; Obtaining the number of medical beauty institutions with the major project category; Obtaining the number of medical beauty institutions with the entity project name of the first candidate entity; Obtaining the ratio of the number of medical beauty institutions with the entity project name of the first candidate entity to the number of medical beauty institutions with the major project category.

3. The medical aesthetic entity alignment method according to claim 1, characterized in that, Constructing mapping keys between the non-standard entities in the non-standard entity set and the first standard entities in the first standard entity set based on the similarity of entity attributes includes: Calculate the entity attribute similarity between the non-standard entity and the first standard entity, where the entity attribute similarity includes at least one of entity semantic vector similarity, item entity vector similarity, and item entity attribute similarity; Based on the entity attribute similarity, select the first standard entity with the highest similarity to the non-standard entity to establish a mapping key.

4. The medical aesthetic entity alignment method according to claim 1, wherein The calculation method of the similarity of the entity attributes includes: For the entity semantic vector or item entity vector in the entity attributes, use the cosine similarity calculation method to calculate the similarity of the entity semantic vector or item entity vector; For the item structure attributes in the entity attributes, use the coincidence degree calculation method to calculate the similarity of the item structure attributes.

5. The medical aesthetic entity alignment method according to claim 1, characterized in that Extracting entities based on the medical beauty project data includes: Classify the medical beauty project data with the same entity item name into an entity dataset of the same entity; Clean the data in the entity dataset, and classify the data in the entity dataset according to the glossary explanation category, item comparison category, case category, medical beauty project category, medical beauty project attribute structured data category, and consultation chat category; Perform knowledge representation based on the classification results of the data in the entity dataset, including: Extract entity semantic vectors based on the glossary explanation category data; Construct an entity exclusion table based on the item comparison category data; Perform document vectorization based on the case category data; Obtain standard entity screening data based on the medical beauty project category; Extract item structure attributes based on the medical beauty project attribute structured data category data; Extract item entity vectors based on the consultation chat category data.

6. A medical beauty entity alignment device, characterized in that: A data collection module, which is used to collect medical beauty project data; An entity extraction module, which is used to extract entities based on the medical beauty project data, and the entity attributes of the entity include at least one of entity semantic vectors, item entity vectors, and item structure attributes; An entity classification module, which is used to screen the entities to obtain a first standard entity set and a non-standard entity set; An entity mapping module, which is used to construct a mapping key between the non-standard entities in the non-standard entity set and the first standard entities in the first standard entity set based on the similarity of entity attributes; In the entity classification module, screening the entities to obtain the first standard entity set includes: Perform a preliminary screening on the entities, and add the entities whose data source is a medical beauty institution to the first candidate set; Perform frequency statistics on the entity item names of the first candidate entities in the first candidate set. If the frequency statistics result is greater than the first preset threshold, add the first candidate entity to the second standard entity set; Remove the second standard entity set from the first candidate set to obtain a second candidate set; Calculate the weights of the second candidate entities in the second candidate set. If the weight calculation result is greater than the second preset threshold, add the second candidate entity to the third standard entity set; Obtain the first standard entity set by combining the second standard entity set and the third standard entity set; The weight calculation for the second candidate entities in the second candidate set includes: Construct a mutual exclusion graph between the second candidate entities in the second candidate set based on the entity recognition model, and perform weight sorting on the mutual exclusion graph to obtain the first weight of the second candidate entities; Obtain the second standard entity with the highest similarity in the second standard entity set for the second candidate entity based on the similarity between the entity semantic vector and the project entity vector in the entity attributes, and the similarity score between the second candidate entity and the second standard entity with the highest similarity in the entity semantic vector and the project entity vector is the second weight; Obtain the second standard entity with the highest similarity in the second standard entity set for the second candidate entity based on the similarity of the project structure attributes in the entity attributes, and the similarity score between the second candidate entity and the second standard entity with the highest similarity in the project structure attributes is the third weight; Subtract the second weight and the third weight of the second candidate entity from its first weight to obtain the weight difference.

7. A computer device, characterized in that, Including a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the medical beauty entity alignment method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, Stores a computer program, and when the computer program is executed by a processor, the processor executes the steps of the medical beauty entity alignment method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Entity linking method and device based on semantic components

    CN111613341A