Institutional standardization method, device, electronic device and storage medium
By identifying and standardizing the institutional information in scientific research literature and building a knowledge graph, the inaccuracy of information and time-consuming query caused by irregular organization names are solved, and the query efficiency and data statistics are improved.
Patent Information
- Application Number
- CN202010417022.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-15
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2040-05-15
AI Technical Summary
When storing and processing a large number of scientific research literature, the writing errors and irregularities of the organization's name lead to inaccurate information, the query and processing take time, and the relevant data calculation and information statistics are not accurate enough.
By obtaining sub-organization fields in organization information, text recognition technology is used to identify and determine the area category hierarchy and sub-organization levels, a knowledge graph is constructed, and the sub-organization fields are standardized using the editing distance algorithm.
It improves the efficiency and accuracy of document and information query, reduces the time for query and processing, and enhances the accuracy of data operations and information statistics.
Smart Images

Figure CN111694823B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method, device, electronic device, and storage medium for institutional standardization. Background Art
[0002] With the development of technology, we have entered an era of information explosion, even in highly specialized scientific research fields. For professional scientific researchers, they often need to read a large number of specialized papers and pay attention to outstanding researchers and research institutions in the industry.
[0003] To focus on important research institutions in a certain research field, the first step is to identify the institution itself. However, in many documents and information, the writing of institution names is often incorrect or irregular, resulting in inaccurate information. In the vast amount of data in the storage system, document or information query and processing are time-consuming, and related data operations and information statistics are also inaccurate. Summary of the Invention
[0004] To solve the above problems, this application provides a method, device, electronic device, and storage medium for institutional standardization, which is conducive to improving the efficiency and accuracy of document and information query and processing.
[0005] In the first aspect of the embodiments of this application, a method for institutional standardization is provided. The method includes:
[0006] Obtain the sub-institution fields in the institutional information, use text recognition technology to recognize each sub-institution field in the sub-institution fields, and determine the regional category level corresponding to each sub-institution field;
[0007] Determine the sub-institution level corresponding to each sub-institution field;
[0008] Use the lowest level among the sub-institution levels corresponding to each sub-institution field as the institutional level of the institutional information, and store the institutional level as a label of the institutional information to complete the construction of the knowledge graph;
[0009] Perform standardization processing on each of the sub-institution fields using the edit distance algorithm.
[0010] In combination with the first aspect, in a possible implementation manner, the performing standardization processing on each of the sub-institution fields using the edit distance algorithm includes:
[0011] Sort each of the sub-institution fields according to the number of each of the sub-institution fields;
[0012] Obtain the edit distance between each of the sub-institution fields;
[0013] Merge the sub - institution fields with an edit distance less than the distance threshold.
[0014] Combined with the first aspect, in a possible implementation manner, the merging process of the sub - institution fields with an edit distance less than the distance threshold includes:
[0015] Store the target sub - institution field with the largest quantity among the sub - institution fields with an edit distance less than the distance threshold as the standardized name of each sub - institution field;
[0016] Before obtaining the sub - institution fields in the institution information and using text recognition technology to recognize each sub - institution field in the sub - institution fields and determine the corresponding regional category level of each sub - institution field, the method further includes:
[0017] Perform data cleaning on the institution data submitted by the terminal to remove noise information.
[0018] Combined with the first aspect, in a possible implementation manner, the data cleaning of the institution data submitted by the terminal to remove noise information includes:
[0019] Extract the institution information and author information from the institution data through semantic recognition technology;
[0020] Match and correct the author information using a preset personal name abbreviation template; and
[0021] Identify the preset conjunctions and preset nouns in the institution information, split the institution information into multiple fields based on the preset conjunctions and the preset nouns, and add preset punctuation marks between the fields.
[0022] Combined with the first aspect, in a possible implementation manner, the method further includes:
[0023] Match the standard name of each sub - institution field according to the corresponding regional category level and sub - institution level of each sub - institution field to obtain a matching result;
[0024] Perform a correction operation on the institution information according to the matching result to obtain standardized institution information.
[0025] Combined with the first aspect, in a possible implementation manner, the method further includes:
[0026] If it is recognized that there are sub - institution fields with the same sub - institution level in the institution information, then reduce the same sub - institution level by one level as the institution level.
[0027] The second aspect of the embodiments of the present application provides an institution standardization device, and the device includes:
[0028] A data acquisition module, configured to acquire sub - institution fields in institutional information, identify each sub - institution field in the sub - institution fields by using text recognition technology, and determine the corresponding regional category level of each sub - institution field;
[0029] A level determination module, configured to determine the sub - institution level corresponding to each sub - institution field;
[0030] A knowledge graph construction module, configured to use the lowest level in the sub - institution levels corresponding to each sub - institution field as the institutional level of the institutional information, store the institutional level as a label of the institutional information, so as to complete the construction of the knowledge graph;
[0031] A standardization module, configured to perform standardization processing on each of the sub - institution fields by using the edit distance algorithm.
[0032] In the third aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes an input device and an output device, and further includes a processor adapted to implement one or more instructions; and,
[0033] A computer storage medium, where the computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the steps in the method described in the first aspect above.
[0034] In the fourth aspect of the embodiments of the present application, a computer storage medium is provided. The computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by a processor to perform the steps in the method described in the first aspect above.
[0035] Compared with the prior art, in the embodiments of the present application, by acquiring sub - institution fields in institutional information, identifying each sub - institution field in the sub - institution fields by using text recognition technology, determining the corresponding regional category level of each sub - institution field; determining the sub - institution level corresponding to each sub - institution field; using the lowest level in the sub - institution levels corresponding to each sub - institution field as the institutional level of the institutional information, storing the institutional level as a label of the institutional information to complete the construction of the knowledge graph; performing standardization processing on each of the sub - institution fields by using the edit distance algorithm. In this way, a large amount of institutional data is used to construct a knowledge graph, the standardized institutional level is stored as a label of the institutional information, and at the same time, each sub - institution field is standardized by the edit distance algorithm. The stored is a common standard name. In subsequent applications for searching institutions, the corresponding standardized institutional names can be matched through labels of the same institutional level, which is beneficial to improving the query efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0037] Figure 1 A network system architecture diagram provided by an embodiment of the present application;
[0038] Figure 2 A flowchart of a method for institutional standardization provided by an embodiment of the present application;
[0039] Figure 3 An example diagram of a regional category level provided by an embodiment of the present application;
[0040] Figure 4 An example diagram of an institutional level provided by an embodiment of the present application;
[0041] Figure 5 An example diagram for determining an institutional level provided by an embodiment of the present application;
[0042] Figure 6 A flowchart of another method for institutional standardization provided by an embodiment of the present application;
[0043] Figure 7 A structural diagram of an institutional standardization device provided by an embodiment of the present application;
[0044] Figure 8 A structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0045] To enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0046] The terms "comprising" and "having" and any variations thereof that appear in the specification, claims, and drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices. In addition, terms such as "first", "second", and "third" are used to distinguish different objects, rather than to describe a specific order.
[0047] The embodiments of this application provide an institutional standardization solution. The so-called institutional standardization is to find the most standard names for scientific research institutions or other entities. This solution is implemented using medical literature as the data set, constructing a standardized data structure for schools, hospitals, laboratories, etc. in the literature, which is beneficial for reducing the time-consuming of document or information query and processing in the large amount of data in the storage system, can more quickly match the accurate institutional name, and determine the institutional level, etc., making the relevant data operations and information statistics more accurate. Of course, in some cases, it can also be implemented using institutional information in other categories of literature or on personal home pages on the web, with a wide range of applications. After subsequent online testing, the accuracy of matching for scientific research institutions reached more than 90%, and the performance of geographical location could reach more than 95%.
[0048] Specifically, this institutional standardization solution can be based on Figure 1 the network system architecture shown in Figure 1 As shown, this network system architecture at least includes a terminal and a server. The entire network system is connected through a wired or wireless network. Parts of the network system not shown may also include a database, a repeater, a switch, etc. The terminal is used to submit a knowledge graph construction request to the server during the knowledge graph construction stage, and this request may include institutional data for constructing the knowledge graph; while in the online standardization stage (application stage), the terminal is used to submit a standardization request to the server, and this request may include institutional data to be matched or standardized. The server is the execution entity of this solution. In some embodiments, the server can perform related steps such as data cleaning of institutional data, sub-institutional field identification, sub-institutional level determination, and edit distance calculation for the knowledge graph construction request submitted by the terminal. Various algorithms such as text recognition and edit distance calculation are integrated in the server to support the implementation of this solution. It can be understood that the terminal in this application can be devices such as a computer, a tablet computer, or a smart phone, and the server can be a local server or a cloud server. Figure 1 This is just an example and does not impose any limitations on the embodiments of this application. In some cases, this solution can also be implemented based on other network architectures, such as: blockchain network.
[0049] Based on Figure 1The network system architecture shown below will be used to elaborate in detail on the institutional standardization method proposed in the embodiments of the present application in combination with relevant attached drawings. Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an institutional standardization method provided in an embodiment of the present application. As Figure 2 shown, it includes steps S21 - S24:
[0050] Step S21: Obtain the sub - institutional fields in the institutional information, use text recognition technology to recognize each sub - institutional field in the sub - institutional fields, and determine the corresponding regional category level of each sub - institutional field;
[0051] In the embodiments of the present application, institutional information refers to the institutional name in institutional data. After basic data cleaning, the sub - institutional fields of different parts in the institutional information can be extracted, such as country, province / state, city university, college affiliated to the university, center, laboratory, etc. In specific implementation, as Figure 3 shown, multiple regional category levels can be preset, and multiple institutional names at the corresponding levels are stored under each regional category level. Specifically, three regional category levels are constructed, corresponding to the word levels of countries, states (provinces), and cities globally, for data matching and correction of institutional information.
[0052] Among them, the obtained sub - institutional fields are matched with the pre - stored institutional names to determine the corresponding regional category level of each sub - institutional field.
[0053] Optionally, before step S21, the institutional data submitted by the terminal is cleaned to remove meaningless noise information, such as special symbols, meaningless words (and, from, etc.). Specifically, through preliminary semantic recognition, institutional information and author information can be respectively extracted. For the names of people in the author information, an abbreviated form is used. The template in the preset database of abbreviated names of people can be used for matching and modification. PubMed is a database that provides searches for biomedical papers and abstracts and is free to search. In PubMed literature, abbreviated names of people are used to indicate different author information. The standard name format of the preset rules for abbreviated names of people in the present application can adopt the abbreviated names of people in the above - mentioned PubMed literature.
[0054] Optionally, all existing literature can be downloaded from PubMed, and then the above - mentioned institutional (affiliation) data can be extracted from it.
[0055] Generally, the information of different paper authors is written together, so it is necessary to split the information of different authors. Specifically, based on the above - mentioned format of abbreviated names of people, the author information is standardized. The author information can be matched and automatically corrected based on the preset rules for abbreviated names of people, or the literature can be associated with the correct abbreviated name label of people.
[0056] For the situation where general nouns in the obtained institutional information are written together, index splitting, semicolon splitting, etc. can be adopted. For example, for "NewYorkCity", it needs to be split, which is the result of in-depth observation of the data.
[0057] In this application, a preset noun library can be established based on a large number of existing sub-institutional fields, storing a large number of commonly used preset nouns, and these data can be certified and sorted. The institutional information can be divided according to the preset nouns and preset conjunctions. Specifically, for a piece of text, when an institutional noun is recognized, the server uses text recognition technology to recognize and extract the preset nouns from it as the splitting fields. For example, it includes certified institutional nouns such as "TshinghuaUniversity", and sorted common nouns such as "school of medicine", etc.; for the remaining fields that cannot correspond to the preset nouns, multiple preset conjunctions such as "of", "and", etc. can be recognized, and then the splitting program is executed: when it is recognized that there are at least two independent nouns between two preset conjunctions, the splitting is performed at the position between the two independent nouns. Among them, for the multiple fields after splitting, a punctuation mark "," is added;
[0058] When there is only one noun between two preset conjunctions, the nouns before and after the preset conjunction are further recognized to determine the institutional type fields among them, such as institutional types like "school", "hospital", etc. For such nouns that can be confirmed as institutional types, the noun after "of" and this noun are classified into one division field; while the field connected with it by "and" is classified into another field.
[0059] For example, set the preset conjunction "of" and the "and" pattern splitting. For example, for "school of medicine of Tshinghua University", when the preset noun "Tshinghua University" is recognized and determined as a division field, the "of" before it can be replaced by ","; while when "school of medicine" is used as a preset noun, it can be directly divided. If it is not recorded as a preset noun, first, the institution type field "school" in it is recognized, and there is a preset conjunction "of" after it. Thus, a noun after "of" is used as its modification and divided into a field "school of medicine". Then, this institution noun is split into multiple fields: school of medicine, Tsinghua university. Another example is "Beijing Biology institute and Beijing Medical Center", which can be processed similarly. The institution type fields "institute" and "center" are recognized and divided into two parts: Beijing Biology institute and Beijing Medical Center through "and".
[0060] Step S22, determine the sub-institution level corresponding to each sub-institution field;
[0061] In a specific embodiment of the present application, while determining the regional category hierarchy, as Figure 4 shown, three sub-institution levels are constructed, that is, the institution is divided into three levels. For example, schools and hospitals become first-level institutions, colleges, branches, etc. become second-level institutions, departments, laboratories, etc. become third-level institutions, and these sub-institution levels can have a subordinate relationship.
[0062] Optionally, the sub-institution fields can be first subjected to field matching, and after determining the standardized sub-institution fields, the hierarchy and level are determined.
[0063] For example, the geographical locations of many countries are written in abbreviations. For example, California is written as CA. Through the preset abbreviation mapping relationship, the standardized sub-institution field corresponding to this abbreviation can be matched.
[0064] Optionally, if it is recognized that there are sub-institution fields with the same sub-institution level in the institution information, the same sub-institution level is lowered by one level as the institution level of the institution information. For example, after dividing the institution information into multiple fields, as Figure 5As shown in the figure, for each complete institutional information field A, if it includes two recognizable sub-institutions b and c, the sub-institutional levels of b and c can be obtained through the sub-institution database. When it is detected that the sub-institutional levels of both b and c are N, the level of the institutional information field A is determined to be N-1. For example, when a piece of institutional information includes an affiliated hospital (a first-level institution) and a school (a first-level institution), it will become a second-level institution. For instance, for Ruijin Hospital, Shanghai Jiao Tong University, it is recognized that "Shanghai Jiao Tong University" is a "university", belonging to the first level, and "Ruijin Hospital" is a "hospital", belonging to the first level. The institutional level of "Ruijin Hospital, Shanghai Jiao Tong University" is determined to be a second-level institution.
[0065] Step S23: Take the lowest level among the sub-institutional levels corresponding to each sub-institutional field as the institutional level of the institutional information, and store the institutional level as the label of the institutional information to complete the construction of the knowledge graph.
[0066] In a specific embodiment of the present application, after determining the sub-institutional levels corresponding to each sub-institutional field, the lowest level can be taken as the institutional level of the institutional information and stored in the form of a label. For institutions not in the database, records can be automatically stored as new institutional information to expand the database information volume. After that, in the application of searching for institutions, in a similar way, the institutional level of the institutional information input by the user can be determined, and the corresponding standardized institutional name can be matched through the label of the same institutional level to improve the query efficiency and accuracy.
[0067] Optionally, in an embodiment of the present application, the regional category level and the sub-institutional level can also be used as the labels of the institutional information and then stored.
[0068] Optionally, in an embodiment of the present application, according to the regional category level and the sub-institutional level corresponding to each sub-institutional field, the standard name of each sub-institutional field can be matched to obtain a matching result.
[0069] Perform a correction operation on the obtained institutional information according to the matching result to obtain standardized institutional information.
[0070] Step S24: Perform a standardization process on each of the sub-institutional fields by using the edit distance algorithm.
[0071] In a specific embodiment of the present application, after constructing the knowledge graph by using steps S21 - S23, continue to standardize each sub-institutional field. The edit distance algorithm can be used to perform a merging process on the sub-institutional fields.
[0072] Optionally, since different people may write the same institution differently. For example, for Shanghai Jiao Tong University, some people may write Jiao Tong University. Therefore, in some embodiments, the TF-IDF (term frequency–inverse document frequency) algorithm can also be used for subsequent standardization processing of the sub-institution fields.
[0073] It can be seen that in the embodiments of the present application, by obtaining the sub-institution fields in the institution information, using text recognition technology to recognize each sub-institution field in the sub-institution fields, determining the regional category level corresponding to each sub-institution field; determining the sub-institution level corresponding to each sub-institution field; taking the lowest level among the sub-institution levels corresponding to each sub-institution field as the institution level of the institution information, and storing the institution level as the label of the institution information to complete the construction of the knowledge graph; using the edit distance algorithm to standardize each sub-institution field. In this way, a large amount of institution data is used to construct the knowledge graph, and the standardized institution level is stored as the label of the institution information. At the same time, each sub-institution field is standardized by the edit distance algorithm, and the stored is the common standard name. In subsequent applications for finding institutions, the corresponding standardized institution names can be matched through labels of the same institution level, which is beneficial to improving the query efficiency and accuracy.
[0074] Please refer to Figure 6 , Figure 6 which is a schematic flow chart of another institution standardization method provided by the embodiments of the present application. As Figure 6 shown, it includes steps S61 - S66:
[0075] Step S61, obtain the sub-institution fields in the institution information, use text recognition technology to recognize each sub-institution field in the sub-institution fields, and determine the regional category level corresponding to each sub-institution field;
[0076] Step S62, determine the sub-institution level corresponding to each sub-institution field;
[0077] Step S63, take the lowest level among the sub-institution levels corresponding to each sub-institution field as the institution level of the institution information, and store the institution level as the label of the institution information to complete the construction of the knowledge graph;
[0078] Step S64, sort each sub-institution field according to the number of each sub-institution field;
[0079] Step S65, obtain the edit distance between each sub-institution field;
[0080] Step S66, merging the sub-organization fields whose edit distance is less than the distance threshold.
[0081] In a specific embodiment of the present application, the edit distance is a quantitative measurement of the degree of difference between two character strings, and the measurement method is to see how many times at least one processing is required to transform one character string into another character string. The edit distance can be used in natural language processing. For example, spell checking can determine which one (or which ones) is a more likely word based on the edit distance between a misspelled word and other correct words. The edit distance between each of the sub-organization fields can be understood as the similarity between each sub-organization field, that is, the similarity between the sub-organization field and the corresponding sub-organization standard name (which may be the correct spelling). Specifically, because some organizations may be written incorrectly due to human factors, the edit distance is used for standardization. Specifically, the data is sorted by quantity, and then the similarity is measured according to the edit distance. The organizations with an edit distance less than the above-mentioned distance threshold (such as 3) are merged, and the target sub-organization field with the largest number of sub-organization fields with an edit distance less than the distance threshold is stored as the standardized name of each of the sub-organization fields. For example, the sub-institution fields used to represent Shanghai Jiao Tong University may include Shanghai Jiao Tong University, Jiao Tong University, Shanghai Jiao Tong University, Jiao Tong University, etc., and Shanghai Jiao Tong University has the largest number of sub-institution fields. Therefore, Shanghai Jiao Tong University is used as the standardized name for the sub-institution fields representing Shanghai Jiao Tong University.
[0082] Optional, because institutions are hierarchical, such as Jiaotong University-School of Computer Science-Department of Software Engineering, etc., different people have different ways of writing, so it is necessary to give a "standard way of writing" (the way most people write). Therefore, the phenomenon of skipping levels at different levels is corrected. For example, in the above example, the School of Computer Science will not be written. After the query and matching of this solution, the missing institution will be filled.
[0083] It should be noted that Figure 6 Some steps in the embodiment shown are in Figure 2 Relevant descriptions have been made in the embodiments shown, and will not be repeated here.
[0084] In the application stage, the process of online standardization of documents and information is similar to the knowledge graph construction stage. When a new institutional data comes in, it will be cleaned and then extracted to obtain the sub-institution fields (i.e., as in the aforementioned steps S61 and S62). The sub-institution fields obtained can then be entered into the knowledge base for matching. After selecting some candidate institutions, they are sorted and finally the best candidate is selected. When matching, the similarity between the candidate institution and the institution to be matched, the consistency of the geographic information, and other measures can be used. Optionally, when the matching criteria are not met, it can be considered as an institution outside the knowledge base, so the extracted information will be directly determined as its standardized institution.
[0085] Based on the description of the above method embodiments, an embodiment of the present application further provides an institutional standardization device, which may be a computer program (including program code) running in a terminal. This institutional standardization device can execute Figure 2 or Figure 6 the method shown. Please refer to Figure 7 , this device includes:
[0086] A data acquisition module 71, configured to acquire sub-institutional fields in institutional information, identify each sub-institutional field in the sub-institutional fields by using text recognition technology, and determine the corresponding regional category level of each sub-institutional field;
[0087] A level determination module 72, configured to determine the sub-institutional level corresponding to each sub-institutional field;
[0088] A knowledge graph construction module 73, configured to use the lowest level among the sub-institutional levels corresponding to each sub-institutional field as the institutional level of the institutional information, and store the institutional level as a label of the institutional information to complete the construction of the knowledge graph;
[0089] A standardization module 74, configured to perform standardization processing on each of the sub-institutional fields by using an edit distance algorithm.
[0090] In an alternative embodiment, in terms of performing standardization processing on each of the sub-institutional fields by using an edit distance algorithm, the standardization module 74 is specifically configured to:
[0091] Sort each of the sub-institutional fields according to the number of each of the sub-institutional fields;
[0092] Obtain the edit distance between each of the sub-institutional fields;
[0093] Perform merging processing on each of the sub-institutional fields with an edit distance less than a distance threshold.
[0094] In an alternative embodiment, in terms of performing merging processing on each of the sub-institutional fields with an edit distance less than a distance threshold, the standardization module 74 is specifically configured to:
[0095] Store the target sub-institutional field with the largest number among each of the sub-institutional fields with an edit distance less than a distance threshold as the standardized name of each of the sub-institutional fields;
[0096] The data acquisition module 71 is further configured to: perform data cleaning on the institutional data submitted by the terminal to remove noise information.
[0097] In an optional implementation, in terms of cleaning the institutional data submitted by the terminal and removing noise information, the data acquisition module 71 is specifically used to:
[0098] Extract the institution information and author information from the institution data through semantic recognition technology;
[0099] Using a preset name abbreviation template to match and correct the author information; and
[0100] Identify the preset conjunctions and the preset nouns in the organization information, split the organization information into multiple fields based on the preset conjunctions and the preset nouns, and add preset punctuation marks between the fields.
[0101] In an optional implementation, the graph construction module 73 is further used to: match the standard name of each sub-organization field according to the regional category level and sub-organization level corresponding to each sub-organization field to obtain a matching result;
[0102] According to the matching result, the organization information is corrected to obtain standardized organization information.
[0103] In an optional implementation, the level determination module 72 is further configured to: if a sub-organization field with the same sub-organization level is identified in the organization information, reduce the same sub-organization level by one level as the organization level.
[0104] According to one embodiment of the present application, Figure 7 The various units in the mechanism standardization device shown can be combined into one or several other units separately or in whole, or one (or some) of the units can be further divided into multiple functionally smaller units, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present invention, the mechanism standardization device can also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.
[0105] According to another embodiment of the present application, the program can be executed by running on a general computing device such as a computer including a central processing unit (CPU), a random access memory medium (RAM), a read-only memory medium (ROM) and other processing elements and storage elements. Figure 2 or Figure 6 The computer program (including program code) of each step involved in the corresponding method shown in is used to construct Figure 7 The device shown, and to implement the above method of the embodiments of the present invention. The computer program can be recorded on, for example, a computer-readable recording medium, and loaded into the above computing device through the computer-readable recording medium and run therein.
[0106] Based on the description of the above method embodiments and device embodiments, embodiments of the present invention also provide an electronic device. Please refer to Figure 8 , the electronic device at least includes a processor 81, an input device 82, an output device 83, and a computer storage medium 84. Among them, the processor 81, the input device 82, the output device 83, and the computer storage medium 84 in the electronic device can be connected through a bus or other means.
[0107] The computer storage medium 84 can be stored in the memory of the electronic device. The computer storage medium 84 is used to store a computer program. The computer program includes program instructions. The processor 81 is used to execute the program instructions stored in the computer storage medium 84. The processor 81 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the electronic device, and is adapted to implement one or more instructions, specifically adapted to load and execute one or more instructions to implement the corresponding method flow or corresponding function.
[0108] In one embodiment, the processor 81 of the electronic device provided by the embodiments of the present application can be used to perform a series of mechanism standardization processes, including:
[0109] Obtain the sub-mechanism fields in the mechanism information, use text recognition technology to recognize each sub-mechanism field in the sub-mechanism fields, and determine the regional category level corresponding to each sub-mechanism field;
[0110] Determine the sub-mechanism level corresponding to each sub-mechanism field;
[0111] Take the lowest level among the sub-mechanism levels corresponding to each sub-mechanism field as the mechanism level of the mechanism information, and store the mechanism level as a label of the mechanism information to complete the construction of the knowledge graph;
[0112] Perform standardization processing on each of the sub-mechanism fields using the edit distance algorithm.
[0113] In one embodiment, the processor 81 executes the step of performing standardization processing on each of the sub-mechanism fields using the edit distance algorithm, including:
[0114] Sort each of the sub-mechanism fields according to the number of each of the sub-mechanism fields;
[0115] Obtain the edit distance between each of the sub-mechanism fields;
[0116] Merge the sub-institution fields with an edit distance less than the distance threshold.
[0117] In one embodiment, the processor 81 executes the merging process for the sub-institution fields with an edit distance less than the distance threshold, including:
[0118] Store the target sub-institution field with the largest quantity among the sub-institution fields with an edit distance less than the distance threshold as the standardized name of each sub-institution field;
[0119] The processor 81 is also used to execute: perform data cleaning on the institution data submitted by the terminal to remove noise information.
[0120] In one embodiment, the processor 81 executes the data cleaning on the institution data submitted by the terminal to remove noise information, including:
[0121] Extract the institution information and author information from the institution data through semantic recognition technology;
[0122] Match and correct the author information using a preset personal name abbreviation template; and
[0123] Identify the preset conjunction and preset noun in the institution information, split the institution information into multiple fields based on the preset conjunction and the preset noun, and add a preset punctuation mark between the fields.
[0124] In one embodiment, the processor 81 is also used to execute: match the standard name of each sub-institution field according to the regional category level and sub-institution level corresponding to each sub-institution field to obtain a matching result;
[0125] Perform a correction operation on the institution information according to the matching result to obtain standardized institution information.
[0126] In one embodiment, the processor 81 is also used to execute: if it is recognized that there are sub-institution fields with the same sub-institution level in the institution information, reduce the same sub-institution level by one level as the institution level.
[0127] In the embodiment of the present application, by obtaining the sub - institution fields in the institution information, using text recognition technology to recognize each sub - institution field in the sub - institution fields, determining the regional category level corresponding to each sub - institution field; determining the sub - institution level corresponding to each sub - institution field; taking the lowest level among the sub - institution levels corresponding to each sub - institution field as the institution level of the institution information, and storing the institution level as a label of the institution information to complete the construction of the knowledge graph; using the edit distance algorithm to standardize each of the sub - institution fields. In this way, a large amount of institution data is used to construct the knowledge graph, and the standardized institution level is stored as the label of the institution information. At the same time, the edit distance algorithm is used to standardize each sub - institution field, and the stored is the common standard name. In the subsequent application of searching for institutions, the corresponding standardized institution name can be matched through the label of the same institution level, which is beneficial to improving the query efficiency and accuracy.
[0128] Exemplarily, the above - mentioned electronic device can be a smart phone, a computer, a laptop, a tablet computer, a palm computer, a server, etc. The electronic device may include but is not limited to a processor 81, an input device 82, an output device 83, and a computer storage medium 84. Those skilled in the art can understand that the schematic diagram is only an example of the electronic device, and does not constitute a limitation on the electronic device. It may include more or fewer components than shown, or combine some components, or different components.
[0129] It should be noted that since the processor 81 of the electronic device implements the steps in the above - mentioned institution standardization method when executing the computer program, the embodiments of the above - mentioned institution standardization method are all applicable to this electronic device and can achieve the same or similar beneficial effects.
[0130] The embodiments of the present application also provide a computer storage medium (Memory). The computer storage medium is a memory device in an electronic device and is used to store programs and data. It can be understood that the computer storage medium here can include both the built-in storage medium in the terminal and, of course, the extended storage medium supported by the terminal. The computer storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, in this storage space, one or more instructions suitable for being loaded and executed by the processor 81 are also stored. These instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer storage medium located far from the aforementioned processor 81. In one embodiment, one or more instructions stored in the computer storage medium can be loaded and executed by the processor 81 to implement the corresponding steps of the above-mentioned institutional standardization method; in a specific implementation, one or more instructions in the computer storage medium are loaded and executed by the processor 81 to perform the following steps:
[0131] Obtain the sub-institutional fields in the institutional information, use text recognition technology to recognize each sub-institutional field in the sub-institutional fields, and determine the corresponding regional category level of each sub-institutional field;
[0132] Determine the sub-institutional level corresponding to each sub-institutional field;
[0133] Take the lowest level among the sub-institutional levels corresponding to each sub-institutional field as the institutional level of the institutional information, and store the institutional level as a label of the institutional information to complete the construction of the knowledge graph;
[0134] Perform standardization processing on each of the sub-institutional fields using the edit distance algorithm.
[0135] In one example, when one or more instructions in the computer storage medium are loaded by the processor 81, the following steps are also performed:
[0136] Sort each of the sub-institutional fields according to the number of each sub-institutional field;
[0137] Obtain the edit distance between each of the sub-institutional fields;
[0138] Perform merging processing on each of the sub-institutional fields with an edit distance less than the distance threshold.
[0139] In one example, when one or more instructions in the computer storage medium are loaded by the processor 81, the following steps are also performed:
[0140] Store the target sub - organization field with the largest quantity among the sub - organization fields whose edit distance is less than the distance threshold as the standardized name of each sub - organization field.
[0141] In one example, when one or more instructions in the computer storage medium are loaded by the processor 81, the following steps are further executed:
[0142] Perform data cleaning on the organization data submitted by the terminal to remove noise information.
[0143] In one example, when one or more instructions in the computer storage medium are loaded by the processor 81, the following steps are further executed:
[0144] Extract the organization information and author information from the organization data through semantic recognition technology;
[0145] Match and correct the author information using a preset short form template for personal names; and
[0146] Identify the preset conjunction and preset nouns in the organization information, split the organization information into multiple fields based on the preset conjunction and the preset nouns, and add preset punctuation marks between the fields.
[0147] In one example, when one or more instructions in the computer storage medium are loaded by the processor 81, the following steps are further executed:
[0148] Match the standard name of each sub - organization field according to the regional category level and sub - organization level corresponding to each sub - organization field to obtain a matching result;
[0149] Perform a correction operation on the organization information according to the matching result to obtain standardized organization information.
[0150] In one example, when one or more instructions in the computer storage medium are loaded by the processor 81, the following steps are further executed:
[0151] If it is recognized that there are sub - organization fields at the same sub - organization level in the organization information, reduce the level of the same sub - organization level by one level as the organization level.
[0152] It should be noted that since the computer program in the computer storage medium realizes the steps in the above - mentioned organization standardization method when executed by the processor, all embodiments or implementation manners of the above - mentioned organization standardization method are applicable to this computer storage medium and can achieve the same or similar beneficial effects.
[0153] The above has introduced the embodiments of the present application in detail. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for institutional standardization, characterized in that The method comprises: A knowledge graph construction request sent by a receiving terminal is included in the construction request, where the organization information used to construct the knowledge graph is the organization name in the organization data; The organization information is divided according to the preset nouns and preset conjunctions in the preset noun library, and the preset nouns are extracted from the organization information as split fields using text recognition technology, and each sub-organization field included in the organization information is obtained. For the remaining fields that cannot correspond to the preset nouns, multiple preset conjunctions are identified. When it is identified that there are at least two independent nouns between two preset conjunctions, the two independent nouns are used as the splitting node for splitting to obtain the sub-organization field; when there is only one noun between the two preset conjunctions, the nouns before and after the preset conjunction are identified to determine the field representing the organization type, and the noun after the preset conjunction and the noun before the preset conjunction are divided into one sub-organization field; Match the acquired sub-institution field with the pre-stored institution name to determine the regional category level corresponding to each sub-institution field; Construct three sub-institutional levels, schools and hospitals are first-level institutions, colleges and branches are second-level institutions, and laboratories, departments, departments, and divisions are third-level institutions. The three sub-institutional levels have a subordinate relationship; Determine the sub-organization level corresponding to each sub-organization field; If it is identified that there is a sub-institution field with the same sub-institution level in the institution information, the same sub-institution level is reduced by one level as the sub-institution level; The lowest level among the sub-organization levels corresponding to each sub-organization field is used as the organization level of the organization information; The regional category level and the sub-institution level are stored as labels of the institution information to complete the construction of the knowledge graph; Sort each of the sub-organization fields according to the number of each sub-organization field in the knowledge graph; Obtaining the edit distance between each of the sub-organization fields, where the edit distance between each of the sub-organization fields is the similarity between each of the sub-organization fields; Merging the sub-organization fields whose edit distance is less than the distance threshold; The target sub-organization field with the largest number among the sub-organization fields whose edit distance is less than the distance threshold is stored as the standardized name of each sub-organization field; Receiving a standardization request sent by a terminal, wherein the standardization request includes information of an organization to be matched or standardized; Determining the institutional level of the institutional information to be matched or standardized; Searching for labels at the same institution level in the knowledge graph to match corresponding standardized institution names; According to the regional category level and sub-organization level corresponding to each sub-organization field of the organization information to be matched or standardized, matching the standard name of each sub-organization field to obtain a matching result; A correction operation is performed on the organization information to be matched or standardized according to the matching result to obtain standardized organization information.
2. The method according to claim 1, characterized in that The method further comprises: Clean the institutional data submitted by the terminal to remove noise information.
3. The method according to claim 2, characterized in that The data cleaning of the institution data submitted by the terminal to remove noise information includes: Extracting the institution information and author information from the institution data by using semantic recognition technology; Using a preset name abbreviation template to match and correct the author information; and Preset conjunctions and preset nouns in the organization information are identified, the organization information is split into multiple fields based on the preset conjunctions and the preset nouns, and preset punctuation marks are added between the fields.
4. A mechanism standardization device, characterized in that: The device comprises a module for executing the method according to any one of claims 1 to 3, and the device comprises: A data acquisition module is used to acquire the sub-institution field in the institution information, identify each sub-institution field in the sub-institution field by using text recognition technology, and determine the regional category level corresponding to each sub-institution field; A level determination module, used to determine the sub-organization level corresponding to each sub-organization field; A graph construction module, used to take the lowest level among the sub-organization levels corresponding to each sub-organization field as the organization level of the organization information, and store the organization level as a label of the organization information to complete the construction of the knowledge graph; The standardization module is used to standardize each of the sub-organization fields using an edit distance algorithm.
5. An electronic device, comprising an input device and an output device, characterized in that: Also includes: a processor adapted to implement one or more instructions; as well as, A computer storage medium storing one or more instructions, wherein the one or more instructions are suitable for being loaded by the processor and executing the steps in the method according to any one of claims 1 to 3.
6. A computer storage medium, characterized in that: The computer storage medium stores one or more instructions, and the one or more instructions are suitable for being loaded by a processor and executing the steps in the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Address comparison method, device and system
CN109739997A