Method, device and medium for identifying a target object in the pharmaceutical industry
Patent Information
- Application Number
- CN202311013409.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2042-09-30
AI Technical Summary
[0004]关于基于纯粹人工的识别方法,其虽然能够识别不规范的医药行业目标对象的原始数据,但是识别效率不高,并且存在因识别主体的经验差异而使得识别结果存在差异性,因此,难以适应大数据量的待识别医药行业目标对象的准确、快速识别,进而难以适应医药行业的服务平台对医药行业目标对象的识别需求
Smart Images

Figure CN117272991B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention application filed on March 3, 2023, with Chinese application number 202211211885.0 and entitled "Method, apparatus and medium for identifying target objects in the pharmaceutical industry". Technical Field
[0002] The embodiments of this disclosure generally relate to the field of data identification, and more specifically to a method, computing device, and computer storage medium for identifying a target object in the pharmaceutical industry to be identified. Background Technology
[0003] Traditional methods for identifying target entities in the pharmaceutical industry (such as, but not limited to, institutions in the pharmaceutical distribution field) typically include: identifying unknown target entities based on purely manual methods; and identifying target entities based on simple word segmentation techniques using natural language processing.
[0004] Regarding purely manual identification methods, while they can identify raw data of non-standard pharmaceutical industry target objects, their identification efficiency is low, and the results vary due to differences in the experience of the identification personnel. Therefore, they are unsuitable for accurately and quickly identifying large volumes of pharmaceutical industry target objects, and consequently, they fail to meet the identification needs of pharmaceutical industry service platforms. As for identification methods based on simple word segmentation technology, given the non-standard representation of raw data of pharmaceutical industry target objects, which often exhibits significant differences in content and structure, coupled with the lack of readily available word segmentation and matching logic in the pharmaceutical industry, the accuracy rate of target object identification is relatively low.
[0005] In summary, the shortcomings of traditional methods for identifying target objects in the pharmaceutical industry are that they are difficult to identify target objects in the pharmaceutical industry quickly and accurately. Summary of the Invention
[0006] To address the aforementioned issues, this disclosure provides a method, computing device, and computer storage medium for identifying target objects in the pharmaceutical industry, enabling rapid and accurate identification of such target objects.
[0007] According to a first aspect of this disclosure, a method for identifying a target object in the pharmaceutical industry is provided, comprising: acquiring raw data to be identified for indicating the target object in the pharmaceutical industry; identifying administrative division information and channel type information in the raw data to be identified; performing noise removal and word segmentation on the raw data to be identified based on the administrative division information, channel type information, and at least one of a noise lexicon, a semantic equivalence lexicon, and a fixed lexicon to generate a word segmentation result, the word segmentation result including multiple keywords; performing hash calculation on the multiple keywords included in the word segmentation result to confirm whether the word segmentation result matches a reference name; and in response to confirming that the word segmentation result does not match the reference name, performing semantic similarity analysis on the reference name and preprocessed data combined based on the word segmentation result to identify the target object in the pharmaceutical industry based on the result of the similarity analysis.
[0008] According to a second aspect of this disclosure, a computing device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of the first aspect of this disclosure.
[0009] In a third aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the method of the first aspect of this disclosure.
[0010] In some embodiments, performing hash calculations on multiple keywords included in the word segmentation result to confirm whether the word segmentation result matches the reference name includes: calculating the sum of the hash values of the multiple keywords included in the word segmentation result to generate a sum of hash values for the word segmentation result; calculating the sum of the hash values of the multiple keywords included in the reference name to generate a sum of hash values for the reference name; confirming whether the sum of hash values for the word segmentation result and the sum of hash values for the reference name are equal; and determining that the word segmentation result matches the reference name in response to confirming that the sum of hash values for the word segmentation result and the sum of hash values for the reference name are equal.
[0011] In some embodiments, the method for identifying a target object in the pharmaceutical industry to be identified further includes: in response to confirming that the word segmentation result matches the reference name, identifying the target object in the pharmaceutical industry to be identified as a target object associated with the reference name.
[0012] In some embodiments, generating word segmentation results includes: obtaining non-administrative division data (excluding administrative division information) from the original data to be identified based on the identified administrative division information; removing noise words and replacing equivalent words for the non-administrative division data; and segmenting the data after noise word removal and equivalent word replacement based on a fixed lexicon to generate word segmentation results corresponding to the original data to be identified. The word segmentation results include multiple keywords and multiple predetermined identifiers indicating segmentation positions.
[0013] In some embodiments, the method for identifying a target object in the pharmaceutical industry further includes: identifying numeric words in the original data to be identified; normalizing the identified numeric words to segment out keywords in the form of numeric words in the original data to be identified; and combining multiple keywords included in the word segmentation results into preprocessed data without geographical information for matching with reference names.
[0014] In some embodiments, normalizing the identified numeric words to segment out keywords in numeric form from the original data to be identified includes: converting uppercase and / or lowercase Chinese numerals in the original data to be identified into Arabic numerals; determining whether the number of digits in the converted Arabic numerals is greater than or equal to a predetermined number of digits threshold; removing the converted Arabic numerals in response to determining that the number of digits in the converted Arabic numerals is greater than or equal to the predetermined number of digits threshold; determining whether the converted Arabic numerals are located at the start or end position of the original data to be identified in response to determining that the converted Arabic numerals are located at the start or end position of the original data to be identified; determining whether the data adjacent to the Arabic numerals at the start or end position indicates a predetermined channel type in response to determining that the data adjacent to the Arabic numerals at the start or end position does not indicate a predetermined channel type; and removing the converted Arabic numerals in response to determining that the data adjacent to the Arabic numerals at the start or end position does not indicate a predetermined channel type.
[0015] In some embodiments, channel type information includes: channel type subcategory name, channel type category name, and channel type category number.
[0016] In some embodiments, identifying administrative division information and channel type information in the original data to be identified includes: determining multiple sets of keywords associated with different priority orders, each set of keywords including multiple predetermined keywords; determining, among the multiple sets of keywords, the target set of predetermined keywords included in the original data to be identified; determining, based on the priority order associated with the target set of keywords, the channel type subclassification name matching the original data to be identified; and determining, based on the determined channel type subclassification name, the channel type classification name and channel type classification number matching the original data to be identified.
[0017] In some embodiments, noise removal and word segmentation of the original data to be identified includes: determining multiple sets of related words, each set of related words including an original word and an equivalent word, the original word and the equivalent word having consistent semantics when indicating a target object in the pharmaceutical industry; determining an association sequence number and a category for each set of related words, the sequence number indicating the priority of each set of related words; and replacing and segmenting the original data to be identified using the equivalent word based on the determined association sequence number, such that the data replaced and segmented by the equivalent word includes the equivalent word and a predetermined identifier, the predetermined identifier indicating the segmentation position.
[0018] In some embodiments, noise removal and word segmentation of the original data to be identified includes: determining the overlapping portion of the preprocessed data and the reference name; deleting the overlapping portion in the preprocessed data to obtain the remaining portion; in response to determining that a first predetermined confidence condition is met, determining the matching confidence level between the original data to be identified and the reference name to be a first level, wherein the matching confidence level of the first level indicates that the original data to be identified and the reference name are matched, and the first predetermined condition includes any one of the following: determining that the number of characters included in the remaining portion is less than or equal to a first character count threshold; determining that the number of characters included in the remaining portion is greater than a second character count threshold and the remaining portion and the overlapping portion are associated with the same channel type information, wherein the second character count threshold is greater than the first character count threshold; the number of characters included in the remaining portion is greater than the first character count threshold and less than the second character count threshold and the remaining portion contains a pair of parentheses; the remaining portion contains "original" or parentheses and "original"; the remaining portion contains a pair of parentheses and the number of characters in the parentheses is less than a third character count threshold, wherein the third character count threshold is greater than the first character count threshold and less than the second character count threshold; determining that there is an overlapping portion between the preprocessed data and the reference name, and that the preprocessed data and the reference name have the same channel type subclassification.
[0019] In some embodiments, semantic similarity analysis of the word segmentation results and the reference name further includes: in response to determining that a second predetermined credibility condition is met, determining the matching credibility level between the original data to be identified and the reference name to be a second level, wherein the second predetermined credibility condition includes: determining that the word segmentation results of the preprocessed data and the reference name have overlapping parts after structural recombination, and that the channel type classification information of the preprocessed data and the reference name is the same; in response to determining that a third predetermined credibility condition is met, determining the mismatch between the original data to be identified and the reference name, wherein the third predetermined credibility condition includes: the word segmentation results of the preprocessed data and the reference name have overlapping parts after structural recombination, and that the channel type classification information of the preprocessed data and the reference name is different.
[0020] In some embodiments, identifying administrative division information and channel type information in the raw data to be identified includes: identifying administrative division information in the name of the organization to be identified based on the full name, abbreviation, former name, and exclusion words of the province, city, district, and county, wherein the administrative division information includes province information, city information, and district / county information; in response to confirming that the identified district / county information or city information does not indicate a unique district / county or city, using the subordinate administrative division information of the identified district / county information or city information, or the administrative division information of the associated target object of the target object to be identified, to identify the administrative division information in the raw data to be identified.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0022] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements.
[0023] Figure 1 A schematic diagram of a system for implementing a method for identifying a target object in the pharmaceutical industry according to an embodiment of the present invention is shown.
[0024] Figure 2 A flowchart of a method for identifying a target object in the pharmaceutical industry according to an embodiment of the present disclosure is shown.
[0025] Figure 3 A flowchart of a method for identifying administrative division information and channel type information in raw data to be identified, according to an embodiment of the present disclosure, is shown.
[0026] Figure 4A flowchart is shown of a method for segmenting keywords in the form of numeric words from raw data to be identified, according to an embodiment of the present disclosure.
[0027] Figure 5 A flowchart is shown for a method for performing semantic similarity analysis on word segmentation results and reference names according to embodiments of the present disclosure.
[0028] Figure 6 A flowchart of a method for generating word segmentation results according to an embodiment of the present disclosure is shown.
[0029] Figure 7 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation
[0030] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0031] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0032] As described earlier, traditional, purely manual identification methods are inefficient and prone to inconsistencies due to differences in the experience of the identification personnel. Therefore, they are ill-suited for accurately and quickly identifying large volumes of pharmaceutical industry target objects, and consequently, cannot meet the identification needs of pharmaceutical industry service platforms. Traditional identification methods based on simple word segmentation lack the word segmentation and pattern logic specific to the pharmaceutical industry, resulting in relatively low accuracy in identifying target objects. Therefore, the shortcomings of traditional methods for identifying pharmaceutical industry target objects lie in their inability to quickly and accurately identify them. For example, traditional methods struggle to quickly and accurately identify "Sanmen Pharmaceutical Co., Ltd." and "China Resources Sanmenxia Pharmaceutical Co., Ltd."
[0033] To at least partially address one or more of the aforementioned problems and other potential issues, exemplary embodiments of this disclosure propose a scheme for identifying target objects in the pharmaceutical industry. In this scheme, administrative division information and channel type information are identified in the acquired raw data used to indicate target objects in the pharmaceutical industry. Based on the identified administrative division information, channel type information, and at least one of a noise lexicon, a semantic equivalence lexicon, and a fixed lexicon, noise removal and word segmentation are performed on the raw data to generate word segmentation results. This disclosure allows the word segmentation results to be standardized by noise removal, semantic equivalence terms, and / or fixed terms, and further aids in the judgment by channel type information. Therefore, it can overcome the problems of structural differences, non-standardized expression, and easy confusion in the original data of target objects in the pharmaceutical industry. Furthermore, this disclosure utilizes hash calculations on multiple keywords included in the word segmentation results to confirm whether the word segmentation results match the reference name; and if it is confirmed that the word segmentation results do not match the reference name, semantic similarity analysis is performed on the preprocessed data generated from the word segmentation results and the reference name, so as to identify the target object in the pharmaceutical industry to be identified based on the results of the similarity analysis. This disclosure can first accurately identify the matching relationship between the word segmentation results and the reference name through hash calculation based on the standardized word segmentation results, and then identify the target object in the pharmaceutical industry to be identified based on the results of semantic similarity analysis when no match can be found. Therefore, this disclosure can identify the target object in the pharmaceutical industry to be identified more quickly and accurately.
[0034] Figure 1 A schematic diagram of a system 100 for implementing a method for identifying a target object in the pharmaceutical industry according to an embodiment of the present invention is shown. Figure 1 As shown, system 100 includes computing device 110, server 130, and network 140. Computing device 110 and server 130 can interact with each other via network 140 (e.g., the Internet).
[0035] Server 130, for example, can send raw data to be identified for indicating target objects in the pharmaceutical industry to computing device 110.
[0036] Regarding computing device 110, it is used, for example, to acquire raw data to be identified provided by server 130 for indicating target objects in the pharmaceutical industry; and to identify administrative division information and channel type information in the raw data to be identified. Computing device 110 can also perform noise removal and word segmentation on the raw data to be identified based on administrative division information, channel type information, and at least one of a noise lexicon, a semantic equivalence lexicon, and a fixed lexicon, to generate word segmentation results; perform hash calculations on multiple keywords included in the word segmentation results to confirm whether the word segmentation results match a reference name; and if it is confirmed that the word segmentation results do not match the reference name, perform semantic similarity analysis on the word segmentation results and the reference name to identify the target object in the pharmaceutical industry to be identified based on the results of the similarity analysis. Computing device 110 may have one or more processing units, including dedicated processing units such as GPUs, FPGAs, and ASICs, and general-purpose processing units such as CPUs. Additionally, one or more virtual machines may run on each computing device 110. In some embodiments, computing device 110 and medical imaging device 110 may be integrated together or set up separately. In some embodiments, the computing device 110 includes, for example, a raw data acquisition unit 112 to be identified, an administrative division and channel type information identification unit 114, a word segmentation result generation unit 116, a hash calculation unit 118, and a pharmaceutical industry target object identification unit 120 to be identified.
[0037] Regarding the raw data acquisition unit 112, it is used to acquire raw data to be identified for indicating target objects in the pharmaceutical industry.
[0038] Regarding the administrative division and channel type information identification unit 114, it is used to identify the administrative division information and channel type information in the original data to be identified.
[0039] Regarding the word segmentation result generation unit 116, it is used to perform noise removal and word segmentation on the original data to be identified based on administrative division information, channel type information, and at least one of the following lexicons: noise lexicon, semantic equivalence lexicon, and fixed lexicon, so as to generate word segmentation results, which include multiple keywords.
[0040] Regarding the hash calculation unit 118, it performs hash calculations based on multiple keywords included in the word segmentation results in order to confirm whether the word segmentation results match the reference name.
[0041] Regarding the target object identification unit 120 in the pharmaceutical industry to be identified, if it is confirmed that the word segmentation result does not match the reference name, it performs semantic similarity analysis on the reference name and the preprocessed data combined based on the word segmentation result, so as to identify the target object in the pharmaceutical industry to be identified based on the result of the similarity analysis.
[0042] The following combination Figure 2 A method for identifying target objects in the pharmaceutical industry is described 200. Figure 2 A flowchart of a method 200 for identifying a target object in the pharmaceutical industry according to an embodiment of the present disclosure is shown. Method 200 may be derived from, for example... Figure 1 The computing device 110 shown can be used for execution, and can also be used in Figure 7 The method is performed at the illustrated electronic device 700. It should be understood that method 200 may also include additional boxes not shown and / or the boxes shown may be omitted, and the scope of this disclosure is not limited in this respect.
[0043] In step 202, computing device 110 acquires raw data to be identified for indicating target objects in the pharmaceutical industry. For example, computing device 110 acquires raw data to be identified from server 130 regarding unknown entities in the pharmaceutical distribution field.
[0044] Regarding the target pharmaceutical industry entity to be identified, this includes, but is not limited to, unknown entities in the pharmaceutical distribution sector. For example, computing device 110 needs to identify which standard entity name represents an unknown company or entity name in a pharmaceutical distribution sector. It should be understood that the same target pharmaceutical industry entity (e.g., but not limited to, the same pharmacy) may have supply relationships with different pharmaceutical entities (e.g., distributors), and this target pharmaceutical industry entity may have inconsistent names or designations at different pharmaceutical entities (e.g., distributors).
[0045] In step 204, the computing device 110 identifies the administrative division information and channel type information in the raw data to be identified.
[0046] Information regarding administrative divisions includes, for example, information about the administrative agencies at the provincial, city, and county levels.
[0047] Methods for identifying administrative division information in raw data to be identified include, for example, computing device 110 identifying administrative division information in the name of an organization to be identified based on the full name, abbreviation, former name, and exclusion words of provinces, cities, districts, and counties, whereby administrative division information includes province information, city information, and district / county information; if it is confirmed that the identified district / county information or city information does not indicate a unique district / county or city, the administrative division information in the raw data to be identified is identified using the subordinate administrative division information of the identified district / county information or city information, or the administrative division information of the associated target object of the target object to be identified. Specifically, if the computing device 110 determines that the province information contained in the original data to be identified includes the full name, abbreviation, and capital city or former name of the capital city, and does not include excluded names related to the province or capital city, then the province information is identified; if the city information contained in the original data to be identified includes the full name, abbreviation, or former name of the city, and does not include excluded names of the city, then the city information is identified; and if the district / county information contained in the original data to be identified includes the full name, abbreviation, or former name of the district / county, and does not include excluded names of the district / county, then the district / county information is identified; if the computing device 110 determines that any of the following conditions are met, then the administrative division information is identified: confirming the identification of province information, city information, and district / county information; confirming the identification of administrative division information and district / county information, confirming the identification of city information and district / county information; confirming that the identified district / county information or city information indicates a unique district / county or city.
[0048] For example, if the computing device 110 determines that the name of the organization to be identified contains administrative agencies at the provincial, municipal, and county levels, or at the provincial and county levels (e.g., province + county / second-level city / district), or at the municipal and county levels (e.g., prefecture-level city + county level), then the administrative division information to which the name of the organization to be identified belongs can be directly identified without further detection, that is, it is considered that the province, municipality, and county to which the name of the organization to be identified belongs has been accurately found.
[0049] For example, if computing device 110 determines that the name of the organization to be identified includes the full name of a district / county or a city, and that the full name of the district / county or city is unique, then the administrative division information to which the name of the organization to be identified belongs is determined. It should be understood that cities and counties nationwide are unique. If the name of the organization to be identified contains a unique full name, abbreviation, or former name of a city / county, then it is considered that the province, city, or county to which the name of the organization to be identified belongs can be uniquely identified.
[0050] If the computing device 110 confirms that the identified district / county or city information does not indicate a unique district / county or city, it uses the lower-level administrative division information of the identified district / county or city information, or the administrative division information of the associated target object, to identify the administrative division information in the original data to be identified. For example, "Guoyuan Village Health Clinic, Yongshun Town, Tongzhou District" and "Traditional Chinese Medicine Pharmacy, Jinsha Town, Tongzhou District," where Tongzhou District does not indicate a unique district / county. For example, Beijing includes Tongzhou, and Jiangsu Province also includes Tongzhou. Therefore, the administrative division information in the name of the institution to be identified can be identified by using lower-level administrative division information (e.g., the relationship between townships and districts / counties). For example, by using "Tongzhou" + "Yongshun," a unique administrative division relationship "Beijing + Tongzhou + Yongshun" can be found, which will then locate Tongzhou District of Beijing; similarly, by using "Tongzhou" + "Jinsha," a unique administrative division relationship "Jiangsu Province + Nantong City + Tongzhou District" can be found.
[0051] For example, as shown in Table 1 below, the name of the target object to be identified (e.g., the buyer's organization) is "Beigou Health Center". However, it is impossible to find the geographical information or administrative division information of its province, city, district, or county from "Beigou Health Center" alone. The computing device 110 can identify that the associated target object (e.g., the seller organization "China Resources Yantai Pharmaceutical Co., Ltd.") is located in Yantai, Shandong. The computing device 110 can then search for downstream towns in the Yantai area to see if "Beigou" exists, ultimately finding a unique town called Beigou within "Penglai District".
[0052] Table 1
[0053]
[0054] In some embodiments, the computing device 110 identifies administrative division information in the name of an organization to be identified based on the full name, abbreviation, former name, and excluded words of the province, city, and district / county. The administrative division information includes province information, city information, and district / county information. Table 2 below exemplarily shows the full name, abbreviation, former name, and excluded words of the city and district / county. The full name, abbreviation, former name, and excluded words of the province are not shown in Table 2.
[0055] Table 2
[0056]
[0057] For example, "Sanmen Pharmaceutical Co., Ltd.", "China Resources Sanmenxia Pharmaceutical Co., Ltd.", and "Jiajiang County Pharmaceutical Company Sanmen City" contain some easily confused abbreviations of districts and counties. The computing device 110 can assist in identifying administrative division information in other organization names based on exclusion words related to provinces, cities, and districts / counties. For example, Sanmen County, whose abbreviation is Sanmen, excludes words such as: Sanmenxia, Sanmen City, Third Gate, and Sanmen Store. By employing the above methods, this invention can accurately identify easily confused administrative division information, thereby improving the accuracy of identifying target objects.
[0058] Regarding channel type information, this includes, for example, the channel type sub-category name, the channel type sub-category name, and the channel type category number. It should be understood that pharmaceutical distribution industry data is divided into three main categories: distributors, medical terminals, and retail terminals. Each category has further subcategories; for example, retail terminals are divided into independent pharmacies and chain pharmacy branches. The institution names typically contain channel type information, which helps improve the accuracy of identifying target pharmaceutical industry entities. For instance, retail terminals cannot identify medical terminals. Therefore, identifying channel type information in the raw data helps improve the accuracy of identifying target pharmaceutical industry entities.
[0059] A method for identifying administrative division information and channel type information in raw data to be identified includes, for example: a computing device 110 determines multiple sets of keywords associated with different priority orders, each set of keywords including multiple predetermined keywords; among the multiple sets of keywords, a target set of keywords containing the predetermined keywords included in the raw data to be identified is determined; based on the priority order associated with the target set of keywords, a channel type subclassification name matching the raw data to be identified is determined; and based on the determined channel type subclassification name, a channel type classification name and a channel type classification number matching the raw data to be identified are determined.
[0060] In step 206, the computing device 110 performs noise removal and word segmentation on the original data to be identified based on administrative division information, channel type information, and at least one of the following word libraries: noise library, semantic equivalence library, and fixed library, in order to generate word segmentation results, which include multiple keywords.
[0061] The method for noise removal and word segmentation of the original data to be identified includes, for example, confirming whether the preprocessed data after noise removal and normalization matches at least one of the full name, alias, and former name of the reference name; if it is confirmed that the preprocessed data after noise removal and normalization does not match the full name, alias, and former name of the reference name, word segmentation is performed on the preprocessed data to generate word segmentation results. If the preprocessed data is equal to an alias or former name of the reference name, or the preprocessed data plus its upstream name equals the reference name or its alias, or the preprocessed data plus its upstream name is a homophone of the reference name or its alias, then the computing device 110 determines that the original data to be identified matches the reference name, and word segmentation is not required on the preprocessed data.
[0062] Methods for generating word segmentation results include, for example: obtaining non-administrative division data (excluding administrative division information) from the original data to be identified based on the identified administrative division information; removing noise words and replacing equivalent words for the non-administrative division data; segmenting the data after noise word removal and equivalent word replacement based on a fixed lexicon to generate word segmentation results corresponding to the original data to be identified, the word segmentation results including multiple keywords and multiple predetermined identifiers indicating segmentation positions; identifying numeric words in the original data to be identified; normalizing the identified numeric words to segment out keywords in numeric form from the original data to be identified; and combining the multiple keywords included in the word segmentation results into preprocessed data without geographical information for matching with reference names. The following will combine... Figure 6 The methods for semantic similarity analysis of word segmentation results and reference names are explained in detail here, but will not be repeated here.
[0063] A method for segmenting keywords in the form of numeric words from raw data to be identified includes, for example, the following steps: The computing device 110 converts uppercase and / or lowercase Chinese numerals in the raw data to be identified into Arabic numerals; determines whether the number of digits in the converted Arabic numerals is greater than or equal to a predetermined digit threshold; in response to determining that the number of digits in the converted Arabic numerals is greater than or equal to the predetermined digit threshold, removes the converted Arabic numerals; in response to determining that the number of digits in the converted Arabic numerals is less than the predetermined digit threshold, determines whether the converted Arabic numerals are located at the start or end position of the raw data to be identified; in response to determining that the converted Arabic numerals are located at the start or end position of the raw data to be identified, determines whether the data adjacent to the Arabic numerals at the start or end position indicates a predetermined channel type; and in response to determining that the data adjacent to the Arabic numerals at the start or end position does not indicate a predetermined channel type, removes the converted Arabic numerals. The following will combine... Figure 4The method for segmenting keywords in the form of numeric words from the original data to be identified is detailed elsewhere and will not be repeated here. Regarding the method for noise removal and equivalent word replacement for non-administrative division data, it includes, for example, the following: a computing device 110 determines multiple sets of related words, each set including an original word and an equivalent word, the original word and the equivalent word having consistent semantics when indicating a target object in the pharmaceutical industry; determines an association sequence number and category for each set of related words, the sequence number indicating the priority of each set of related words; and based on the determined association sequence number, replaces and segments the original data to be identified using equivalent words, such that the data replaced and segmented by equivalent words includes the equivalent word and a predetermined identifier, the predetermined identifier indicating the segmentation position.
[0064] In step 208, the computing device 110 performs hash calculations on the multiple keywords included in the word segmentation results in order to confirm whether the word segmentation results match the reference name.
[0065] The method for confirming whether the word segmentation result matches the reference name includes, for example: the computing device 110 calculates the sum of the hash values of multiple keywords included in the word segmentation result to generate a sum of hash values for the word segmentation result; calculates the sum of the hash values of multiple keywords included in the reference name to generate a sum of hash values for the reference name; confirms whether the sum of hash values for the word segmentation result and the sum of hash values for the reference name are equal; and in response to confirming that the sum of hash values for the word segmentation result and the sum of hash values for the reference name are equal, determines that the word segmentation result matches the reference name. By using the algorithm logic of the sum of hash values of the keywords after word segmentation, the matching result can be ensured that it is not affected by different keyword word orders.
[0066] The following formula (1) schematically illustrates the algorithm for confirming whether the word segmentation result matches the reference name.
[0067]
[0068] In the above formula (1), ora_hash(key) reference i) represents the hash value calculated for the i-th keyword included in the word segmentation results of the reference data. i represents the keyword index. This represents the sum of the hash values of the reference names. 'n' represents the total number of keywords; for example, in Table 3 or Table 4, the total number of keywords, n, is 19. ora_hash(key) original i) represents the hash value calculated for the i-th keyword included in the word segmentation results of the original data to be identified. This represents the sum of the hash values of the word segmentation results.
[0069] For example, Table 3 below illustrates the word segmentation results of a reference name. The reference name, for example, is "Fuyang Yansheng Pharmacy Retail Chain Co., Ltd. Menglian Branch," and is segmented into nineteen keywords, from keyword 1 to keyword 19 in Table 3. Only nine of these keywords are shown in Table 3.
[0070] Table 3
[0071]
[0072] For example, Table 4 below illustrates the word segmentation results of the original data to be identified. The original data to be identified is, for example, "Fuyang Yansheng Pharmacy Retail Chain Company (Menglian)", and is segmented into nineteen keywords, from keyword 1 to keyword 19, as shown in Table 4. Only nine of these keywords are illustrated in Table 4.
[0073] Table 4
[0074]
[0075] The original data to be identified, “Fuyang Yansheng Pharmacy Retail Chain Company (Menglian)”, was processed through noise reduction and word segmentation, breaking the word order and normalizing the capitalization, generating nineteen keywords from keyword 1 to keyword 19 in Table 4. The sum of the hash values of all keywords from keyword 1 to keyword 19 in the word segmentation results of the original data to be identified (i.e., the sum of the hash values of the word segmentation results) equals the sum of the reference names (the reference name refers to the standard target object name). Thus, the computing device 110 determines that the word segmentation result matches the reference name.
[0076] The following example shows illustrative program code for implementing an algorithm to confirm whether the word segmentation result matches the reference name.
[0077] select*
[0078] from(select a.collatejobdetailid,a.orgname,o.ovalmasterid asstdorgid,o.orgcode as stdorgcode,
[0079] o.orgname as stdorgname,2as status,2as gradelevel,length(o.orgname)asorglen,
[0080] case when a.channelname = o.channel and substr(a.keyword05,-1) in ('-','Store','Pharmacy','Life','Clinic','Hospital') then '99%'
[0081] when a.channelname = o.channel then '98%' else '95%' end as grade, 'Word-segmented full matching recommendation' as splitstatus_std
[0082] from collatejobdetail a, ovalmaster o
[0083] where a.jobid = v_jobid......
[0084] and a.keyword01 = o.keyword01
[0085] and a.keyword02 = o.keyword02
[0086] and a.keyword03 = o.keyword03
[0087] and a.hashvalue = o.hashvalue
[0088] / *hashvalue
[0089] ora_hash(a.keyword04)+ora_hash(a.keyword05)+
[0090] ora_hash(a.keyword06)+ora_hash(a.keyword07)+
[0091] ora_hash(a.keyword08)+ora_hash(a.keyword09)+
[0092] ora_hash(a.keyword10)+ora_hash(a.keyword11)+
[0093] ora_hash(a.keyword12)+ora_hash(a.keyword19) =
[0094] ora_hash(o.keyword04)+ora_hash(o.keyword05)+
[0095] ora_hash(o.keyword06)+ora_hash(o.keyword07)+
[0096] ora_hash(o.keyword08)+ora_hash(o.keyword09)+
[0097] ora_hash(o.keyword10)+ora_hash(o.keyword11)+
[0098] ora_hash(o.keyword12)+ora_hash(o.keyword19)* / )
[0100] order by orgname,orglen
[0101] In step 210, if the computing device 110 confirms that the word segmentation result does not match the reference name, it performs semantic similarity analysis on the reference name and the preprocessed data combined based on the word segmentation result, so as to identify the target object in the pharmaceutical industry to be identified based on the result of the similarity analysis. For example, if the computing device 110 confirms that the word segmentation result matches the reference name, it identifies the target object in the pharmaceutical industry to be identified as the target object associated with the reference name.
[0102] A method for semantic similarity analysis of word segmentation results and reference names includes, for example: identifying overlapping portions of preprocessed data and reference names; deleting overlapping portions in the preprocessed data to obtain remaining portions; in response to determining that a first predetermined confidence level condition is met, determining a first level of matching confidence between the original data to be identified and the reference name, wherein a first level of matching confidence indicates that the original data to be identified and the reference name are matched, and the first predetermined condition includes any of the following: determining that the number of characters included in the remaining portion is less than or equal to a first character count threshold; determining that the number of characters included in the remaining portion is greater than a second character count threshold and the remaining portion and the overlapping portion are associated with the same channel type information, wherein the second character count threshold is greater than the first character count threshold; the number of characters included in the remaining portion is greater than the first character count threshold and less than the second character count threshold and the remaining portion contains a pair of parentheses; the remaining portion contains "original" or parentheses and "original"; the remaining portion... The part contains a pair of parentheses and the number of characters in the parentheses is less than the third character count threshold, the third character count threshold is greater than the first character count threshold and less than the second character count threshold; in response to meeting the second predetermined confidence condition, the matching confidence level between the original data to be identified and the reference name is determined to be the second level. The second predetermined confidence condition includes any one of the following: determining that the preprocessed data and the reference name have overlapping parts, and the preprocessed data and the reference name have the same channel type subclassification; determining that the word segmentation results of the preprocessed data and the reference name have overlapping parts after structural recombination, and the channel type classification information of the preprocessed data and the reference name is the same; in response to meeting the third predetermined confidence condition, the mismatch between the original data to be identified and the reference name is determined. The third predetermined confidence condition includes: the word segmentation results of the preprocessed data and the reference name have overlapping parts after structural recombination, and the channel type classification information of the preprocessed data and the reference name is different. The following will combine Figure 5 The methods for semantic similarity analysis of word segmentation results and reference names are explained in detail here, but will not be repeated here.
[0103] In some embodiments, if the semantic similarity analysis of the word segmentation results and the reference name or the hash calculation of the word segmentation results cannot accurately identify the target object in the pharmaceutical industry to be identified, the computing device 110 can adjust the weight of the semantic similarity analysis of the word segmentation results and the reference name based on the channel type information, so as to perform semantic similarity analysis on the word segmentation results and the reference name based on the adjusted weight.
[0104] In the above scheme, by identifying administrative division information and channel type information of the acquired original data to be identified, which is used to indicate the target objects in the pharmaceutical industry, noise removal and word segmentation are performed on the original data to be identified based on the identified administrative division information, channel type information, and at least one of the following word libraries: noise library, semantic equivalence library, and fixed word library, so as to generate word segmentation results. This disclosure can make the word segmentation results the result after noise removal, standardization by semantic equivalence words and / or fixed words, and assist in the judgment by channel type information. Therefore, it can overcome the problems of differences in the original data structure, non-standard expression, and easy confusion of the target objects in the pharmaceutical industry. Furthermore, this disclosure utilizes hash calculations on multiple keywords included in the word segmentation results to confirm whether the word segmentation results match the reference name; and if it is confirmed that the word segmentation results do not match the reference name, semantic similarity analysis is performed on the preprocessed data generated from the word segmentation results and the reference name, so as to identify the target object in the pharmaceutical industry to be identified based on the results of the similarity analysis. This disclosure can first accurately identify the matching relationship between the word segmentation results and the reference name through hash calculation based on the standardized word segmentation results, and then identify the target object in the pharmaceutical industry to be identified based on the results of semantic similarity analysis when no match can be found. Therefore, this disclosure can identify the target object in the pharmaceutical industry to be identified more quickly and accurately.
[0105] The following combination Figure 3 This describes the method used to identify administrative division information and channel type information in the raw data to be identified. Figure 3 A flowchart of a method 300 for identifying administrative division information and channel type information in raw data to be identified, according to an embodiment of the present disclosure, is shown. Method 300 may be derived from, for example... Figure 1 The computing device 110 shown can be used for execution, and can also be used in Figure 7 The method is performed at the illustrated electronic device 700. It should be understood that method 300 may also include additional boxes not shown and / or the boxes shown may be omitted, and the scope of this disclosure is not limited in this respect.
[0106] In step 302, computing device 110 determines multiple sets of keywords associated with different priority orders, each set of keywords including multiple predetermined keywords.
[0107] The keyword sets used to identify channel type classifications include, for example, keyword sets for identifying chain pharmacies, keyword sets for identifying independent pharmacies, keyword sets for identifying chain companies, keyword sets for identifying hospitals, and keyword sets for identifying health supervision bureaus.
[0108] Table 3 below provides examples of the keyword sets used to identify chain pharmacies and the keyword sets used to identify individual pharmacies.
[0109] Table 5
[0110]
[0111]
[0112] In step 304, the computing device 110 determines, from multiple keyword sets, the target keyword set containing the predetermined keywords included in the original data to be identified. For example, if the computing device 110 determines that the original data to be identified includes the predetermined keyword "%retail center%" but does not include "%chain%store%", then the target keyword set containing the included predetermined keywords is the keyword set in the second row of Table 5.
[0113] In step 306, the computing device 110 determines the channel type sub-category name that matches the original data to be identified based on the priority order associated with the target keyword set. For example, if the priority order associated with the keyword set in the second row of Table 5 is 18, the computing device 110 determines the channel type sub-category name that matches the original data to be identified as "individual pharmacies" based on this priority order 18. It should be understood that when identifying channel types, based on the priority order associated with the target keyword set, those that meet the conditions first are considered to have "located the channel type sub-category name that matches the original data to be identified".
[0114] In step 308, the computing device 110 determines the channel type classification name and channel type classification number that match the original data to be identified, based on the determined channel type subclass name. For example, based on the determined channel type subclass name "individual pharmacy", the computing device 110 determines the channel type classification name that matches the original data to be identified as "terminal pharmacy" and the channel type classification number as "114".
[0115] In the above scheme, this disclosure can accurately determine the channel type to which the original data to be identified belongs, which is conducive to improving the accuracy of identifying target objects in the pharmaceutical industry based on the accurate channel type.
[0116] The following combination Figure 4 This describes a method for segmenting keywords in the form of numeric words from the raw data to be identified. Figure 4 A flowchart of a method 400 for segmenting keywords in the form of numeric words from raw data to be identified, according to an embodiment of the present disclosure, is shown. Method 400 may be derived from, for example... Figure 1 The computing device 110 shown can be used for execution, and can also be used in Figure 7 The method is performed at the illustrated electronic device 700. It should be understood that method 400 may also include additional boxes not shown and / or the boxes shown may be omitted, and the scope of this disclosure is not limited in this respect.
[0117] In step 402, the computing device 110 converts uppercase Chinese numerals and / or lowercase Chinese numerals in the original data to be identified into Arabic numerals. It should be understood that the original data to be identified may contain uppercase and lowercase numerals, telephone numbers, and postal codes. These numerals may appear before, after, or in the middle of the name of the target object to be identified. For example, the computing device 110 can uniformly convert all uppercase and lowercase Chinese numerals appearing in the original data to be identified into Arabic numerals. For example, "一百五十一" (one hundred and fifty-one), "一五一" (one five one), or "壹佰五十一" (one hundred and fifty-one in formal Chinese numerals) are all finally converted into the Arabic numeral "151". By using the above method, it is beneficial to normalize the numeric words in the original data to be identified.
[0118] In step 404, the computing device 110 determines whether the number of digits of the converted Arabic numeral is greater than or equal to a predetermined digit threshold.
[0119] For the predetermined digit threshold, it is, for example but not limited to, a number of 6 or above.
[0120] If the computing device 110 determines that the number of digits of the converted Arabic numeral is greater than or equal to the predetermined digit threshold, in step 406, the converted Arabic numeral is removed. For example, if it is determined that the number of digits of the converted Arabic numeral is greater than or equal to 6 (or 6 digits or more), the converted Arabic numeral is directly removed regardless of where the converted Arabic numeral appears. This is because the converted Arabic numeral may be information such as a telephone number or a postal code.
[0121] In step 408, if the computing device 110 determines that the number of digits of the converted Arabic numeral is less than the predetermined digit threshold, it determines whether the converted Arabic numeral is located at the start position or the end position of the original data to be identified.
[0122] In step 410, if the computing device 110 determines that the converted Arabic numeral is located at the start position or the end position of the original data to be identified, it determines whether the data adjacent to the Arabic numeral located at the start position or the end position indicates a predetermined channel type. For example, if the computing device 110 determines that the converted Arabic numeral is less than the predetermined digit threshold and appears at the start position of the name of the target object to be identified, the converted Arabic numeral can also be removed, because it is very likely that the converted Arabic numeral is a serial number accidentally added when providing the name of the target object.
[0123] If the computing device 110 determines that the data adjacent to the Arabic numeral located at the start position or the end position does not indicate a predetermined channel type, the process jumps to step 406 to remove the converted Arabic numeral.
[0124] In step 412, if the computing device 110 determines that the data adjacent to the Arabic numerals at the start or end position indicates a predetermined channel type, the converted Arabic numerals are not removed.
[0125] For example, if the computing device 110 determines that the converted Arabic numeral is located at the start or end position of the original data to be identified, and the type immediately following the converted Arabic numeral is not a pharmacy type or a medical institution type (the channel type is not a pharmacy), then the Arabic numeral can be removed. If the computing device 110 determines that the Arabic numeral at the end position appears at the end position, and the type immediately following the converted Arabic numeral is a pharmacy type or a medical institution type, and the converted Arabic numeral is less than or equal to a predetermined number threshold, then the Arabic numeral cannot be removed. For example, in the example "56 store Wang Zhiheng" in Table 6 below, the Arabic numeral "56" is located at the start position of the original data to be identified, and the type immediately following the converted Arabic numeral "56" is a pharmacy type. At the same time, the Arabic numeral "56" is less than the predetermined number threshold associated with a municipal pharmaceutical company. In this case, the computing device 110 determines that the Arabic numeral "56" cannot be removed.
[0126] For example, in the example of "56 Store Wang Zhiheng" in Table 6 below, the Arabic numeral "56" is located at the beginning of the original data to be identified, and the pharmacy type is immediately following the converted Arabic numeral "56". At the same time, the Arabic numeral "56" is less than the predetermined numerical threshold associated with the municipal pharmaceutical company. In this case, the computing device 110 determines that the Arabic numeral "56" cannot be removed.
[0127] For example, in Table 6 below, the examples "Qihe Township Wanglou Health Clinic 50" or "Qihe Township Wanglou Health Clinic 1" show that the Arabic numeral "50" or "1" is located at the end of the original data to be identified, and the Arabic numeral "50" or "1" following the converted Arabic numeral "50" is the type of medical institution. Assuming that the Arabic numeral "50" is greater than the predetermined numerical threshold associated with the township health clinic, the computing device 110 determines to remove the Arabic numeral "50". However, the Arabic numeral "1" is less than the predetermined numerical threshold associated with the township health clinic, so the computing device 110 determines that the Arabic numeral "1" cannot be removed.
[0128] Table 6
[0129]
[0130] In the above scheme, this disclosure can accurately identify and remove numerical noise in the original data to be identified, which is conducive to accurately segmenting numerical keywords that are conducive to identifying the target object.
[0131] The following combination Figure 5This describes the method used for semantic similarity analysis of word segmentation results and reference names. Figure 5 A flowchart of a method 500 for semantic similarity analysis of word segmentation results and reference names according to embodiments of the present disclosure is shown. Method 500 may be derived from, for example... Figure 1 The computing device 110 shown can be used for execution, and can also be used in Figure 7 The method is performed at the illustrated electronic device 700. It should be understood that method 500 may also include additional boxes not shown and / or the boxes shown may be omitted, and the scope of this disclosure is not limited in this respect.
[0132] In step 502, computing device 110 determines the overlapping portion of preprocessed data and reference names.
[0133] In step 504, computing device 110 removes overlapping portions from the preprocessed data to obtain the remaining portions.
[0134] For example, if the preprocessed data consists of "Tian Zhidong Clinic 1, Xuanhua District" and the reference name "Tian Zhidong Clinic, Xuanhua District", the overlapping part between the preprocessed data and the reference name is "Tian Zhidong Clinic, Xuanhua District". After deleting the overlapping part from the preprocessed data, the remaining part is "1".
[0135] In step 506, if the computing device 110 determines that a first predetermined confidence condition is met, it determines the matching confidence level between the original data to be identified and the reference name to be a first level. The matching confidence level of the first level indicates that the original data to be identified and the reference name are matched. The first predetermined condition includes any one of the following: determining that the number of characters included in the remaining part is less than or equal to a first character count threshold; determining that the number of characters included in the remaining part is greater than a second character count threshold and the remaining part and the overlapping part are associated with the same channel type information, the second character count threshold being greater than the first character count threshold; the number of characters included in the remaining part is greater than the first character count threshold and less than the second character count threshold and the remaining part contains a pair of parentheses; the remaining part contains "original" or parentheses and "original"; or the remaining part contains a pair of parentheses and the number of characters in the parentheses is less than a third character count threshold, the third character count threshold being greater than the first character count threshold and less than the second character count threshold; determining that there is an overlapping part between the preprocessed data and the reference name, and that the preprocessed data and the reference name have the same channel type subclassification.
[0136] Regarding the first word count threshold, it is, for example but not limited to, 2. For example, the number of words included in the above remaining part "1" is less than the first word count threshold, and it is determined that the matching credibility level between the original data to be recognized and the reference name is the first level. For example, the matching similarity is 100%, that is, the original data to be recognized matches the reference name. Regarding the second word count threshold, it is, for example but not limited to, 10. For example, the overlapping part between the preprocessed data "Qitaihe Yuanfu Pharmacy (Qitaihe Yuanhongfu Medical Device Store)" and the reference name "Qitaihe Yuanfu Pharmacy" is "Qitaihe Yuanfu Pharmacy". The remaining part after deleting the overlapping part from the preprocessed data is "(Qitaihe Yuanhongfu Medical Device Store)". If the number of words included in the remaining part is greater than 10, and the channel type information of the remaining part is the same as that of the overlapping part, then it is determined that the matching credibility level between the original data to be recognized and the reference name is the first level. For example, the matching similarity is 98%, that is, the original data to be recognized is highly similar to the reference name, and therefore matches.
[0137] For example, the overlapping part between the preprocessed data "Health Room of Kangbei Village, Kangcun Town, Huojia County (Formerly Kangbei United Health Room)" and the reference name "Health Room of Kangbei Village, Kangcun Town" is "Health Room of Kangbei Village, Kangcun Town". The remaining part after deleting the overlapping part from the preprocessed data is "Huojia County (Formerly Kangbei United Health Room)". If the remaining part contains "former" or "(former", it is determined that the matching credibility level between the original data to be recognized and the reference name is the first level. For example, the matching similarity is 99%, that is, the original data to be recognized is highly similar to the reference name, and therefore matches.
[0138] Regarding the third word count threshold, it is, for example but not limited to, 4. For example, the overlapping part between the preprocessed data "Xianghongtang Pharmaceutical Retail Store, Sanxiang Town, Zhongshan City (06)" and the reference name "Xianghongtang Pharmaceutical Retail Store, Sanxiang Town, Zhongshan City" is "Xianghongtang Pharmaceutical Retail Store, Sanxiang Town, Zhongshan City". If the remaining part after deleting the overlapping part from the preprocessed data contains a pair of parentheses, and the content length in the parentheses is less than 4 characters, it is determined that the matching credibility level between the original data to be recognized and the reference name is the first level. For example, the matching similarity is 96%, that is, the original data to be recognized is highly similar to the reference name, and therefore matches.
[0139] For example, the overlapping part between the preprocessed data "Xiamen Huli Dingling Doctor First Outpatient Department Co., Ltd." and the reference name "Xiamen Huli Dingling Doctor First Outpatient Department" is "Xiamen Huli Dingling Doctor First Outpatient Department". If the two have a completely overlapping part, and they have the same channel subclassification name, it is determined that the matching credibility level between the original data to be recognized and the reference name is the first level. For example, the matching similarity is 70%, that is, the original data to be recognized has high similarity with the reference name, and therefore matches.
[0140] The matching credibility level is Level 1, which, for example, indicates that the matching similarity is between 70% and 100%.
[0141] In step 508, if the computing device 110 determines that the second predetermined credibility condition is met, the matching credibility level between the original data to be identified and the reference name is determined to be the second level. The second predetermined credibility condition includes: determining that the word segmentation results of the preprocessed data and the reference name have overlapping parts after structural recombination, and that the channel type classification information of the preprocessed data and the reference name are the same.
[0142] For example, there is overlap between the preprocessed data "Huainan City Panji District Luji Town Health Center (Chengbei Village Health Clinic)" and the reference name "Luji Town Chengbei Village Health Clinic". The word segmentation result of the preprocessed data "Huainan City Panji District Luji Town Health Center (Chengbei Village Health Clinic)" is restructured into "Luji Health (Chengbei Health)". The restructured word segmentation result "(Chengbei Health)" is contained in the reference name "Luji Chengbei Health". Furthermore, the channel type classification information of the preprocessed data and the reference name is the same. Therefore, the matching confidence level between the original data to be identified and the reference name is determined to be level two, for example, a matching similarity of 65%, meaning that there is a certain degree of similarity between the original data to be identified and the reference name.
[0143] In some embodiments, a matching confidence level of 2 can be considered as a match between the preprocessed data and the reference name.
[0144] In step 510, if the computing device 110 determines that the third predetermined credibility condition is met, it determines the mismatch between the original data to be identified and the reference name. The third predetermined credibility condition includes: the word segmentation results of the preprocessed data and the reference name have overlapping parts after structural recombination, and the channel type classification information of the preprocessed data and the reference name are different.
[0145] For example, there is overlap between the preprocessed data "China Resources Qingdao Pharmaceutical Co., Ltd. Laoshan Road Branch" and the reference name "China Resources Qingdao Pharmaceutical Co., Ltd." However, the channel type information associated with the preprocessed data "China Resources Qingdao Pharmaceutical Co., Ltd. Laoshan Road Branch" and the reference name "China Resources Qingdao Pharmaceutical Co., Ltd." is different; the reference name "China Resources Qingdao Pharmaceutical Co., Ltd." is a commercial company, while the preprocessed data "China Resources Qingdao Pharmaceutical Co., Ltd. Laoshan Road Branch" is a pharmacy. Therefore, they are not similar. This indicates a mismatch between the original data to be identified and the reference name.
[0146] By employing the above methods, this disclosure is able to quickly and accurately identify whether the original data to be identified and the reference name match, even when there are differences between the preprocessed data and the reference name.
[0147] The following combination Figure 6 This describes the method used to generate word segmentation results. Figure 6 A flowchart of a method 600 for generating word segmentation results according to an embodiment of the present disclosure is shown. Method 600 may be derived from, for example... Figure 1 The computing device 110 shown can be used for execution, and can also be used in Figure 7 The method is performed at the illustrated electronic device 700. It should be understood that method 600 may also include additional boxes not shown and / or the boxes shown may be omitted, and the scope of this disclosure is not limited in this respect.
[0148] In step 602, the computing device 110 obtains non-administrative division data from the original data to be identified, excluding the administrative division information, based on the identified administrative division information.
[0149] In step 604, the computing device 110 performs noise word removal and equivalent word replacement on the non-administrative division data.
[0150] Regarding equivalent terms, these are, for example, words that can be considered equivalent in the medical industry, or words that are semantically equivalent. For example, "Lizhuang Health Clinic in Qihe Township" and "Lizhuang Health Station in Qihe Township" differ only in their word segmentation structure, with "health clinic" and "health station" being different. In reality, these two names usually belong to the same category and can be considered equivalent in the medical industry, thus belonging to "equivalent terms".
[0151] The equivalence thesaurus contains a large number of equivalence terms, which are classified, for example, through manual annotation or machine learning. Table 6 below illustrates some of the equivalence terms in the equivalence thesaurus.
[0152] Table 6
[0153]
[0154]
[0155] A method for equivalent term replacement of non-administrative division data includes, for example: a computing device 110 determines multiple sets of related terms, each set including an original term and an equivalent term, the original term and the equivalent term having consistent semantics when indicating a target object in the pharmaceutical industry; determines an association sequence number and a category for each set of related terms, the sequence number indicating the priority of each set of related terms; and based on the determined association sequence number, replaces and segments the original data to be identified using equivalent terms, such that the data replaced and segmented by equivalent terms includes equivalent terms and a predetermined identifier, the predetermined identifier indicating the segmentation position.
[0156] Regarding the predefined identifier, which is, for example but not limited to, "%", the predefined identifier indicates that the corresponding position is a separator. There is a priority order when using equivalent word replacement and segmentation. For example, the computing device 110 uses equivalent words to replace and segment the original data to be identified according to the "sequence number" (for example, as shown in Table 6; generally, the longer the original word, the higher the priority). For example, the original word "store company" in "Suzhou Huiren Pharmaceutical Store Co., Ltd." is replaced and segmented by the equivalent word % company % and the data after equivalent word replacement and segmentation is, for example, "Suzhou Huiren Pharmaceutical % company".
[0157] In step 606, the computing device 110 segments the data after noise word removal and equivalent word replacement based on a fixed vocabulary to generate a word segmentation result corresponding to the original data to be identified. The word segmentation result includes multiple keywords and multiple predetermined identifiers indicating segmentation positions. Fixed words include, for example, at least the full and abbreviations of provinces, cities, and districts / counties, as well as other conventional fixed phrases. For example, in the institution name "Beijing Normal University Affiliated Middle School Health Station," the words "normal university," "affiliated," "middle school," and "health station" are fixed words and do not need further segmentation. In some embodiments, after performing noise word removal and equivalent word replacement, the computing device 110, based on the ASCII code table, confirms whether the ASCII value of each character in the data after noise word removal and equivalent word replacement is outside a first predetermined numerical range (e.g., 48-57), so that all characters outside the first predetermined numerical range (e.g., 48-57) are removed. The main reason for adopting the above methods is that while a batch of noisy words can be removed by noise word removal and equivalent word replacement, letters, hyphens and other content often appear in the names of institutions. Medical retail institutions and medical terminals in China do not contain uppercase and lowercase letters or other English symbols. Therefore, in addition to Chinese characters and numbers, other symbols need to be removed. Thus, through the above methods, this disclosure can further filter noisy words.
[0158] Regarding the method for segmenting data after noise word removal and equivalent word replacement, it includes, for example, the computing device 110, after completing the noise word and equivalent word replacement, using fixed words to "segment" the original data to be identified. Taking "Suzhou Huiren Pharmaceutical Store Co., Ltd." as an example, after equivalent word replacement and segmentation, it becomes "Suzhou Huiren Pharmaceutical% Company"; after being segmented using a fixed word library, it becomes "Suzhou%Huiren%Pharmaceutical% Company", where "Suzhou" is geographical information, which will be extracted and stored separately, and the remaining parts will be separated one by one to generate word segmentation results corresponding to the original data to be identified. For example, Table 7 below shows the analysis results of the original data to be identified "Suzhou Huiren Pharmaceutical Store Co., Ltd.", Table 8 shows the word segmentation results of the original data to be identified "Pharmaceutical Store Co., Ltd. (Suzhou Huiren)", and Table 9 shows the word segmentation results of the original data to be identified "Suzhou Huiren Pharmaceutical Trading Co., Ltd."
[0159] Table 7
[0160]
[0161] Table 8
[0162]
[0163] Table 9
[0164]
[0165]
[0166] In step 608, the computing device 110 identifies numeric words in the raw data to be identified.
[0167] In step 610, the computing device 110 performs normalization processing on the identified numeric words in order to segment out keywords in numeric form from the original data to be identified. (The above has already been combined with...) Figure 4 The method for segmenting keywords in the form of numeric words in the original data to be identified has been explained, and will not be repeated here.
[0168] In step 612, the computing device 110 combines the multiple keywords included in the word segmentation results into preprocessed data without geographical information for matching with reference names.
[0169] In some embodiments, the computing device 110 can combine the segmented keywords into preprocessed data without geographic information for matching with reference names. For example, taking "Suzhou Huiren Pharmaceutical Store Co., Ltd." as an example, the preprocessed data without geographic information could be "Huiren Pharmaceutical Company" for matching with reference names.
[0170] Figure 7 A block diagram schematically illustrates an electronic device 700 suitable for implementing embodiments of the present invention. The electronic device 700 may be used to implement... Figures 2 to 6 The methods shown are 200 to 600. (For example...) Figure 7 As shown, the electronic device 700 includes a central processing unit (i.e., CPU 701), which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (i.e., ROM 702) or loaded from storage unit 708 into random access memory (i.e., RAM 703). The RAM 703 may also store various programs and data required for the operation of the electronic device 700. The CPU 701, ROM 702, and RAM 703 are interconnected via bus 704. An input / output interface (i.e., I / O interface 705) is also connected to bus 704.
[0171] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, output unit 707, and storage unit 708. CPU 701 executes the various methods and processes described above, such as executing methods 200 to 600. For example, in some embodiments, methods 200 to 600 may be implemented as computer software programs stored in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more operations of methods 200 to 600 described above may be performed. Alternatively, in other embodiments, CPU 701 may be configured to execute one or more actions of methods 200 to 600 by any other suitable means (e.g., by means of firmware).
[0172] It should be further noted that the present invention can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the present invention.
[0173] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0174] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0175] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.
[0176] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0177] These computer-readable program instructions can be provided to a processor in a voice interaction device, a general-purpose computer, a special-purpose computer, or a processing unit of another programmable data processing device, thereby producing a machine such that, when executed by the processing unit of the computer or other programmable data processing device, these instructions create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing device, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0178] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0179] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0180] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0181] The above are merely optional embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for identifying target objects in the pharmaceutical industry, comprising: Acquire raw data to be identified for use in indicating target objects in the pharmaceutical industry; Identify the administrative division information and channel type information in the original data to be identified. The channel type information includes: channel type subclass name, channel type class name, and channel type class number. Based on administrative division information, channel type information, and at least one of the following lexicons: a noise lexicon, a semantic equivalence lexicon, and a fixed lexicon, noise removal and word segmentation are performed on the original data to be identified to generate word segmentation results, which include multiple keywords. Generating word segmentation results includes: normalizing the identified numeric words to segment out keywords in numeric form from the original data to be identified, which includes: converting uppercase and / or lowercase Chinese numerals in the original data to be identified into Arabic numerals; determining whether the number of digits in the converted Arabic numerals is greater than or equal to a predetermined number of digits threshold; and responding to the determination of the number of digits in the converted Arabic numerals... If the number of digits is greater than or equal to a predetermined threshold, the converted Arabic numerals are removed; in response to determining that the number of digits of the converted Arabic numerals is less than the predetermined threshold, it is determined whether the converted Arabic numerals are located at the start or end position of the original data to be identified; in response to determining that the converted Arabic numerals are located at the start or end position of the original data to be identified, it is determined whether the data adjacent to the Arabic numerals located at the start or end position indicates a predetermined channel type; and in response to determining that the data adjacent to the Arabic numerals located at the start or end position does not indicate a predetermined channel type, the converted Arabic numerals are removed. Hash calculations are performed on multiple keywords included in the word segmentation results to confirm whether the word segmentation results match the reference name; and In response to the confirmation that the word segmentation result does not match the reference name, a semantic similarity analysis is performed on the reference name and the preprocessed data combined based on the word segmentation result, so as to identify the target object in the pharmaceutical industry to be identified based on the results of the similarity analysis.
2. The method according to claim 1, wherein performing hash calculation on multiple keywords included in the word segmentation result to confirm whether the word segmentation result matches the reference name includes: Calculate the sum of the hash values of multiple keywords included in the word segmentation result in order to generate the sum of hash values of the word segmentation result; Calculate the sum of the hash values of the multiple keywords included in the reference name in order to generate the sum of the reference name hash values; Confirm whether the sum of the hash values of the word segmentation results and the sum of the hash values of the reference names are equal; and In response to the fact that the sum of the hash values of the confirmed word segmentation results and the sum of the hash values of the reference name are equal, it is determined that the word segmentation results match the reference name.
3. The method according to claim 1 or 2, further comprising: In response to the confirmation that the word segmentation result matches the reference name, the target object in the pharmaceutical industry to be identified is identified as the target object associated with the reference name.
4. The method according to claim 1, wherein generating the word segmentation result includes: Based on the identified administrative division information, obtain the non-administrative division data from the original data to be identified, excluding the administrative division information; For non-administrative division data, noise word removal and equivalent word replacement are performed; and Based on a fixed lexicon, the data after noise word removal and equivalent word replacement is segmented to generate word segmentation results corresponding to the original data to be identified. The word segmentation results include multiple keywords and multiple predetermined identifiers indicating the segmentation positions.
5. The method according to claim 4, further comprising: Identify numeric words in the raw data to be identified; The identified numeric words are normalized in order to segment out keywords in numeric form from the original data to be identified; as well as The multiple keywords included in the word segmentation results are combined into preprocessed data without geographical information for matching with reference names.
6. The method according to claim 1, wherein identifying the administrative division information and channel type information in the original data to be identified includes: Identify multiple sets of keywords that are associated with different priority orders, each set of keywords including multiple predefined keywords; Among multiple keyword sets, determine the target keyword set containing the predetermined keywords included in the original data to be identified; Based on the priority order associated with the target keyword set, determine the channel type sub-category name that matches the original data to be identified; and Based on the determined channel type subclass name, determine the channel type category name and channel type category number that match the original data to be identified.
7. The method according to claim 1, wherein noise removal and word segmentation of the original data to be identified comprises: Multiple sets of related terms are identified, each set of related terms includes original terms and equivalent terms, and the original terms and equivalent terms have consistent semantics when indicating target objects in the pharmaceutical industry; Assign a sequence number and category to each group of related terms; the sequence number indicates the priority of each group of related terms. as well as Based on the determined sequence number of the association, the original data to be identified is replaced and segmented using equivalent terms, such that the data replaced and segmented by equivalent terms includes equivalent terms and predetermined identifiers, the predetermined identifiers indicating the segmentation bits.
8. The method according to claim 1, wherein noise removal and word segmentation of the original data to be identified comprises: Identify the overlapping portions of the preprocessed data and the reference names; Remove overlapping portions from the preprocessed data to obtain the remaining portion; In response to determining that a first predetermined confidence level condition is met, the matching confidence level between the original data to be identified and the reference name is determined to be a first level. A matching confidence level of the first level indicates that the original data to be identified and the reference name are matched. The first predetermined condition includes any of the following: Determine that the number of characters in the remaining portion is less than or equal to the first character count threshold; The remaining part is determined to have a character count greater than the second character count threshold and the remaining part and the overlapping part are associated with the same channel type information, and the second character count threshold is greater than the first character count threshold; The remaining part contains more than the first word count threshold and less than the second word count threshold, and the remaining part contains a pair of parentheses; The remaining part contains "original" or parentheses and "original"; The remaining part contains a pair of parentheses and the number of characters in the parentheses is less than the third character threshold, which is greater than the first character threshold and less than the second character threshold; It was determined that there was overlap between the preprocessed data and the reference name, and that the preprocessed data and the reference name had the same channel type subclass.
9. The method according to claim 8, wherein semantic similarity analysis of the word segmentation results and reference names further includes: In response to determining that a second predetermined confidence condition is met, the confidence level of the match between the original data to be identified and the reference name is determined to be the second level. The second predetermined confidence condition includes: It was determined that the word segmentation results of the preprocessed data and the reference name had overlapping parts after structural recombination, and the channel type classification information of the preprocessed data and the reference name were the same. In response to determining that a third predetermined confidence condition is met, a mismatch is identified between the original data to be identified and the reference name. The third predetermined confidence condition includes: The word segmentation results of the preprocessed data and the reference name have overlapping parts after structural recombination, and the channel type classification information of the preprocessed data and the reference name are different.
10. The method according to claim 1, wherein identifying the administrative division information and channel type information in the original data to be identified includes: The administrative division information in the name of the organization to be identified is based on the full name, abbreviation, former name and exclusion words of the province, city and district / county. The administrative division information includes province information, city information and district / county information. In response to the confirmation that the identified district or city information does not indicate a unique district or city, the administrative division information in the original data to be identified is identified using the lower-level administrative division information of the identified district or city information, or the administrative division information of the associated target object of the target object to be identified.
11. A computing device, comprising: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-10.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method of any one of claims 1-10.
Citation Information
Patent Citations
Method for processing medical institution data and method and device for constructing database
CN111899821A
Name matching method and device
CN114911999A
Method and equipment for identifying target object to be identified in pharmaceutical industry, and medium
CN115730595A