Target data determination method, device, computer device and storage medium

By receiving source data and matching with the pre-data in the mapping relationship set, and querying in the candidate data set after the matching failure, the problem of low data cleaning efficiency in traditional technology is solved, and efficient and accurate data matching and cleaning is achieved.

CN114020733BActive Publication Date: 2025-05-30SOFTIUM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111308637.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-05
Publication Date
2025-05-30
Estimated Expiration
2041-11-05

AI Technical Summary

Technical Problem

In traditional technology, the product flow data cleaning process is not efficient, resulting in low data matching efficiency and insufficient accuracy.

Method used

By receiving the source data, it matches it with the pre-data in the corresponding mapping relationship set. In the case of matching failure, the source data is queried in the candidate data set to obtain the query result set, and the target data matching the source data is determined in the target candidate data.

Benefits of technology

Automatic data cleaning is realized, which improves data cleaning efficiency and accuracy, reduces manual participation, and improves the accuracy of data matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114020733B_ABST
    Figure CN114020733B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a method, apparatus, computer device, storage medium, and computer program product for determining standard data. This method receives source data and matches the source data with the pre-data in the corresponding mapping relationship set. In the case of successful matching, without manual participation, it efficiently matches some of the source data with the target data to achieve automatic data cleaning and improve data cleaning efficiency. In the case of failed matching, it continues to query the source data in the candidate data set to obtain a query result set including at least one target candidate data, and determines the target data that matches the source data among the target candidate data, improving the accuracy of data matching. Further, when an enterprise uses the cleaned flow data for data analysis, it improves the accuracy of the data analysis results, which is beneficial for the enterprise to make decisions and control products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the technical field of pharmaceutical industry data processing, and particularly to a method, device, computer device, storage medium, and computer program product for determining target data. Background Art

[0002] Looking at most industrial pharmaceutical enterprises, the classic sales models of drugs or medical devices generally include the following two methods: self-operation and agency. In the distribution channels of pharmaceutical enterprises, accurate product (such as drugs or medical devices) flow data has become the basis for pharmaceutical enterprises to make decisions and control.

[0003] Pharmaceutical enterprises can collect flow data for marketing managers to understand the inventory data in each distribution channel and the sales volume data at the market terminal every month, so as to formulate sales plans adapted to future product demands. Among them, the flow data collected by pharmaceutical enterprises has the characteristics of huge data volume, multiple data sources, and low data quality. Summary of the Invention

[0004] In view of this, the embodiments of this specification are committed to providing a method, device, computer device, storage medium, and computer program product for determining target data to solve the technical problem of low efficiency in the process of cleaning product flow data in traditional technologies.

[0005] The embodiments of this specification provide a method for determining target data, the method comprising: receiving source data; the source data is used to represent the attribute data of pharmaceutical and medical products; dividing into a plurality of mapping relationship sets according to the attribute data of the pharmaceutical and medical products; the mapping relationship set includes pre-data and candidate data with an associated relationship;

[0006] matching the source data with the pre-data in the corresponding mapping relationship set; in the case of a matching failure, querying the source data in the candidate data set to obtain a query result set; wherein, the candidate data set includes a plurality of candidate data; wherein, the query result set includes at least one target candidate data; determining the target data that matches the source data from the target candidate data.

[0007] An embodiment of this specification provides a method for determining target data. The method includes: providing a data processing page for source data, where the source data is used to represent the attribute data of a medical device product, and there are multiple mapping relationship sets divided according to the attribute data of the medical device product. The mapping relationship set includes pre-data and candidate data with an associated relationship; displaying the processing information of the source data on the data processing page, where the processing information is generated based on the matching result between the source data and the pre-data in the corresponding mapping relationship set, and the processing information corresponds to a matching operation control. When the matching operation control is triggered, at least one target candidate data included in the query result set is displayed, where the target candidate data corresponds to a matching confirmation control. The query result set is obtained by querying the source data in the candidate data set when the source data fails to match, and the candidate data set includes multiple candidate data. When the matching confirmation control is triggered, the target data that matches the source data is determined from the target candidate data.

[0008] An embodiment of this specification provides a device for determining target data. The device includes: a source data receiving module for receiving source data, where the source data is used to represent the attribute data of a medical device product, and there are multiple mapping relationship sets divided according to the attribute data of the medical device product. The mapping relationship set includes pre-data and candidate data with an associated relationship; a source data matching module for matching the source data with the pre-data in the corresponding mapping relationship set; a source data query module for querying the source data in the candidate data set to obtain a query result set when the matching fails, where the candidate data set includes multiple candidate data, and the query result set includes at least one target candidate data; a target data determining module for determining the target data that matches the source data from the target candidate data.

[0009] An embodiment of this specification provides a target data determination device, which includes: a processing page providing module for providing a data processing page for source data; wherein the source data is used to represent the attribute data of a medical device product; multiple mapping relationship sets are divided according to the attribute data of the medical device product; each mapping relationship set includes pre-data and candidate data with an associated relationship; a processing information display module for displaying the processing information of the source data in the data processing page; wherein the processing information is generated based on the matching result between the source data and the pre-data in the corresponding mapping relationship set; the processing information corresponds to a matching operation control; a query result display module for displaying at least one target candidate data included in a query result set when the matching operation control is triggered; wherein the target candidate data corresponds to a matching confirmation control; the query result set is obtained by querying the source data in a candidate data set when the source data fails to match; wherein the candidate data set includes multiple candidate data; a target data determination module for determining, when the matching confirmation control is triggered, the target data that matches the source data from the target candidate data.

[0010] An embodiment of this specification provides a computing device, including a memory and a processor, where the memory stores a computer program, and the processor implements the method steps in the above embodiment when executing the computer program.

[0011] An embodiment of this specification provides a computer-readable storage medium, on which a computer program is stored, and the computer program implements the method steps in the above embodiment when executed by a processor.

[0012] An embodiment of this specification provides a computer program product, which includes instructions that, when executed by a processor of a computer device, enable the computer device to execute the method steps in the above embodiment.

[0013] In an embodiment of this specification, by receiving source data, matching the source data with the pre-data in the corresponding mapping relationship set, in the case of successful matching, without manual participation, efficiently matching some source data with target data to achieve automatic data cleaning and improve data cleaning efficiency. In the case of failed matching, continue to query the source data in the candidate data set to obtain a query result set including at least one target candidate data, and determine the target data that matches the source data from the target candidate data to improve the accuracy of data matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1a Shown is an interaction diagram of a target data determination method in a scenario example provided by an embodiment.

[0015] Figure 1b The figure shows an application environment diagram of a target data determination method provided in an embodiment;

[0016] Figure 2 The figure shows a schematic flowchart of a target data determination method provided in an embodiment;

[0017] Figure 3 The figure shows a schematic flowchart of a target data determination method provided in an embodiment;

[0018] Figure 4 The figure is a structural block diagram of a target data determination device provided in an embodiment;

[0019] Figure 5 The figure is a structural block diagram of a target data determination device provided in an embodiment;

[0020] Figure 6 The figure is an internal structure diagram of a computer device provided in an embodiment. Specific Embodiments

[0021] Next, the technical solutions in the embodiments of this specification will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0022] The following explains some terms involved in this specification. Flow data is generated in the sales link of the pharmaceutical industry. The flow data includes at least one of sales-related data, inventory-related data, shipping-related data, and procurement-related data. On the premise of obtaining the authorization of medical device manufacturers and institutions such as distributors, agents, and pharmacies, the flow data in the sales link of the pharmaceutical industry can be collected and stored in a data warehouse. A medical device manufacturer can be a pharmaceutical factory or enterprise that produces and sells medical devices. Medical devices (which can be simply referred to as products) include pharmaceutical products and medical device products. An institution can be understood as an institutional entity involved in the circulation process of products, and the institution types include at least one of hospitals, pharmacies, distributors, agents, and other institutions. A hospital can be understood as a medical and health institution in the real world. A pharmacy can be understood as a pharmacy in the real world, including a chain headquarters. A distributor can generally be understood as an institution responsible for the circulation and distribution of drugs. An agent can generally be understood as an institution responsible for the sales of drugs. Other institutions can be understood as some institutions not involved in the above in the product circulation process.

[0023] The source data may be the source organization data in the flow data, and the source organization data includes at least one of the source organization name, source organization code, and source organization address information. The source data may be the source product data in the flow data, and the source product data includes at least one of the source product name, source product code, source product grade, and source product trade name. The source data may be the source unit data in the flow data, and the source unit data includes the product unit.

[0024] The flow data may include that distributor A delivers X units of product C to agent B, and the unit of product C may be any one of box, package, and box. Distributor A and agent B may be the source organization data in the flow data. Product C may be the source product data in the flow data. The X units may be the source unit data in the flow data.

[0025] In some embodiments, different people may have different names for the same institutional entity. For example, for the "Suzhou No. 10 People's Hospital", among the people in Suzhou, through the abbreviated name "No. 10 Hospital", it can be understood as the Suzhou No. 10 People's Hospital; among some people outside Suzhou, through another abbreviated name "Suzhou No. 10 Hospital", it can be understood as the Suzhou No. 10 People's Hospital; however, for some other people, hearing "No. 10 Hospital" or "Suzhou No. 10 Hospital" may not be able to think of the Suzhou No. 10 People's Hospital. It can be understood that "No. 10 Hospital", "Suzhou No. 10 People's Hospital", and "Suzhou No. 10 Hospital" all refer to the same institutional entity. Therefore, it is necessary to construct a mapping relationship set. For example, establish a mapping relationship between "No. 10 Hospital" and "Suzhou No. 10 People's Hospital", and establish a mapping relationship between "Suzhou No. 10 Hospital" and "Suzhou No. 10 People's Hospital". The constructed mapping relationship set includes pre-data and candidate data with an associated relationship. For example, the pre-data can be "No. 10 Hospital", "Suzhou No. 10 Hospital", and the candidate data can be "Suzhou No. 10 People's Hospital". The pre-data can be understood as different names or calls of the same entity by different people, which may cause some ambiguities due to the change of the population. The candidate data can be understood as a relatively complete name for the entity to a certain extent, which can uniquely point to the entity and will not cause some ambiguities due to the change of the population. It can be understood that in this embodiment, only the institutional name is used as an example for illustrative purposes. Similarly, the names of the same product entity may also not be unified. By establishing a mapping relationship to form a product mapping relationship set, different names can all point to the same product entity. In some embodiments, the units of the products can also be unified. Agents, distributors, and pharmacies have different selling objects, and the demand of different selling objects is different, so different product units are used. For example, when a distributor transfers products to an agent, it can be in units of boxes; when an agent transfers products to a pharmacy, it can be in units of packages; when a pharmacy sells products, it can be in units of boxes. In summary, the attribute data of pharmaceutical and medical products can be institutional names, product names, and product units. According to the attribute data of pharmaceutical and medical products, there are multiple mapping relationship sets, which can be respectively a name mapping relationship set, a product mapping relationship set, and a unit mapping relationship set. It should be noted that the name mapping relationship set includes pre-institutional data and standard institutional names with an associated relationship. The product mapping relationship set includes pre-product data and standard product names with an associated relationship. The unit mapping relationship set includes pre-unit data and standard product units with an associated relationship.

[0026] In some embodiments, the candidate data set belongs to the target enterprise that produces product C. The candidate data set can be understood as the enterprise database of the target enterprise. The candidate data can be understood as the data stored in the enterprise database. For example, the data stored in the enterprise database may include information such as standard institution name, standard institution code, standard product name, standard product code, standard unit, etc. The enterprise database can also store institution attribute information. For example, when the institution is a hospital, the enterprise database can store information such as grade-A tertiary, public, private, etc.

[0027] Please refer to Figure 1a . In a specific scenario example, the candidate data set and the mapping relationship set are pre-deployed on the server. The user can access the web page provided by the server through the terminal. A data upload page is displayed on the operation interface of the terminal. The data upload page has a data import control. Through the data import control, the flow data can be uploaded to the server by specifying the flow direction to the data storage path or dragging the flow direction of the data. The server obtains the flow data, and the flow data can enter the data cleaning process. Among them, the file type of the flow data can be an Excel file or a ZIP package. The flow data includes source data, and the source data is used to represent the attribute data of the medical device product. There are multiple mapping relationship sets divided according to the attribute data of the medical device product; the mapping relationship set includes the pre-data and the candidate data with an associated relationship. The source data can be at least one of source institution data, source product data, and source unit data. The mapping relationship set can be at least one of the name mapping relationship set, the product mapping relationship set, and the unit mapping relationship set. The data cleaning process can include at least one of the institution matching process, the product matching process, and the unit matching process.

[0028] The mapping relationship set pre-deployed on the server includes the pre-data and the candidate data with an associated relationship. The server obtains the source data in the flow data, and the source data is used to represent the attribute data of the medical device product. The attribute data represented by the source data corresponds to a mapping relationship set. The server uses the source data to query in the mapping relationship set corresponding to the source data, that is, matches the source data with the pre-data in the corresponding mapping relationship set, and returns the matching result to the terminal.

[0029] The terminal can provide a data processing page for the source data. The processing information of the source data is displayed on the data processing page. Among them, the processing information is generated based on the matching result between the source data and the pre-data in the corresponding mapping relationship set. In some embodiments, the processing information includes the processed result and the to-be-processed result. The processed result includes the number of source data that have successfully matched the pre-data in the corresponding mapping relationship set. The processed data can be the source data that has successfully matched the pre-data in the corresponding mapping relationship set. The to-be-processed result includes the number of source data that have not matched the pre-data in the corresponding mapping relationship set. The to-be-processed data can be the source data that has not matched the pre-data in the corresponding mapping relationship set. The to-be-processed data corresponds to a matching operation control.

[0030] The terminal monitors the matching operation control. When the terminal monitors that the matching operation control is triggered, the terminal sends a matching operation instruction to the server. The matching operation instruction carries the query keyword corresponding to the source data. The server queries in the candidate data set according to the query keyword corresponding to the source data to obtain a query result set. Among them, the candidate data set includes multiple candidate data;

[0031] The server returns the query result set to the terminal. The query result set includes at least one target candidate data; the terminal displays at least one target candidate data included in the query result set. Among them, the target candidate data corresponds to a matching confirmation control.

[0032] The terminal monitors the matching confirmation control. When it monitors that the matching confirmation control is triggered, it selects the target data that matches the source data from the target candidate data, and the terminal sends a matching confirmation instruction to the server. The matching confirmation instruction carries the determined target data and the source data. The server establishes an association relationship between the source data and the determined target data to update the mapping relationship set.

[0033] Please refer to Figure 1b, embodiments of this specification provide a flow data cleaning system, and the target data determination method provided in this specification is applied to this flow data cleaning system. This flow data cleaning system may include a hardware environment formed by a terminal 110 and a server 120. The terminal 110 communicates with the server 120 through a network. A mapping relationship set and a candidate data set are pre-deployed on the server 120. The candidate data set includes multiple candidate data. The mapping relationship set corresponds to the attribute data of medical devices and products, and the mapping relationship set includes pre-data and candidate data with an associated relationship. Specifically, the terminal 110 uploads source data to the server 120, and the server 120 matches the source data with the pre-data in the corresponding mapping relationship set. In the case of a matching failure, the server 120 queries the source data in the candidate data set to obtain a query result set. Among them, the query result set includes at least one target candidate data. Determine the target data that matches the source data from the target candidate data.

[0034] Among them, the terminal 110 can be, but is not limited to, various personal computers, laptop computers, smartphones, tablet computers, and portable wearable devices. The server 120 can be implemented by an independent server or a server cluster composed of multiple servers. With the development of science and technology, some new types of computing devices may appear, such as quantum computing servers, and these new types of computing devices can also be applied to the embodiments of this specification.

[0035] Please refer to Figure 2 , embodiments of this specification provide a target data determination method. This target data determination method includes the following steps.

[0036] S210. Receive source data.

[0037] Among them, the source data is used to represent the attribute data of medical devices and products. Specifically, in some embodiments, the source data uploaded by the terminal to the server can be received. The file type of the source data can be an EXCEL file or a ZIP package.

[0038] In some embodiments, a cleaning instruction for flow data sent by the terminal to the server can be received. The cleaning instruction carries the source data, and the source data can be the name of the institution associated with the medical device and product. The source data can be the name of the medical device and product. The source data can be the unit of the medical device and product. In some embodiments, the source data can be keywords. Among them, the keywords can include the name of the institution. The keywords can include the name of the product, the name of the institution, and the product specification. The keywords can include the name of the product, the unit of the product, and the product specification.

[0039] S220. Match the source data with the pre-data in the corresponding mapping relationship set.

[0040] Among them, the source data is used to represent the attribute data of the medical device product. Different attributes correspond to different sets of mapping relationships, that is, there are multiple sets of mapping relationships divided according to the attribute data of the medical device product. The set of mapping relationships can be at least one of the name mapping relationship set, the product mapping relationship set, and the unit mapping relationship set. The set of mapping relationships includes pre-data and candidate data with an associated relationship. In some embodiments, the pre-data and the candidate data point to the same medical device institution entity. It should be noted that the pre-data and the candidate data can be different names of the same medical device institution. The pre-data and the candidate data can be different names of the same medical device product. The pre-data and the candidate data can be different packaging units of the same medical device product.

[0041] The pre-data can be understood as the "dirty data" of the flow data generated in the sales link of the pharmaceutical industry. The pre-data can be an abbreviated name of the institution name of the medical device institution, an alias of the medical device institution. The pre-data can be an abbreviated name of the product name of the medical device product, the trade name of the medical device product. The pre-data can be the name of the standard quantity for measuring the medical device product. The candidate data can be the name used by the pharmaceutical enterprise to refer to the medical device institution. To a certain extent, the candidate data can also be understood as the standard institution name of the medical device institution, the standard product name of the medical device product, and the standard product unit of the medical device product.

[0042] Specifically, the source data is used to represent the attribute data of the medical device product. The source data is queried in the set of mapping relationships corresponding to the source data, and the source data is matched with the pre-data in the corresponding set of mapping relationships to determine whether the source data is consistent with the pre-data in the corresponding set of mapping relationships. If the two are consistent, it is determined that the source data matches successfully in the corresponding set of mapping relationships. If the two are inconsistent, it is determined that the source data fails to match in the corresponding set of mapping relationships.

[0043] S230. In the case of a matching failure, query the source data in the candidate data set to obtain a query result set.

[0044] Among them, the candidate data set can be a set of candidate data. The candidate data set includes multiple candidate data. The candidate data can be understood as the standard institution name of the medical device institution defined by the medical device enterprise, the standard product name of the medical device product, and the standard product unit of the medical device product. The candidate data can also be understood as the standard institution name, standard product name, and standard product unit required by the drug regulatory department. In some embodiments, the candidate data set includes the candidate data in the set of mapping relationships.

[0045] Specifically, in the case of a matching failure, the source data is used to query the candidate data, and at least one target candidate data corresponding to the source data is recalled. Among them, at least one target candidate data can form a query result set. In some embodiments, at least one target candidate data in the query result set is sorted by confidence level.

[0046] S240. Determine the target data that matches the source data from the target candidate data.

[0047] Specifically, by querying the source data in the candidate data set, a query result set is obtained. In some embodiments, the terminal can display the target candidate data in the query result set. In response to a matching confirmation operation on any target candidate data, the confirmed target candidate data in the query result set is used as the target data that matches the source data. In some embodiments, the terminal can display the target candidate data in the query result set. In response to a matching confirmation operation on any target candidate data, the terminal sends a matching confirmation instruction to the server. The matching confirmation instruction carries the confirmed target candidate data, and the server uses the confirmed target candidate data as the target data that matches the source data. Further, the source data can be cleaned using the target data.

[0048] The above target data determination method, by receiving the source data, matches the source data with the pre-data in the corresponding mapping relationship set. In the case of a successful match, without manual participation, it efficiently matches some source data with the target data, realizes automatic data cleaning, and improves data cleaning efficiency. In the case of a matching failure, continue to query the source data in the candidate data set to obtain a query result set including at least one target candidate data, and determine the target data that matches the source data from the target candidate data, improving the accuracy of data matching. Further, when an enterprise uses the cleaned flow data for data analysis, the accuracy of the data analysis results is improved, which is beneficial for the enterprise to make decisions and control products.

[0049] In some embodiments, the target data determination method may further include: establishing an association relationship between the source data and the determined target data to update the mapping relationship set.

[0050] Specifically, during the current data cleaning process, the source data fails to successfully match the target data in the corresponding mapping relationship set. It is necessary to query the source data in the candidate data set to obtain a query result set, and determine the target data that matches the source data in the query result set. It can be understood that by querying the source data in the candidate data set, the target data that matches the source data has been identified. Further, it is known that the source data matches the determined target data, or rather, there is an association relationship between the source data and the determined target data. Therefore, an association relationship is established between the source data and the determined target data to update the mapping relationship set. In this embodiment, by updating the mapping relationship set, when the source data enters the data cleaning process next time, the source data is matched in the updated mapping relationship set, and the target data can be successfully matched without querying in the candidate data set, thus improving the data cleaning efficiency.

[0051] In some embodiments, the source data includes source institution data, the mapping relationship set includes a name mapping relationship set, and the candidate data set includes the enterprise data set of the target enterprise. Matching the source data with the pre-data in the corresponding mapping relationship set may include: matching the source institution data with the pre-institution data in the name mapping relationship set. Correspondingly, in the case of a matching failure, querying the source data in the candidate data set to obtain a query result set may include: in the case of a failure to match the source institution data, querying the source institution data in the enterprise data set and recalling the set of institution names corresponding to the source institution data.

[0052] Specifically, the source data is received, and the source data includes source institution data. The mapping relationship set corresponding to the source institution data is the name mapping relationship set. The name mapping relationship set includes pre-institution data and candidate institution names with an association relationship. The source institution data is used to match the pre-institution data in the name mapping relationship set. When there is no pre-institution data in the name mapping relationship set that is the same as the source institution data, it indicates a matching failure. It is necessary to further query the source institution data in the enterprise data set and recall at least one target institution name corresponding to the source institution data. The at least one target institution name constitutes the set of institution names. Further, the set of institution names includes at least one target candidate institution data, and the target institution data is determined from the target candidate institution data.

[0053] In some embodiments, the candidate data set is the enterprise data set of the target enterprise, and the enterprise data set includes candidate institution names. The enterprise data set includes the candidate institution names in the name mapping relationship set.

[0054] In some embodiments, when there is pre - institution data in the name mapping relationship set that is consistent with the source institution data, it indicates a successful match, and the candidate institution data associated with the pre - institution data can be determined as the target institution data. This embodiment can reduce the data cleaning work of the flowing data and improve the data cleaning efficiency. Further, after determining the target institution data, a name association relationship can be automatically established between the source institution data and the determined target institution data to update the name mapping relationship set. In some embodiments, the established name association relationship can be displayed, and the name association relationship corresponds to a relationship modification control. Among them, the relationship modification control includes a re - matching control and an undo - matching control. Monitor the relationship modification control. If it is monitored that the relationship modification control is triggered, in response to the relationship modification instruction, modify the name association relationship.

[0055] In this embodiment, first, the source institution data is matched with the pre - institution data in the name mapping relationship set to quickly match the target institution data corresponding to the source institution data. Then, in the case where the source institution data fails to match, query the source institution data in the enterprise dataset and recall the set of institution names corresponding to the source institution data to ensure the accurate determination of the target institution name corresponding to the source institution data and improve the accuracy of data cleaning.

[0056] In some embodiments, the name mapping relationship set can also be referred to as institution matching relationship data. The name mapping relationship set can be stored in the server in a batch import manner. When storing the institution matching relationship data in a batch import manner, an institution matching relationship template can be obtained from the server, and the institution matching relationship data can be generated according to the institution matching relationship template. In some embodiments, the institution matching relationship template can be provided to the user in the form of an EXCEL file. The institution matching relationship template can include dealer code, dealer name, original institution name, standard institution code, and standard institution name. In some embodiments, the name mapping relationship set can also be stored in the server in a single - item addition manner. A matching relationship addition control is provided on the interface of the terminal. When it is monitored that the matching relationship addition control is triggered on the terminal, a new matching relationship page is displayed. Determine the source institution data and the target institution data through the new matching relationship page and establish a name mapping relationship between the two.

[0057] In some embodiments, the target enterprise also has a set of institution aliases. This method for determining the target data can also include: in the case where the set of institution names cannot be queried in the enterprise dataset, query according to the source institution data in the institution alias dataset and recall the set of institution names corresponding to the source institution data.

[0058] Among them, the institutional alias dataset is used to store institutional aliases and the corresponding relationships between institutional aliases and candidate institutional names. Specifically, when receiving source data which includes source institutional data, the set of mapping relationships corresponding to the source institutional data is the name mapping relationship set. The name mapping relationship set includes pre-institutional data and candidate institutional names with an associated relationship. The source institutional data is used to match the pre-institutional data in the name mapping relationship set. When there is no pre-institutional data in the name mapping relationship set that is the same as the source institutional data, it indicates that the matching fails. It is necessary to further query in the enterprise dataset according to the source institutional data. When no set of institutional names is obtained by querying in the enterprise dataset, the source institutional data is used to query in the institutional alias dataset to obtain the set of institutional names corresponding to the source institutional data. Further, the set of institutional names includes at least one target candidate institutional data, and the target institutional data is determined from the target candidate institutional data. In this embodiment, by further querying in the institutional alias dataset, the set of institutional names corresponding to the source institutional data is recalled, providing accurate candidate institutional names to the user comprehensively, which is conducive to establishing a more accurate mapping relationship.

[0059] In some embodiments, the target enterprise belongs to the target industry, and the target industry has an industry dataset. The method for determining the target data may further include: when the matching of the source institutional data fails, query the source institutional data in the industry dataset to recall the set of institutional names corresponding to the source institutional data.

[0060] Specifically, when receiving source data which includes source institutional data, the set of mapping relationships corresponding to the source institutional data is the name mapping relationship set. The name mapping relationship set includes pre-institutional data and candidate institutional names with an associated relationship. The source institutional data is used to match the pre-institutional data in the name mapping relationship set. When there is no pre-institutional data in the name mapping relationship set that is the same as the source institutional data, it indicates that the matching fails. It is necessary to further query in the industry dataset according to the source institutional data to recall the set of institutional names corresponding to the source institutional data. Further, the set of institutional names includes at least one target candidate institutional data, and the target institutional data is determined from the target candidate institutional data.

[0061] In some embodiments, when no set of institutional names is obtained by querying in both the enterprise dataset and the institutional alias dataset, the source institutional data is used to query in the industry dataset to obtain the set of institutional names corresponding to the source institutional data.

[0062] In some embodiments, the target data determination method may further include: in the case of successful matching, determining the candidate data associated with the pre-data successfully matched with the source data as the target data. This enables quick matching to the target institutional data corresponding to the source institutional data, reduces human intervention in the data cleaning process, and improves the accuracy and efficiency of data cleaning.

[0063] In some embodiments, the source data includes at least one of source product data and source unit data. The set of mapping relationships includes at least one of a product mapping relationship set and a unit mapping relationship set. Matching the source data with the pre-data in the corresponding set of mapping relationships includes at least one of the following. Matching the source product data with the pre-product data in the product mapping relationship set. Or, matching the source unit data with the pre-unit data in the unit mapping relationship set.

[0064] Specifically, in some embodiments, the source data includes source product data. The set of mapping relationships corresponding to the product name attribute is the product mapping relationship set. The product mapping relationship set includes pre-product data and candidate product names with an associated relationship. Matching the source product data with the pre-product data in the product mapping relationship set. When there is no pre-product data in the product mapping relationship set that is the same as the source product data, it indicates a failed match. It is necessary to further query in the enterprise dataset based on the source product data, institutional name, and product specifications, and recall at least one corresponding target product name. The at least one target product name constitutes the product name set. Further, the product name set includes at least one target candidate product data, and the target product data is determined from the target candidate product data.

[0065] In some embodiments, the source data includes source unit data. The set of mapping relationships corresponding to the product unit attribute is the unit mapping relationship set. The unit mapping relationship set includes pre-unit data and candidate product units with an associated relationship. Matching the source unit data with the pre-unit data in the unit mapping relationship set. When there is no pre-unit data in the unit mapping relationship set that is the same as the source unit data, it indicates a failed match. It is necessary to further query in the enterprise dataset based on the source product name, product unit, and product specifications, and recall at least one corresponding target product unit. The at least one target product unit constitutes the product unit set. Further, the product unit set includes at least one target candidate product unit, and the target product unit is determined from the target candidate product units.

[0066] In this embodiment, the source product data in the flow data is directly cleaned into target product data through the product mapping relationship set; the source unit data in the flow data is directly cleaned into target product units through the unit mapping relationship set, reducing manual intervention and improving the accuracy and efficiency of data cleaning.

[0067] The embodiments of this specification provide a method for determining target data, and the method for determining target data includes the following steps.

[0068] S302. Receive source data.

[0069] Among them, the source data is used to represent the attribute data of the medical device product. There are multiple mapping relationship sets divided according to the attribute data of the medical device product. The mapping relationship set includes pre-data and candidate data with an associated relationship. The pre-data and the candidate data point to the same medical device institution entity.

[0070] In some embodiments, the pre-data and the candidate data point to the same medical device institution entity. The candidate data set includes the candidate data in the mapping relationship set. The source data includes source institution data, source product data, and source unit data.

[0071] In some embodiments, the mapping relationship set includes a name mapping relationship set, a product mapping relationship set, and a unit mapping relationship set. The name mapping relationship set includes pre-institution data and candidate institution data with an associated relationship. The product mapping relationship set includes pre-product data and candidate product data with an associated relationship. The unit mapping relationship set includes pre-unit data and candidate unit data with an associated relationship.

[0072] In some embodiments, the candidate data set includes the enterprise data set of the target enterprise. The target enterprise also has an institution alias data set. The target enterprise belongs to the target industry, and the target industry has an industry data set.

[0073] S304. Match the source institution data with the pre-institution data in the name mapping relationship set.

[0074] S306. When the source institution data is successfully matched, determine the candidate institution data associated with the pre-institution data successfully matched with the source institution data as the target institution data.

[0075] S308. When the source institution data fails to be matched, query the source institution data in the enterprise data set and recall the set of institution names corresponding to the source institution data.

[0076] Among them, the set of institution names includes at least one target candidate institution data.

[0077] In some embodiments, when the set of institution names is not retrieved from the enterprise data set, query according to the source institution data in the institution alias data set and recall the set of institution names corresponding to the source institution data.

[0078] In some embodiments, in the case where the source institution data matching fails, query the source institution data in the industry dataset, and recall the set of institution names corresponding to the source institution data.

[0079] S310. Determine the target institution data that matches the source institution data in the target candidate institution data.

[0080] S312. Establish an association relationship between the source institution data and the determined target institution data to update the set of name mapping relationships.

[0081] S314. Match the source product data with the pre-product data in the product mapping relationship set.

[0082] S316. In the case where the source product data matches successfully, determine the candidate product data associated with the pre-product data that successfully matches the source product data as the target product data.

[0083] S318. In the case where the source product data matching fails, query the source product data in the enterprise dataset, and recall the set of product names corresponding to the source product data.

[0084] Wherein the set of product names includes at least one target candidate product data.

[0085] S320. Determine the target product data that matches the source product data in the target candidate product data.

[0086] S322. Establish an association relationship between the source product data and the determined target product data to update the product mapping relationship set.

[0087] S324. Match the source unit data with the pre-unit data in the unit mapping relationship set.

[0088] S326. In the case where the source unit data matches successfully, determine the candidate unit data associated with the pre-unit data that successfully matches the source unit data as the target unit data.

[0089] S328. In the case where the source unit data matching fails, query the source unit data in the enterprise dataset, and recall the set of product units corresponding to the source unit data.

[0090] Wherein, the set of product units includes at least one target candidate product unit.

[0091] S330. Determine the target product unit that matches the source unit data in the target candidate unit data.

[0092] S332. Establish an association relationship between the source unit data and the determined target product unit to update the unit mapping relationship set.

[0093] Please refer to Figure 3 , the embodiments of this specification provide a method for determining target data. The method for determining target data includes the following steps.

[0094] S410. Provide a data processing page for source data.

[0095] Among them, the source data is used to represent the attribute data of the medical device product; there are multiple mapping relationship sets divided according to the attribute data of the medical device product; the mapping relationship set includes pre-data and candidate data with an associated relationship.

[0096] S420. Display the processing information of the source data in the data processing page.

[0097] Among them, the processing information is generated based on the matching result between the source data and the pre-data in the corresponding mapping relationship set; the processing information corresponds to a matching operation control.

[0098] S430. When the matching operation control is triggered, display at least one target candidate data included in the query result set.

[0099] Among them, the target candidate data corresponds to a matching confirmation control; the query result set is obtained by querying the source data in the candidate data set when the source data fails to match; among them, the candidate data set includes multiple candidate data.

[0100] S440. When the matching confirmation control is triggered, determine the target data that matches the source data among the target candidate data.

[0101] In some embodiments, the processing information includes: processed results and / or to-be-processed results; among them, the processed results are generated based on the number of source data that have successfully matched the pre-data in the corresponding mapping relationship set; the to-be-processed results are generated based on the number of source data that have not matched the pre-data in the corresponding mapping relationship set.

[0102] For the specific limitations on the method for determining target data applied to the terminal, reference can be made to the limitations on the method for determining target data in the above text, which will not be elaborated here.

[0103] It should be understood that although the steps in the above flowcharts are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0104] Please refer to Figure 4 , the embodiment of this specification provides a target data determination device 400. The determination device 400 includes a source data receiving module 410, a source data matching module 420, a source data query module 430, and a target data determination module 440.

[0105] The source data receiving module 410 is configured to receive source data; the source data is used to represent the attribute data of a medical device product; there are multiple mapping relationship sets divided according to the attribute data of the medical device product; the mapping relationship set includes pre-data and candidate data with an associated relationship.

[0106] The source data matching module 420 is configured to match the source data with the pre-data in the corresponding mapping relationship set.

[0107] The source data query module 430 is configured to, in the case of a matching failure, query the source data in the candidate data set to obtain a query result set; wherein, the candidate data set includes multiple candidate data; wherein, the query result set includes at least one target candidate data.

[0108] The target data determination module 440 is configured to determine the target data that matches the source data from the target candidate data.

[0109] Please refer to Figure 5 , the embodiment of this specification provides a target data determination device 500. The determination device 500 includes a processing page providing module 510, a processing information display module 520, a query result display module 530, and a target data determination module 540.

[0110] The processing page providing module 510 is configured to provide a data processing page for the source data; wherein, the source data is used to represent the attribute data of a medical device product; there are multiple mapping relationship sets divided according to the attribute data of the medical device product; the mapping relationship set includes pre-data and candidate data with an associated relationship.

[0111] The processing information display module 520 is used to display the processing information of the source data on the data processing page; wherein, the processing information is generated based on the matching result between the source data and the pre-data in the corresponding mapping relationship set; the processing information correspondingly has a matching operation control.

[0112] The query result display module 530 is used to display at least one target candidate data included in the query result set when the matching operation control is triggered; wherein, the target candidate data correspondingly has a matching confirmation control; the query result set is obtained by querying the source data in the candidate data set when the source data fails to match; wherein, the candidate data set includes multiple candidate data.

[0113] The target data determination module 540 is used to determine the target data that matches the source data among the target candidate data when the matching confirmation control is triggered.

[0114] For the specific limitations of the target data determination device, reference can be made to the limitations of the target data determination method in the above text, which will not be elaborated here. Each module in the above target data determination device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0115] In some embodiments, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements a target data determination method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0116] Those skilled in the art can understand, Figure 6The structure shown is only a block diagram of some of the structures related to the solution disclosed in this specification, and does not constitute a limitation on the computer device to which the solution disclosed in this specification is applied. Specifically, the computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0117] In some embodiments, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the method steps in the above embodiments are implemented.

[0118] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method steps in the above embodiments are implemented.

[0119] In some embodiments, a computer program product is further provided. The computer program product includes instructions, and when the instructions are executed by a processor of a computer device, the method steps in the above embodiments are implemented.

[0120] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this specification may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0121] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0122] The above are only the preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent replacements, etc. made within the spirit and principles of this specification shall be included within the protection scope of this specification.

Claims

1. A method for determining target data, characterized in that, the method includes: Receiving source data; the source data is used to represent the attribute data of a medical device product; there are multiple mapping relationship sets divided according to the attribute data of the medical device product; the mapping relationship set includes pre-data and candidate data with an associated relationship; the source data includes product name, product unit, and product specification; the pre-data is the dirty data of the flow data generated in the medical industry sales link; the candidate data includes the standard institution name of the medical institution defined by the medical device enterprise, the standard product name of the medical device product, and the standard product unit of the medical device product; Matching the source data with the pre-data in the corresponding mapping relationship set; In the case of a failed match, querying the source data in the candidate data set to obtain a query result set; wherein, the candidate data set includes multiple candidate data; wherein, the query result set includes at least one target candidate data; Determining the target data that matches the source data among the target candidate data, and cleaning the source data with the target data.

2. The method according to claim 1, characterized in that, the method further includes: Establishing an association relationship between the source data and the determined target data to update the mapping relationship set.

3. The method according to claim 1, characterized in that, the source data includes source institution data, the mapping relationship set includes a name mapping relationship set, and the candidate data set includes the enterprise data set of the target enterprise; the matching the source data with the pre-data in the corresponding mapping relationship set includes: Matching the source institution data with the pre-institution data in the name mapping relationship set; The querying the source data in the candidate data set to obtain a query result set in the case of a failed match includes: In the case of a failed match of the source institution data, querying the source institution data in the enterprise data set and recalling the set of institution names corresponding to the source institution data.

4. The method according to claim 3, characterized in that, the target enterprise also has a set of institution aliases; the method further includes: In the case that the set of institution names is not retrieved from the enterprise data set, querying according to the source institution data in the set of institution aliases and recalling the set of institution names corresponding to the source institution data.

5. The method according to claim 3, characterized in that, the target enterprise belongs to the target industry, and the target industry has an industry data set; the method further includes: In the case of a failed match of the source institution data, querying the source institution data in the industry data set and recalling the set of institution names corresponding to the source institution data.

6. The method according to any one of claims 1 to 5, characterized in that, the method further includes: In the case of a successful match, determining the candidate data associated with the pre-data that successfully matches the source data as the target data.

7. The method according to any one of claims 1 to 5, characterized in that, The source data includes at least one of source product data and source unit data; the mapping relationship set includes at least one of a product mapping relationship set and a unit mapping relationship set; the matching of the source data with the pre-data in the corresponding mapping relationship set includes at least one of the following: Matching the source product data with the pre-product data in the product mapping relationship set; Matching the source unit data with the pre-unit data in the unit mapping relationship set.

8. The method according to any one of claims 1 to 5, characterized in that, The pre-data and the candidate data point to the same medical device institution entity; The candidate data set includes the candidate data in the mapping relationship set.

9. A method for determining target data, characterized in that, The method includes: Providing a data processing page for the source data; wherein, the source data is used to represent the attribute data of medical devices; there are multiple mapping relationship sets divided according to the attribute data of the medical devices; the mapping relationship set includes pre-data and candidate data with an associated relationship; the source data includes product name, product unit, and product specification; the pre-data is the dirty data of the flow data generated on the sales link of the pharmaceutical industry; the candidate data includes the standard institution name of the medical device institution defined by the medical device enterprise, the standard product name of the medical device product, and the standard product unit of the medical device product; Displaying the processing information of the source data on the data processing page; wherein, the processing information is generated based on the matching result between the source data and the pre-data in the corresponding mapping relationship set; the processing information corresponds to a matching operation control; When the matching operation control is triggered, displaying at least one target candidate data included in the query result set; wherein, the target candidate data corresponds to a matching confirmation control; the query result set is obtained by querying the source data in the candidate data set when the source data fails to match; wherein, the candidate data set includes multiple candidate data; When the matching confirmation control is triggered, determining the target data that matches the source data from the target candidate data, and cleaning the source data with the target data.

10. The method according to claim 9, characterized in that, The processing information includes: processed result and / or to-be-processed result; wherein, the processed result is generated based on the number of source data that successfully match the pre-data in the corresponding mapping relationship set; the to-be-processed result is generated based on the number of source data that do not match the pre-data in the corresponding mapping relationship set.

11. A target data determination device, characterized in that, The device includes: A source data receiving module, configured to receive source data; the source data is used to represent the attribute data of a medical device product; there are multiple mapping relationship sets divided according to the attribute data of the medical device product; the mapping relationship set includes pre-data and candidate data with an associated relationship; the source data includes product name, product unit, and product specification; the pre-data is the dirty data of the flow data generated in the medical industry sales link; the candidate data includes the standard institution name of the medical device institution defined by the medical device enterprise, the standard product name of the medical device product, and the standard product unit of the medical device product. A source data matching module, configured to match the source data with the pre-data in the corresponding mapping relationship set. A source data query module, configured to query the source data in the candidate data set in case of a matching failure to obtain a query result set; wherein, the candidate data set includes multiple candidate data; wherein, the query result set includes at least one target candidate data. A target data determination module, configured to determine target data that matches the source data from the target candidate data, and clean the source data with the target data.

12. A target data determination device Characterized in that The device includes: A processing page providing module, configured to provide a data processing page for the source data; wherein, the source data is used to represent the attribute data of a medical device product; there are multiple mapping relationship sets divided according to the attribute data of the medical device product; the mapping relationship set includes pre-data and candidate data with an associated relationship; the source data includes product name, product unit, and product specification; the pre-data is the dirty data of the flow data generated in the medical industry sales link; the candidate data includes the standard institution name of the medical device institution defined by the medical device enterprise, the standard product name of the medical device product, and the standard product unit of the medical device product. A processing information display module, configured to display the processing information of the source data in the data processing page; wherein, the processing information is generated based on the matching result between the source data and the pre-data in the corresponding mapping relationship set; the processing information corresponds to a matching operation control. A query result display module, configured to display at least one target candidate data included in the query result set when the matching operation control is triggered; wherein, the target candidate data corresponds to a matching confirmation control; the query result set is obtained by querying the source data in the candidate data set in case of a failure in matching the source data; wherein, the candidate data set includes multiple candidate data. A target data determination module, configured to determine target data that matches the source data from the target candidate data when the matching confirmation control is triggered, and clean the source data with the target data.

13. A computer device, including a memory and a processor, the memory stores a computer program, Characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.

15. A computer program product, comprising instructions in the computer program product, characterized in that, when the instructions are executed by a processor of a computer device, the computer device is enabled to execute the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Food material entity linking method and device between multi-source food material data

    CN111708891A

  • Medical data processing method and system and computer readable medium

    CN113127473A