Data processing method and device, electronic equipment and storage medium

By using artificial intelligence to identify and process data element tags, the problem of inconsistent data naming across different systems has been solved, data standardization has been achieved, the efficiency and accuracy of data governance have been improved, and the data exchange and analysis process has been simplified.

CN115391339BActive Publication Date: 2025-11-28BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210967080.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-11
Publication Date
2025-11-28
Estimated Expiration
2042-08-11

AI Technical Summary

Technical Problem

Because different systems use different naming conventions or definitions for data, data sharing and exchange are difficult, and existing technologies struggle to effectively standardize data processing.

Method used

By employing artificial intelligence, the system acquires the attributes of target elements, determines their primary and secondary labels, identifies data categories, and performs standardization processing based on these categories, thereby achieving automatic data conversion from non-standardized to standardized.

Benefits of technology

It improves the efficiency and accuracy of data governance, frees up human and material resources, simplifies data exchange and analysis processes, and provides visualization and ease of use for standardized data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115391339B_ABST
    Figure CN115391339B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data processing method and device, electronic equipment, storage medium and computer program product, relating to the technical fields of artificial intelligence, big data, machine learning, natural language processing and the like. The specific implementation scheme is: obtaining target data, the target data comprising at least one target element; based on the attributes of the target element, obtaining a first label and a second label of the target element, wherein the first label represents the semantics of the target element, and the second label represents the value characteristics of the target element; based on the first label and the second label of the target element, determining the data category of the target element in an artificial intelligence manner; and based on the data category of the target element, performing standardization processing on the target element to obtain standardized data of the target element. The present disclosure provides technical support for the automatic processing of data from non-standardization to standardization, and can improve the data governance efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of artificial intelligence, big data, machine learning, natural language processing, and the like, and more particularly to a data processing method and device, an electronic device, a storage medium, and a computer program product. BACKGROUND

[0002] In related technologies, there are scenarios where the same data appears in different systems or the same system with different naming methods or definitions, making data sharing and exchange extremely difficult. Therefore, how to standardize data has become a technical problem to be solved. SUMMARY

[0003] The present disclosure provides a data processing method, device, electronic device, storage medium, and computer program product.

[0004] According to an aspect of the present disclosure, a data processing method is provided, comprising:

[0005] obtaining target data, the target data comprising at least one target element;

[0006] based on the attributes of the target element, obtaining a first label and a second label of the target element, wherein the first label represents the semantics of the target element, and the second label represents the value characteristics of the target element;

[0007] based on the first label and the second label of the target element, determining the data category of the target element using an artificial intelligence method;

[0008] based on the data category of the target element, performing standardization processing on the target element to obtain standardized data of the target element.

[0009] According to another aspect of the present disclosure, a data processing device is provided, comprising:

[0010] a first obtaining unit configured to obtain target data, the target data comprising at least one target element;

[0011] a second obtaining unit configured to obtain a first label and a second label of the target element based on the attributes of the target element, wherein the first label represents the semantics of the target element, and the second label represents the value characteristics of the target element;

[0012] a first determining unit configured to determine the data category of the target element using an artificial intelligence method based on the first label and the second label of the target element;

[0013] The third acquisition unit is configured to perform standardization processing on the target element based on a data category of the target element to obtain standardized data of the target element.

[0014] According to still another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method in any embodiment of the present disclosure.

[0016] According to still another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, the computer instructions being used to cause a computer to perform the method in any embodiment of the present disclosure.

[0017] According to still another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method in any embodiment of the present disclosure.

[0018] The technical solution of the present disclosure provides technical support for automatic processing of data from non-standardization to standardization, and improves data governance efficiency.

[0019] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0020] The accompanying drawings serve to better understand the present solution and do not constitute a limitation of the present disclosure. Among them:

[0021] Figure 1 is an application scenario of an embodiment of the present disclosure Figure 1 ;

[0022] Figure 2 is an application scenario of an embodiment of the present disclosure Figure 2 ;

[0023] Figure 3 is a flowchart of a data processing method of an embodiment of the present disclosure Figure 1 ;

[0024] FIG. 4(a) is a schematic diagram of a first target table according to an embodiment of the present disclosure;

[0025] FIG. 4(b) is a schematic diagram of a second target table according to an embodiment of the present disclosure;

[0026] Figure 5is a flowchart of a data processing method of an embodiment of the present disclosure Figure 2 ;

[0027] Figure 6 is a flowchart of a data processing method of an embodiment of the present disclosure Figure 3 ;

[0028] Figure 7 is an implementation architecture diagram of a data processing method of an embodiment of the present disclosure Figure 1 ;

[0029] Figure 8 is an implementation architecture diagram of a data processing method of an embodiment of the present disclosure Figure 2 ;

[0030] Figure 9 is a composition diagram of a data processing device of an embodiment of the present disclosure

[0031] Figure 10 is a block diagram of an electronic device for implementing a data processing method of an embodiment of the present disclosure. DETAILED DESCRIPTION

[0032] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, and should be considered as merely exemplary. Thus, those of ordinary skill in the art will recognize various changes and modifications of the embodiments described herein, without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.

[0033] The present disclosure is directed to a scenario in which there is a large amount of business data. In actual applications, the scenario can be a scenario in which patients seek medical treatment, in which business data such as patient outpatient medical records, inpatient medical records, surgical medical records, and death medical records can be generated. Each hospital stores the generated business data in its own medical system. Among them, patient outpatient medical records, inpatient medical records, surgical medical records, and death medical records record patient information in the form of a table. For example, in the table of outpatient medical records, the name, gender, birth date, past medical history, and current medical condition of the patient are recorded. Figure 1 As shown, when needed, a unit in need, such as the Health and Family Planning Commission or a research unit, collects the data stored in the medical systems of all hospitals to analyze the medical condition. For example, the age of a high-risk group of a certain disease, the region where the high-risk group of a certain disease is located, and the like.

[0034] Figure 1 N in the above formula represents the number of hospitals, which is a positive integer greater than or equal to 1. It can be understood that the medical system of each hospital can be one, or two or more.

[0035] The scenario of a large amount of business data can also be a scenario of storing student information, in which personal information of students such as name, gender, age, political face, address, etc. is stored as business data into the business system of each school. Among them, the personal information of each student of different schools is recorded in the form of a table. For example, the aforementioned personal information is recorded in the student personal information table. As shown in the table, when needed, the unit in need such as the education committee collects the data stored in the business system of all schools to analyze the academic situation. For example, analyze the gender of the children ranked in the top 1000 in the middle school entrance examination; the district or county where the family of the child admitted to the first batch of college entrance examination is located. Figure 2

[0036] Figure 2 In the formula, M represents the number of schools, which is a positive integer greater than or equal to 1. It can be understood that the business system of each school can be one, or two or more.

[0037] Taking the medical system as an example, in the medical systems of different hospitals, the same business data will be stored in different ways such as naming or writing. For example, for outpatient medical records, some hospitals name it as “outpatient medical record”, some hospitals name it as “outpatient and emergency medical record”, some hospitals name it as “XX Children's Hospital Outpatient and Emergency Medical Record”, etc. The naming of outpatient medical records is various. For example, for the gender item in the outpatient medical record, it is written as “male” in some hospital business systems, and it is written as “male, male adult or boy” in some hospital business systems. The writing of the gender item is different.

[0038] Taking the business system of a school as an example, in some business systems, the name item is named as “name”, in some business systems, the name item is named as “name”, and in some business systems, the name item is named as “surname”.

[0039] From the foregoing, the naming or writing of the same business data in different medical systems is different, which causes great difficulty in data exchange, analysis and integration in different medical systems.

[0040] The technical solution of the present disclosure can solve the inconsistency problem of the business data in the aforementioned different medical systems or school business systems, can standardize the business data in different medical systems or different school business systems, obtain standardized business data, and can provide the unit in need such as the health and construction committee or the education committee with standard business data, thereby facilitating the efficient analysis of the disease or academic situation of the unit in need, and improving the analysis efficiency.

[0041] ​If the business data in different medical systems or different school business systems is regarded as non-standardized data, the technical solution of the present disclosure can be regarded as a solution for processing non-standardized data into standardized data. And the solution realized by relying on artificial intelligence (AI) is a solution for automatically realizing standardization.

[0042] If the solution for processing non-standardized data into standardized data can be regarded as a data governance solution, the solution of the present disclosure realized by relying on artificial intelligence (AI) can be regarded as an intelligent data governance solution. Compared with the solution for manually governing data in the related art, human and material resources can be liberated, and the governance efficiency and accuracy can be improved.

[0043] The processing logic of the data processing method of the present disclosure can be deployed in any reasonable terminal or server. Among them, the server includes a general server, a cloud server, a server for a professional field such as a (disease or academic) data analysis server. The terminal includes but is not limited to a tablet computer, an all-in-one machine, a desktop computer, a mobile phone, etc. It is preferred to be deployed in a server.

[0044] The technical solution of the present disclosure will be further described below.

[0045] Figure 3 is the flowchart of the data processing method of the embodiment of the present disclosure Figure 1 . As Figure 3 shown, the method comprises:

[0046] S301: obtaining target data, the target data comprising at least one target element;

[0047] It can be understood that in the case of storing data in the form of a table in a medical system or a school business system, the target element can be obtained by collecting each item of data stored in the form of a table in the medical system or the school business system. For example, collecting the name data, collecting the gender data. The target element can also be obtained by collecting the table stored in the medical system or the school business system and reading part or all of the item data in the table. For example, collecting the outpatient medical record table and reading the name (or surname), gender and other item data in the table as the target element.

[0048] In this step, each item of data in the table can be regarded as an element of different kinds. The target element is an element of at least one of the different kinds.

[0049] Exemplarily, taking the outpatient medical record table stored in the medical system as an example, the table includes several different kinds of data such as name, gender, date of birth, previous medical history, current medical condition, etc. The target element can be the data of the name kind. Or, the target element is the data of the name and gender kinds. Or, the target element is the data of all kinds.

[0050] The at least one target element included in the target data can be single or multiple data of the same kind. Taking the outpatient medical record table stored in the medical system as an example, the at least one target element can be the name or gender in a certain outpatient medical record table, or the name or gender in multiple outpatient medical record tables.

[0051] The at least one target element included in the target data can be single or multiple data of different kinds. Taking the outpatient medical record table stored in the medical system as an example, the at least one target element can be the name and gender in a certain outpatient medical record table, or the name and gender in multiple outpatient medical record tables.

[0052] In a general way, when the collected target element is data of a single kind, the number of the single kind can be one or multiple. Taking the name as an example, the collected target element can be the name in a certain outpatient medical record table, or the name in each of multiple outpatient medical record tables. When the collected target element is data of multiple kinds, the data of each kind can be one or multiple. Taking the name and gender as an example, the collected target element can be the name and gender in a certain outpatient medical record table, or the name and gender in each of multiple outpatient medical record tables.

[0053] In an optional solution, the target data can be acquired in response to the collection instruction. The acquisition instruction can be automatically generated when the collection period arrives. For example, the collection instruction can be generated when the preset collection period arrives, and the data in the medical system or the school business system can be collected in response to the collection instruction. The collection period can be in units of seconds, minutes, hours, days, months, or years, such as 30 days, and the data collection can be performed every 30 days.

[0054] The acquisition instruction can be generated based on the input operation of the user. For example, the processing logic of the data processing method is deployed in the server, and the display interface of the server can present a function key. The user clicks, presses, or performs other input operations on the function key. When the input operation is detected, the collection instruction is generated, and the data in the medical system or the school business system is collected in response to the collection instruction.

[0055] S302: Based on the attribute of the target element, a first label and a second label of the target element are obtained, wherein the first label represents the semantic of the target element, and the second label represents the value feature of the target element;

[0056] If each kind of data included in the table is regarded as an item of data, and each item of data is regarded as an element, the table records the relevant information of each element, such as the naming information and the value information of each element. Based on this, the attribute of an element includes the naming information of the element. Since the naming methods used by different hospitals or schools may be different, for the element of name, the naming information can be one or more of name, surname, name, etc. The attribute of an element also includes the value information of the element. Since the writing methods used by different hospitals or schools may be different, for the element of gender, the value can be one or more of male (female), male (female), adult male (female) gender, male (female) child, etc.

[0057] The semantics of an element refers to what kind of element it is, such as the element of "name" or the element of "gender". The value feature of an element represents the range of values that the element can take, such as the value range of the element of gender, which is male and female, including the two values of male and female.

[0058] S303: Based on the first label and the second label of the target element, the data category of the target element is determined by using an artificial intelligence method;

[0059] In this step, based on the semantics and value feature of the target element, the data category of the target element is determined by using an AI method. It is equivalent to combining the semantics and value of the element, and using an AI method to determine or identify the data category.

[0060] Generally, the data category corresponds to the element. If the target element is the element of name, the data category is the category of name, and if the target element is the element of gender, the data category is the category of gender.

[0061] It can be understood that in actual application, there are many data categories, and for a target element, it is necessary to determine what kind of data category corresponds to it by using an AI method. Among them, the AI method is any reasonable AI algorithm, such as classification algorithm, logistic regression algorithm, decision tree, etc. In implementation, the two labels of the element can be input into at least one of the above algorithm models, and the algorithm model is processed and outputs the data category corresponding to the target element.

[0062] S304: Based on the data category of the target element, the target element is standardized to obtain the standardized data of the target element.

[0063] In this step, based on the data category of the target element, the target element of the target table is standardized, and the non-standard data in the target table is standardized, and the standardized data of the target element is obtained. Considering that the related information of the target element includes naming information and value information, the standardization of the target element includes standardization of the naming information and / or standardization of the value information, and the standardized data of the target element with standardized naming information and / or value information is obtained.

[0064] In S301-S304, based on the obtained attribute of the target element, the first label representing the semantic of the element and the second label representing the value feature of the element are obtained. Based on the first label and the second label of the target element, the data category of the target element is determined by AI, and the target element of the target table is standardized based on the data category of the target element, and the standardization of the target element is realized. It is a scheme for processing non-standardized data into standardized data, and it is an automatic processing scheme for standardization. Compared with the scheme of manually managing data (manually processing standardized data) in related art, it can liberate manpower and material resources and improve management efficiency.

[0065] It can be seen that the technical scheme of the present disclosure provides technical support for the automatic processing of data from non-standardization to standardization.

[0066] In addition, the two labels are combined, and the AI method is used to determine or identify the data category. Considering that the AI algorithm has strong stability and robustness, the determination or identification based on the AI method can improve the determination or identification accuracy of the data category. The combination of the two labels and the use of robust algorithms can greatly improve the accuracy, thereby improving the accuracy of standardization and achieving accurate management of data.

[0067] For example, assume that the collected target element is "name" (the name is named in English in the medical system), based on the naming information and value information of the target element "name", the semantic and value feature of the target element "name" are obtained, and the semantic and value feature are input into the AI model, and the AI model is used to identify the data category of the target element. The AI model outputs the result that it is a name data category. The name belonging to the name data category is standardized, and the name element named "name" in the medical system is changed to be named "name". Or, the name element named "surname" in the medical system is changed to be named by the unified standard "name", realizing the standardized naming of the name element.

[0068] It can be understood that if the table stored in the medical system or school business system is regarded as the first target table, the data recorded in the first target table can be data named and valued in a standard form, and can also be data named and / or valued in a non-standard form. For example, there can be a table in the medical system, and the element of the name of the table is named surname or English name, and the standard name is name. After the technical solutions of S301-S304, the standardized data of the target element is obtained, and based on the standardized data of the target element, the second target table is obtained. In the second target table, the target element is recorded in the form of standardized data. It can be understood that the data in the first target table can be recorded in a non-standard form, and the data in the second target table is recorded in a standard form. For example, in the second target table, the name of the element is named name, not surname or name. In the second target table, both the naming of the element and the value of the element are standardized data. This scheme of processing non-standardized data into standardized data can effectively improve the data governance efficiency.

[0069] In actual application, a standard table can be set in advance, which can be regarded as a third target table, which is different from the first and second target tables. The element in the standard table can be designed as a standardized name, and the value of the element is required to be a standardized value. After the scheme of S301-S304, the standardized data of the target element, such as the standardized value, is obtained, and the standardized value is added to the standard table, together with the standardized name of the element already existing in the standard table, so that the standard table with the standardized name and the standardized value is obtained.

[0070] For example, for the case that the value of the element in the first target table is not standard or uniform, as shown in FIG. 4(a), the value of the element of gender is “boy”, and the value of the element of mobile phone number is 111-111-111 (there are dashes between the numbers). It can be known from common sense that gender is usually male (sex) or female (sex), and there is no need to have dashes between the numbers of the mobile phone number. This writing method for boy and mobile phone number is not standard compared with common sense. After the scheme of S301-S304, the standardized value of the element of gender is male, and “male” is added to the element of gender with the standardized name in the standard table as the value of gender. Or after the scheme of S301-S304, the standardized value of the element of mobile phone number is 111111111, and “111111111” is added to the element of mobile phone number with the standardized name in the standard table as the value of mobile phone number, as shown in FIG. 4(b).

[0071] It can be understood that if the collected table is a table stored in a medical system or a school business system, the collected table is regarded as a target table, and the number of the collected target tables is multiple. There can be the same target element in the multiple target tables. For example, the gender is included in the multiple collected outpatient medical record tables.

[0072] In an optional embodiment, after obtaining the target data, the method further comprises:

[0073] obtaining a value distribution feature of the same target element in each target table;

[0074] In a case where the value distribution feature of the same target element meets a predetermined distribution feature, obtaining a first label and a second label of the same target element in each target table based on the attribute of the same target element in each target table.

[0075] In the foregoing optional solution, the values of the same target element are obtained from each target table, and the value distribution feature of the same target element is calculated based on the values. If the value distribution feature of the same target element meets a predetermined distribution feature such as a normal distribution or an average distribution, it indicates that the collected values of the same target element are relatively balanced, which is beneficial to data governance. Then, the semantic and value feature of the same target element in each target table are obtained based on the naming information and / or value information of the same target element in each target table.

[0076] If the value distribution feature of the same target element does not meet the predetermined distribution feature, it indicates that the collected values of the same target element are not balanced enough, which is not conducive to data governance. It is necessary to return to S301 to reacquire the target data, so that the values of various types of target elements are as balanced as possible.

[0077] The foregoing solution is considered in view of the fact that data governance needs to govern data with balanced values, such as governing disease data with balanced male and female proportions. Based on this, balanced data such as data with balanced male and female proportions or data with balanced proportions of middle-aged people, old people, and children needs to be collected, so as to avoid that the data collection is not balanced enough and causes a unit in need to be unable to realize normal data analysis.

[0078] As an implementable embodiment, as shown in Figure 5 The foregoing solution of (S302) obtaining the first label of the target element based on the attribute of the target element can be implemented by one of the following:

[0079] S302a comprises:

[0080] Mode one: obtaining the first label of the target element based on the naming information of the target element;

[0081] In the first mode, a semantic recognition method of natural language is used to perform semantic recognition on the naming information of the target element to obtain a first label representing the semantics of the target element. This method of obtaining the first label is simple and easy to implement in engineering.

[0082] In the second mode, a first label of the target element is obtained based on the naming information of the target table in which the target element is located and the naming information of the target element.

[0083] In the second mode, a semantic recognition method of natural language is used to perform semantic recognition on the naming information of the target table and the naming information of the target element to obtain a first label representing the semantics of the target element. In actual applications, the elements appearing in the (target) table are elements related to the naming of the table. For example, in a clinic medical record table, the naming of the table is “clinic medical record table”, and in order to record the information of the patient in detail, elements such as gender and age are essential and should appear in the table. Therefore, the semantic information of the naming of the table can help to identify the semantic information of the elements appearing in the table to some extent. Compared with the scheme of performing semantic recognition on the naming information of the target element in the first mode, the second mode combines the semantic information of the naming of the target table to obtain the semantic information of the target element, so that the semantic information of the element is more accurate, thereby improving the accuracy of data governance.

[0084] As an implementable mode, as shown in Figure 5 The scheme of obtaining the second label of the target element based on the attribute of the target element in the foregoing (S302) can be implemented by the following scheme:

[0085] S302b: obtaining the second label of the target element based on the value information of the target element and the first label.

[0086] In this step, the value characteristics of the target element are identified by combining the semantic information of the target element and the value information of the target element.

[0087] For example, the value of the element of gender has the characteristics of two values of male and female. Based on the semantic information of the target element of gender and the value information such as “male, man, boy, female, woman, girl” of the element, it is identified that the value characteristics are male and female.

[0088] This method of obtaining the second label is simple and easy to implement in engineering. Moreover, the value characteristics of the element are identified by combining the semantic information of the element, which can greatly improve the identification accuracy of the second label by considering the influence of the semantic information of the element on the identification of the value characteristics of the element.

[0089] The target table in the disclosure is the first target table as described above, is a table stored in a medical system or a school business system, or is a table collected from the aforementioned system, and thus the number of tables is multiple. The target element in the technical solution of the disclosure is an element in the target table, and based on this, as an implementable manner, the data processing method of the disclosure further comprises:

[0090] Based on the first label and the second label of the target element of each target table in the multiple target tables, a first preset algorithm is used to determine the target tables having an association relationship from the multiple target tables; wherein the association relationship between the target tables is used for output.

[0091] In actual application, there may be tables having a certain association relationship in the multiple target tables. For example, the multiple target tables include an outpatient medical record table of Zhang San and an inpatient medical record table of Zhang San. The outpatient medical record table of Zhang San and the inpatient medical record table of Zhang San as separate tables can execute the flow as shown in Figure 3 to realize the standardization of the elements in the two tables.

[0092] In addition, the outpatient medical record table of Zhang San and the inpatient medical record table of Zhang San as two different tables of the same patient, Zhang San, can use a first preset algorithm, a social relationship mining algorithm, to mine the different tables of the same patient from the multiple collected target tables based on the element semantics and element value characteristics of the target elements in the two tables. In implementation, the element semantics and element value characteristics of the target elements in the tables of the different tables can be input into a mining model using the social relationship mining algorithm, the mining model mines the target tables having an association relationship according to the input information, and gives a mining result. For example, the target tables having an association relationship or the target elements in the target tables having an association relationship are output. Thus, the target tables having an association relationship can be obtained from the multiple target tables. The mining of the target tables having an association relationship can facilitate the analysis of the illness of the same patient by a relevant unit such as the health commission, and has strong practicability.

[0093] In actual application, the association relationship between the target tables can be output, for example, the target table 1 and the target table 2 are two different tables for the same patient, or when one of the target tables is output, the other target table having an association relationship is also output. The visualization of the standardized data having an association relationship is realized, and the practicability and ease of use of the standardized data are embodied.

[0094] In practical applications, the aforementioned outpatient medical record table, inpatient medical record table, operation medical record table, and student personal information table are used as business tables in a medical system or a school business system. In addition to storing business tables, the medical system or the school business system also stores dictionary tables. The dictionary table can be regarded as a table set for the value or naming of a certain item or element in the business table to save the storage space of the medical system or the school business system.

[0095] Taking the gender element in the business table as an example, if the value of the name element in the business table is written and saved in the system in the form of Chinese characters, male (sex) and female (sex), it will occupy a large storage space. In order to reduce the occupation of storage space, a dictionary table is designed for the gender element, and the dictionary table records that the number “0” represents male and the number “1” represents female. In this way, in the medical system, the gender element in the business table can be stored in the form of numbers instead of Chinese characters, such as writing the number “0” after the gender item in an outpatient medical record table as the value of the gender element. Through the reference of the dictionary table, it can be known that the gender of the patient with the outpatient medical record table is male. That is, by using the dictionary table, the naming of the element in the business table or the simplified storage of the value of the element can be realized, and the occupation of the storage space can be greatly reduced.

[0096] Therefore, there are two types of tables in the medical system or the school business system: business tables and dictionary tables. When collecting data in the medical system or the school business system, the target element collected may be an element in the business table or an element in the dictionary table. How to identify whether the target element collected is an element in the business table or an element in the dictionary table can be realized by using the following scheme.

[0097] Based on the first label and the second label of the target element, a second preset algorithm is used to determine the type of the target table.

[0098] Here, the second preset algorithm is any reasonable AI model algorithm, such as a binary classification algorithm, a logistic regression algorithm, a decision tree algorithm, etc. In implementation, the first label and the second label of the target element collected can be input into an AI model using at least one of the aforementioned algorithms. The AI model analyzes the semantic and value characteristics of the target element to determine whether the target element comes from the business table or the dictionary table, and gives the analysis result. In this way, the identification of whether the target element is an element in the business table or an element in the dictionary table is realized.

[0099] Since the AI model has strong robustness and robustness, the accurate identification of the type of the target table can be realized. Therefore, the accuracy of data governance can be provided.

[0100] It can be understood that the application value of the business table is greater than that of the stored dictionary table, and the application significance of standardizing the elements in the business table is stronger than that of standardizing the elements in the dictionary table.

[0101] In actual application, if it is determined that the type of the target table is the first type table, i.e., the business table, the target element is standardized based on the data category of the target element, and the standardized data of the target element is obtained. If it is determined that the type of the target table is the second type table, i.e., the dictionary table, no standardization is performed or the data category of the target element does not need to be determined.

[0102] Of course, the target element can be standardized regardless of whether the type of the target table is the first type or the second type. Further, the standardization of the elements in different types of tables is realized, and the standardization is more comprehensive.

[0103] As an optional manner, the foregoing (S304) can be: the target element is standardized according to the target processing mode corresponding to the data category of the target element, and the standardized data of the target element is obtained.

[0104] In which, a corresponding target processing mode is set for each data category, and the target processing mode is used to realize the standardization of the data of the category. For example, the target processing mode set for the name data category is to process the element named with "name" or "surname" into the element named with "name". Based on the processing mode, the element named with "name" or "surname" in the (first) target table is processed into the element named with "name", and the standardized name of the element is obtained.

[0105] This standardization method is simple, easy to implement, and practical. It effectively improves the data management efficiency.

[0106] As an implementable manner, the standardization of the target element based on the data category of the target element includes:

[0107] The target element of the target table is standardized based on the data category of the target element and the type of the target table, and the standardized data of the target element is obtained.

[0108] This case is aimed at the same element appearing in the business table and the dictionary table. For example, the gender element, which is defined in the dictionary table as number 0 representing male and number 1 representing female. In the business table, the gender is represented by number 0 or number 1. The types of target tables are different, and different standardization processing is performed on the data of the same data category in different types of target tables to obtain standardized data. Based on the data category and the type of the target table, the standardization of the target element is realized, and the respective standardization of the same element in different tables is realized, so that the standardization is more comprehensive.

[0109] Exemplarily, if the target table is a business table, the target processing mode adopted for the gender data category in the business table is to restore the gender from digital representation to Chinese character representation according to the provisions of the dictionary table. According to this target processing mode, the value of the gender element with a value of "0" in the target table can be processed as "male".

[0110] If the target table is a dictionary table, no standardization processing is performed on the target elements in the dictionary table. Alternatively, the target elements in the dictionary table are standardized. For example, for the gender data category in the dictionary table, the target processing mode adopted is to transform the gender representation of number 0 and number 1, such as transforming the original number 0 representing male to number 0 representing female, and transforming the original number 1 representing female to number 1 representing male. According to this target processing mode, it can be known that the gender represented by number 0 and 1 in the dictionary table is interchanged. Subsequently, the medical treatment system or the school business system records the gender according to the provisions of the dictionary table.

[0111] As shown in Figure 6 As an implementable manner, the data processing method of the present disclosure further comprises:

[0112] S305: Output the standardized data of the target element.

[0113] For example, the name target element named by "name" or "surname" in the first target table is output as "name" standardized data, so as to realize the output of the standardized naming of the name target element.

[0114] This standardization output of data realizes the visualization of the standardized data, and embodies the practicability and ease of use of the standardized data. It can also greatly facilitate the viewing of relevant personnel such as the health commission or the education commission personnel, and improve the user experience.

[0115] The technical solutions of the embodiments of the present disclosure will be described in detail below. Figure 7- Figure 8 The technical solutions of the embodiments of the present disclosure will be described in detail below.

[0116] As shown in Figure 7As shown, the system implementing the data processing method of the embodiments of the present disclosure includes a collection subsystem, a processing subsystem and a display subsystem. The collection subsystem is configured to collect target data from a medical system or a school business system. The analysis subsystem is configured to process the target data collected by the collection subsystem, mainly to process the naming of the collected non-standard elements and / or the values of the elements into standardized naming and / or standard values, i.e., mainly to implement the processing of data from non-standardization to standardization. The display subsystem is configured to display the processing result of standardization.

[0117] The above several subsystems will be described in detail below.

[0118] Taking the outpatient medical record table in the medical system as an example, the collection subsystem can collect data once when the collection period comes or when the input operation of the user requiring the collection subsystem to collect data is detected.

[0119] Considering that there are various tables in the medical system, such as outpatient medical record tables, inpatient medical record tables, and surgical medical record tables, the type of table collected each time and the number of each type of table can be set according to actual conditions.

[0120] For example, L outpatient medical record tables are collected in the current collection, and the name, gender, age, mobile phone number, medical history, and current illness of the outpatient medical record table are taken as target elements. The naming information of various target elements and the value information of the target elements existing in the L outpatient medical record tables are read.

[0121] For example, the naming information and value information of L elements about the name are read, the naming information and value information of L elements about the gender are read, and the naming information and value information of L elements about the age are read.

[0122] Taking the name as an example, the naming of the element in the outpatient medical record table can be "name", "surname", or "name", and the value information can be Zhang San, Li Si, or Wang Wu.

[0123] The naming of the gender in the outpatient medical record table can be "gender", "Sexy", or "Sex" (English for gender). Because of the existence of the dictionary table, the value information can be a number 0 or a number 1.

[0124] As can be seen, the naming information of the collected target elements such as the name and the gender is diverse and not unified.

[0125] It can be understood that the number of naming information and value information of the name is L, and the number of naming information and value information of the gender is L. If only the name and the gender are taken as the target elements, only the L naming information and value information about the name can be standardized, and the L naming information and value information about the gender can be standardized.

[0126] As will be described below, the L naming information about the name is standardized as "name", and the L naming information about the gender is standardized as "gender". In this way, the data items originally written as "surname" or "name" in the collected outpatient medical record table can be unified as "name", and the data items originally written as "Sexy" or "Sex" can be unified as "gender". In this way, the standardization of elements such as name and gender can effectively improve the data management efficiency.

[0127] As shown in the following table, the processing subsystem mainly includes the following modules: Figure 8

[0128] 1) Exploration module

[0129] It is assumed that L tables are collected in the current time, and a specific element such as gender or age in the L tables is taken as an example. In the case of reading the value of the gender in each table in the L tables, the distribution characteristics of the value of the gender are calculated by using statistical methods. If the distribution characteristics conform to the normal distribution or the average distribution, it means that the male and female gender ratio collected in the current time is relatively balanced, and the value information of each target element of the L tables read can be input to the subsequent module. If the distribution characteristics do not conform to the normal distribution or the average distribution, it means that the male and female gender ratio collected in the current time is not balanced, and the acquisition subsystem is triggered to re-collect the target data until the distribution characteristics of the value of the gender conform to the predetermined distribution characteristics.

[0130] In some application scenarios, data of each age group also needs to be collected. Referring to the foregoing description of whether the value of the gender is balanced, the distribution characteristics of the value of the age as the target element can be calculated to identify whether the age in the collected target data is balanced.

[0131] 2) Semantic restoration (or identification) module

[0132] The semantic recognition method of natural language is used to perform semantic recognition on the naming information of the target elements such as the name and the gender in the L tables, or the naming information of the target elements in combination with the naming information of the table in which the target elements are located, such as "outpatient medical record table", to obtain the semantics of the target elements. In this way, it can be known that the "name" and "surname" in the table refer to the name, and the "Sexy" and "Sex" in the table refer to the gender.

[0133] ​In practical application, the data in the table can be written in English or in Chinese pinyin, such as "NL (the first letter of the Chinese pinyin of the age)". In this case, the naming information written in English and / or Chinese pinyin is recognized in terms of semantics to facilitate subsequent standardization.

[0134] 3) Feature analysis module

[0135] For each element in the table, the value feature of the element is identified in combination with the semantic of the element and the value information of the element.

[0136] For example, the value of the gender element in the L tables is 0 or 1, which is two values. The semantic of the element taking these two values is gender. Therefore, the element with the semantic of gender has the value feature of taking 0 and 1.

[0137] For example, for the element of medical history, the value of the element in the L tables can be the past medical history of the L patients. Based on the analysis of the value of the element, it is known that the value of the element with the semantic of medical history is multiple diseases.

[0138] For example, for the element of age, the value of the element in the L tables can be the age of the L patients. Based on the analysis of the value of the element, it is known that the value of the element with the semantic of age is multiple numerical ages.

[0139] Therefore, the value feature of each element in the table can be obtained.

[0140] 4) Recognition module

[0141] The semantic of the element and the value feature are combined, and AI is used to identify the data category.

[0142] For example, the semantic of the element with the semantic of gender and the value feature of taking 0 and 1 are input into an AI model such as a logistic regression model, and the AI model identifies that the data category of the element is the category of gender.

[0143] For example, the semantic of the element with the semantic of age and the value feature of taking multiple numerical ages are input into an AI model such as a logistic regression model, and the AI model identifies that the data category of the element is the category of age.

[0144] The semantic of the element and the value feature are combined, and AI is used to identify the type of the table in which the element is located.

[0145] In actual application, there can be a dictionary table and a business table in each collected table. The semantics and value characteristics of the target elements read from the table 1 are input to an AI model such as a binary classification model, and the binary classification model identifies the result of whether the table 1 is a dictionary table or a business table based on the input information and outputs the result. Thus, the identification of whether each collected table is a dictionary table or a business table can be realized.

[0146] 5) Mining module

[0147] The semantics and value characteristics of the target elements in each of the L tables can be input to a mining model using a social relationship mining algorithm. The mining model mines the tables having a correlation relationship according to the input information, and gives a mining result.

[0148] For example, the table 1 and the table 2 are tables having a correlation relationship, and are both tables for the same patient. For example, the table 1 is an outpatient medical record table for the year 2021, and the table 2 is an outpatient medical record table for the year 2022.

[0149] The principle of the mining model to mine the tables having a correlation relationship is mainly that in different tables having a correlation relationship, the values of two or more elements are the same. For example, in two different tables for the same patient, the values of the elements such as name, gender, age, date of birth, and mobile phone number are the same. The mining model identifies which tables among the L tables have a correlation relationship through the semantics and value characteristics of the elements.

[0150] The aforementioned correlation relationship can be any reasonable correlation relationship, such as mining tables of patients having a kinship relationship based on patient family information in the table, inpatient medical record tables and surgical medical record tables of the same patient, and the like.

[0151] 6) Standardization module

[0152] The standardization processing of the data of each data category is realized according to the target processing mode set for the category.

[0153] For example, the name element in the table named with “name” or “surname” is processed to be named with “name”, and the standardized naming of the name element is obtained. The gender element in the table named with “Sexy” is processed to be named with “gender”, and the standardized naming of the gender element is obtained.

[0154] Alternatively, according to the provisions of the dictionary table, the element values of the numbers 0 and 1 used for the gender element in the table are transformed into the element values of male and female.

[0155] Thus, the standardized naming and standardized value of the data in the table can be realized.

[0156] It is understandable that data standardization can be performed on different categories based on the table type. If the table is a business table, then standardization should be performed according to the aforementioned scheme. If the table is a dictionary table, standardization may or may not be performed. For the standardization process of dictionary tables, please refer to the aforementioned related explanations.

[0157] In practical applications, different types of tables, such as dictionary tables and business tables, may contain the same elements. Standardizing these identical elements based on the table type can effectively avoid data standardization errors.

[0158] The aforementioned processing subsystem implements the process of transforming non-standardized data in the table into standardized data. That is, the processing subsystem transforms non-standardized data in the table into standardized data, providing an automated standardization solution. Compared with manual data processing solutions in related technologies, this approach frees up manpower and resources, improving processing efficiency.

[0159] Furthermore, in the aforementioned scheme, the processing subsystem combines element semantics and value features, employing AI to identify data categories, mine table relationships, and identify table types. The combination of these two types of labels (element semantics and value features) and the use of AI algorithms significantly improves accuracy, thereby achieving accurate data governance.

[0160] The above processing procedure of the processing subsystem (including the processing procedure of each module of the processing subsystem) can be performed on a table-by-table basis, processing each target element in each table one by one. Alternatively, it can be performed on an element-by-element basis, grouping identical elements in each table into a group to be processed, and performing the above processing on each group to be processed, which can speed up the processing efficiency and thus improve the efficiency of data governance.

[0161] The processing subsystem can use the identified data categories, mined table relationships, identified table types, and processed standardized data as recommended content and then recommend them to the display subsystem for output.

[0162] like Figure 8 As shown, the demonstration subsystem includes the following modules:

[0163] 1) Table Display Module: This module displays tables that have standardized the naming and / or values ​​of target elements. For example, it displays a standardized outpatient medical record table where the name element is named "Name" and the gender element is named "Gender" with a value of "Male," all of which are standardized data. Data named "name" or "Sex" does not exist.

[0164] If a standardized surgery medical record form of a patient is used as a master form, a standardized outpatient medical record form of the patient is used as a slave form, and the processing subsystem recommends slave forms for the master form, the number of recommended slave forms can be one or more. If the form display module can output only one or a small number of slave forms each time, if it is found through manual verification that the slave form is incorrect and cannot be used as a slave form of the current output master form, the correct slave form can be selected from other recommended slave forms and output manually.

[0165] The master form can be determined in combination with the acquisition task input by the user. If the acquisition task input by the user is to acquire an outpatient medical record form, the outpatient medical record form is used as the master form. Other standardized forms such as a surgery medical record form are used as slave forms.

[0166] 2) Element display module: one of the naming information and the value information of the standardized element can be displayed. That is, the entire form is not displayed, and only the naming information and the value information of the standardized element in the form are displayed.

[0167] For example, the naming information and the value information (male or female) of the standardized element of gender in the L forms are displayed.

[0168] 3) Form relationship display module: different forms or data in the form having a correlation relationship can be displayed. The data in the form can be standardized data or non-standardized data, and is preferably standardized data.

[0169] Alternatively, when a form or data in the form is output, the ID (identification) of the form having a correlation relationship with the form, such as the naming information of the form file, is also output.

[0170] 4) Dictionary form display module: when it is identified that the type of the form is a dictionary form, the dictionary form can be displayed. During the display of the dictionary form, the dictionary form can be changed manually, such as the rules described in the dictionary form. Alternatively, when the master form is output, the dictionary form is output as a slave form, and if it is found through manual verification that the output dictionary form is not the dictionary form of the current output master form, the correct dictionary form can be selected manually from the multiple dictionary forms recommended by the display subsystem.

[0171] It can be understood that the foregoing scheme is a scheme for automatically implementing data governance, and is used to improve the accuracy of governance. At the same time, the data can be visualized. The contents output by the above display modules can be further verified manually, and if errors are found, the correct contents can be selected from the recommended contents and displayed.

[0172] In practical applications, after the naming information and the value information of the target elements in a collected table are standardized, the related information (the naming information and the value information) of the target elements after the standardization is collected to obtain a new table, and the naming information and the value information of each element recorded in the table are standardized data.

[0173] In practical applications, a standard table about the outpatient medical record can be designed in advance, and each target element in the standard table is named by using standard naming information, so that only the value of the element written by using non-standard data in the collected table needs to be processed from non-standard to standard, and added to the back of the corresponding item in the standard table. For example, the data originally taking the value 0 of the gender element is processed as male, and the two words male are added to the back of the gender item in the standard table.

[0174] The technical solution of the present disclosure replaces the artificial governance solution, greatly shortens the artificial governance time, improves the data governance efficiency, and at the same time makes the data more intuitive. In addition, the technical solution of the present disclosure provides a modification opportunity for artificial based on the intuitive display of data, which can further ensure the accuracy of data governance.

[0175] The embodiment of the present disclosure provides a data processing device, as shown in Figure 9 The device comprises:

[0176] The first acquisition unit 901 is configured to acquire target data, and the target data comprises at least one target element;

[0177] The second acquisition unit 902 is configured to obtain a first label and a second label of the target element based on the attribute of the target element, wherein the first label represents the semantic of the target element, and the second label represents the value feature of the target element;

[0178] The first determination unit 903 is configured to determine the data category of the target element by using an artificial intelligence method based on the first label and the second label of the target element;

[0179] The third acquisition unit 904 is configured to perform standardization processing on the target element based on the data category of the target element to obtain the standardized data of the target element.

[0180] In an optional solution, the second acquisition unit 902 is configured to

[0181] obtain the first label of the target element based on the naming information of the target element;

[0182] Alternatively,

[0183] The first label of the target element is obtained based on naming information of a target table in which the target element is located and naming information of the target element.

[0184] In an optional implementation, the second obtaining unit 902 is configured to obtain a second label of the target element based on the value information of the target element and the first label.

[0185] In an optional implementation, the target element is an element in a target table, and the number of the target tables is a plurality, and the apparatus further includes a second determining unit configured to

[0186] The target tables having the association relationship are determined from the plurality of target tables based on the first label and the second label of the target element in each of the plurality of target tables by using a first preset algorithm, and the association relationship between the target tables is used for output.

[0187] In an optional implementation, the target element is an element in a target table, and the apparatus further includes a third determining unit configured to

[0188] The type of the target table is determined based on the first label and the second label of the target element by using a second preset algorithm.

[0189] In an optional implementation, the third obtaining unit 904 is configured to

[0190] The target element in the target table is standardized based on the data category of the target element and the type of the target table, to obtain standardized data of the target element.

[0191] In an optional implementation, the third obtaining unit 904 is configured to

[0192] The target element is standardized according to a target processing manner corresponding to the data category of the target element, to obtain standardized data of the target element.

[0193] In an optional implementation, the target element is an element in a target table, and the number of the target tables is a plurality, and the first obtaining unit 901, after obtaining the target data, is further configured to

[0194] Obtain value distribution characteristics of the same target element in each target table.

[0195] In a case where the value distribution characteristics of the same target element meet predetermined distribution characteristics, obtain the first label and the second label of the same target element in each target table based on attributes of the same target element in each target table.

[0196] In an optional implementation, the apparatus further includes an output unit configured to output the standardized data of the target element.

[0197] The functions of each component unit in the data processing apparatus of the embodiments of the present disclosure can be referred to the description of the data processing method, which will not be repeated here. The data processing apparatus of the embodiments of the present disclosure, since the solving principle is similar to the foregoing data processing method, therefore, the implementation process and implementation principle, beneficial effects of the data processing apparatus can be referred to the foregoing related method implementation process and implementation principle, beneficial effects description, and the repeated part will not be repeated.

[0198] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, comprising at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the foregoing business configuration parameter obtaining method.

[0199] The description of the processor and the memory of the electronic device can be referred to the related description of the computing unit 1001 and the storage unit 1008 in Figure 10 .

[0200] According to the embodiments of the present disclosure, the present disclosure also provides a non-transitory computer readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the foregoing traffic control method and the training method of the traffic control model. The description of the computer readable storage medium can be referred to the related description in Figure 10 .

[0201] According to the embodiments of the present disclosure, the present disclosure also provides a computer program product, comprising a computer program which, when executed by a processor, implements the foregoing business configuration parameter obtaining. The description of the computer program product can be referred to the related description in Figure 10 .

[0202] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0203] Figure 10is a block diagram of an electronic device that implements the data processing apparatus of the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit the implementations of the present disclosure described and / or claimed in this document.

[0204] As shown in Figure 10 The electronic device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes in accordance with a computer program stored in a ROM 1002 or a computer program loaded into a RAM 1003 from a storage unit 1008. Various programs and data required for the operation of the electronic device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0205] Various components in the electronic device 1000 are connected to the I / O interface 1005, including an input unit 1006, such as a keyboard, a mouse, etc., an output unit 1007, such as various types of displays, speakers, etc., a storage unit 1008, such as a magnetic disk, an optical disk, etc., any device that can be used as a memory, and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0206] The storage unit 1008 in the embodiments of the present disclosure can be embodied as at least one memory among a portable computer disk, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM), or a flash memory, an optical fiber, a CD-ROM, an optical storage device, a magnetic storage device.

[0207] The computing unit 1001 can be various general purpose and / or special purpose processing components having processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, CPUs, graphics processing units (GPUs), artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, or processor of any kind. The computing unit 1001 performs various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded onto the RAM 1003 and executed by the computing unit 1001, one or more steps of the data processing method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the data processing method by any other suitable means, such as by means of firmware.

[0208] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0209] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0210] In the context of this disclosure, a machine-readable medium (storage medium) can be any tangible medium that can contain or store program code for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM or Flash memory, an optical fiber, a CD-ROM, an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0211] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0212] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0213] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0214] It should be understood that the various forms of flow shown above can be used to reorder, add, or remove steps. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.

[0215] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A data processing method, comprising: Acquire target data, wherein the target data includes at least one target element; Based on the attributes of the target element, a first label and a second label of the target element are obtained, wherein the first label represents the semantics of the target element and the second label represents the value characteristics of the target element. Based on the first and second tags of the target element, the data category of the target element is determined using artificial intelligence. Based on the data category of the target element, the target element is standardized to obtain standardized data of the target element; The step of obtaining the first tag of the target element based on its attributes includes: Based on the naming information of the target element, the first tag of the target element is obtained; or, Based on the naming information of the target table where the target element is located and the naming information of the target element, the first tag of the target element is obtained; The step of obtaining the second tag of the target element based on its attributes includes: Based on the value information of the target element and the first tag, the second tag of the target element is obtained.

2. The method according to claim 1, wherein the target element is an element in a plurality of target tables, and the method further comprises: Based on the first and second labels of the target elements in each of the multiple target tables, a first preset algorithm is used to determine the target tables with related relationships from the multiple target tables; The relationships between the target tables are used for output.

3. The method according to claim 1, wherein the target element is an element in a target table, and the method further comprises: Based on the first and second tags of the target element, the type of the target table is determined using a second preset algorithm.

4. The method according to claim 3, wherein, The standardization process based on the data category of the target element to obtain standardized data of the target element includes: Based on the data category of the target element and the type of the target table, the target elements of the target table are standardized to obtain standardized data of the target elements.

5. The method according to any one of claims 1 to 2, wherein, The standardization process based on the data category of the target element to obtain standardized data of the target element includes: According to the target processing method corresponding to the data category of the target element, the target element is standardized to obtain the standardized data of the target element.

6. The method according to claim 1, wherein, The target element is an element from multiple target tables. After obtaining the target data, the method further includes: Obtain the value distribution characteristics of the same target elements in each target table; When the value distribution characteristics of the same target element meet the predetermined distribution characteristics, the first label and the second label of the same target element in each target table are obtained based on the attributes of the same target element in each target table.

7. The method according to claim 1, further comprising: Output the standardized data of the target element.

8. A data processing apparatus, comprising: The first acquisition unit is used to acquire target data, wherein the target data includes at least one target element; The second acquisition unit is used to obtain a first tag and a second tag of the target element based on the attributes of the target element, wherein the first tag represents the semantics of the target element and the second tag represents the value characteristics of the target element. The first determining unit is used to determine the data category of the target element based on the first label and the second label of the target element using artificial intelligence. The third acquisition unit is used to perform standardization processing on the target element based on the data category of the target element to obtain standardized data of the target element; The second acquisition unit is used for: Based on the naming information of the target element, the first tag of the target element is obtained; or, Based on the naming information of the target table where the target element is located and the naming information of the target element, the first tag of the target element is obtained; The second acquisition unit is used for: Based on the value information of the target element and the first tag, the second tag of the target element is obtained.

9. The apparatus according to claim 8, wherein, The target element is an element from multiple target tables, and the device further includes: The second determining unit is used to determine, based on the first and second labels of the target elements in each of the plurality of target tables, a first preset algorithm is used to determine the target tables with related relationships from the plurality of target tables; wherein the relationship between the target tables is used for output.

10. The apparatus according to claim 8, wherein, The target element is an element in the target table, and the device further includes: The third determining unit is used to determine the type of the target table based on the first and second tags of the target element using a second preset algorithm.

11. The apparatus according to claim 10, wherein, The third acquisition unit is used for: Based on the data category of the target element and the type of the target table, the target elements of the target table are standardized to obtain standardized data of the target elements.

12. The apparatus according to any one of claims 8 to 9, wherein, The third acquisition unit is used for: According to the target processing method corresponding to the data category of the target element, the target element is standardized to obtain the standardized data of the target element.

13. The apparatus according to claim 8, wherein, The target element is an element from multiple target tables, and the first acquisition unit is further configured to: After obtaining the target data, Obtain the value distribution characteristics of the same target elements in each target table; When the value distribution characteristics of the same target element meet the predetermined distribution characteristics, the first label and the second label of the same target element in each target table are obtained based on the attributes of the same target element in each target table.

14. The apparatus according to claim 8 further includes an output unit for outputting standardized data of the target element.

15. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.

16. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method of any one of claims 1-7.

17. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • A medical record document standardize processing system and method

    CN109408635A

  • Dependency relationship recognition method and device based on data table and computer equipment

    CN110889286A