Multi-source heterogeneous master data mapping method and system for medical material supply chain
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]本发明提供了一种面向医疗物资供应链的多源异构主数据映射方法及系统,以解决现有技术中存在医疗物资多源异构主数据难以准确生成可追溯映射关系的问题
(1)本发明通过对采购系统和仓储系统的物资主数据记录进行字段结构提取、异构字段筛选和复合属性识别,能够从多源字段中定位包含物资名称、规格、批号、效期、单位等属性的物资复合字段,避免仅依赖固定字段名称或固定字段位置进行匹配而造成字段识别错误,使后续主数据映射建立在更准确的异构字段识别结果之上。
Smart Images

Figure CN122547754A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical supplies data processing technology, and in particular to a method and system for mapping multi-source heterogeneous master data for the medical supplies supply chain. Background Technology
[0002] Currently, the medical supply chain continuously generates a large amount of master data during the processes of material circulation, acceptance and warehousing, inventory management, and system integration. With the expansion of business scale and the increasing demand for system collaboration, fields such as material name, specifications, batch number, expiration date, unit of measurement, and supplier identifier are gradually exhibiting characteristics of multiple sources, multiple formats, and multiple levels, making medical supply master data processing a typical big data management process. The ability to accurately identify, uniformly represent, and stably convert master data across different business systems directly affects the reliability of subsequent warehousing verification, inventory traceability, and cross-system data retrieval.
[0003] In existing technologies, medical supply master data mapping typically relies on manually maintained field mapping tables or fixed rule templates. The system matches supply fields in the procurement system with those in the warehousing system based on field names or fixed field locations, and completes basic data alignment through preset format conversion rules. However, different systems record the same medical supplies in inconsistent ways. The procurement system might combine supply name, specifications, batch number, expiration date, or supplier information into a single long text field, while the warehousing system might split the same attributes into multiple independent fields. Especially in scenarios with differences in unit descriptions and inconsistent field naming, fixed field rules struggle to accurately identify the true attribute boundaries within composite fields and fail to retain the source relationship between attribute fragments and original fields, resulting in unstable traceability of the field extraction process in subsequent mapping relationships.
[0004] Existing technologies present the problem that it is difficult to accurately generate traceable mapping relationships from multi-source heterogeneous master data of medical supplies. Summary of the Invention
[0005] This invention provides a method and system for mapping multi-source heterogeneous master data for the medical supplies supply chain, in order to solve the problem in the prior art that it is difficult to accurately generate traceable mapping relationships from multi-source heterogeneous master data of medical supplies.
[0006] Firstly, to address the aforementioned technical problems, this invention provides a method for mapping multi-source heterogeneous master data in the medical supplies supply chain, comprising: Obtain the master data records of materials from the procurement system and warehousing system in the medical supplies supply chain, extract the field structure of the master data records of materials, and obtain the material field features; Heterogeneous fields are filtered based on the characteristics of the material fields to obtain a set of candidate heterogeneous fields. Composite attributes are then identified from the set of candidate heterogeneous fields to obtain composite material fields. The composite material field is segmented and located to obtain field segmentation points, and the composite material field is sliced according to the field segmentation points to obtain material attribute fragments with field source paths and slice location information. Based on the fragment content of the material attribute fragment, the slice positioning information, and the relationship between adjacent fragments, attribute attribution is marked to obtain structural attribute fragments, and cross-system format alignment is performed on the structural attribute fragments to obtain standard attribute features; Based on the standard attribute features, a standard mapping association is performed to obtain a mapping relationship, which includes the field source path, slice location information, and target field path.
[0007] Secondly, the present invention provides a multi-source heterogeneous master data mapping system for the medical supplies supply chain, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described above.
[0008] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described above.
[0009] Compared with the prior art, the present invention has the following beneficial effects: (1) By extracting field structure, filtering heterogeneous fields and identifying composite attributes from the master data records of the procurement system and the warehousing system, this invention can locate composite fields of materials containing attributes such as material name, specifications, batch number, expiration date and unit from multiple source fields, avoiding field identification errors caused by relying solely on fixed field names or fixed field positions for matching, and enabling subsequent master data mapping to be based on more accurate heterogeneous field identification results.
[0010] (2) This invention divides and locates the composite field of materials and slices the attribute, and retains the source path of the field and the location information of the slice during the attribute slicing process. This enables the material attribute fragments to not only have attribute value content, but also to correspond to the specific source location in the original field, thereby providing traceable data basis for subsequent standard mapping association and avoiding the problem that the mapping relationship cannot reproduce the field extraction process.
[0011] (3) By performing cross-system format alignment on structural attribute fragments and performing standard mapping association based on standard attribute features, this invention can convert medical supplies attributes that are inconsistent in different systems into a unified mapping relationship, so that the mapping relationship carries the field source path, slice location information and target field path, thereby enabling multi-source heterogeneous material master data to be accurately mapped according to a unified relationship. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the multi-source heterogeneous master data mapping method for the medical supplies supply chain provided in the first embodiment of the present invention. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] Reference Figure 1 The first embodiment of the present invention provides a method for mapping multi-source heterogeneous master data for the medical supplies supply chain, including the following steps: S11, Obtain the master data records of materials from the procurement system and warehousing system in the medical supplies supply chain, extract the field structure of the master data records of materials, and obtain the material field features; S12, perform heterogeneous field filtering based on the material field characteristics to obtain a candidate heterogeneous field set, and perform composite attribute identification on the candidate heterogeneous field set to obtain the material composite field; S13, the composite material field is segmented and located to obtain the field segmentation point, and the composite material field is sliced according to the field segmentation point to obtain the material attribute fragment with the field source path and slice location information. S14. Attribute attribution is marked according to the fragment content of the material attribute fragment, the slice positioning information and the relationship between adjacent fragments to obtain structural attribute fragments, and cross-system format alignment is performed on the structural attribute fragments to obtain standard attribute features. S15, perform standard mapping association based on the standard attribute features to obtain a mapping relationship, which includes field source path, slice location information and target field path.
[0015] In step S11, the master data records of the procurement system and warehousing system in the medical supplies supply chain are obtained, and the field structure of the master data records is extracted to obtain the field features of the supplies.
[0016] Specifically, the field structure of the material master data record is extracted to obtain material field features, including: Parse the field names, field values, field source paths, and source system identifiers in the material master data records to obtain basic field information; Character type statistics are performed on the basic information of the fields to obtain character structure features; Field path features are obtained by performing field-level identification on the basic information of the fields. By combining the character structure features and the field path features, the material field features are obtained.
[0017] In one implementation, this embodiment reads material master data records through the data interface between the procurement system and the warehousing system. The material master data record is a set of fields corresponding to a single piece of material business data. Each material master data record contains at least one field name and a corresponding field value. The field name is a field identifier that records the meaning of the field in the business system; the field value is the actual recorded content corresponding to the field name; the field source path is the hierarchical position of the field in the source system; and the source system identifier is a data source marker that distinguishes between the procurement system and the warehousing system.
[0018] It should be noted that this embodiment generates a field source path simultaneously when reading the material master data record. Specifically, this embodiment extracts the data object name, record identifier, and field name of the material master data record, and combines the source system identifier, the data object name, the record identifier, and the field name in hierarchical order to obtain the field source path. The record identifier is a business record number or data row number that distinguishes different material master data records. The field source path indicates the specific data object and record location from which the field originates in the procurement system or warehousing system.
[0019] It should be noted that the basic information of a field consists of a data unit formed by the field name, field value, field source path, and source system identifier. The basic information is stored using a key-value structure, with each of the field name, field value, field source path, and source system identifier occupying an independent key. If the field value is empty, this embodiment writes a null value marker into the basic information and sets the field value length to zero. If the field value is not empty, this embodiment preserves the original character order of the field value and does not truncate it.
[0020] It is worth noting that when performing character type statistics on the basic information of the field, this embodiment reads each character in the field value and classifies them according to Chinese characters, English letters, Arabic numerals, date separators, specification separators, unit characters, and other visible symbols. Date separators include hyphens, forward slashes, and decimal points. Specification separators include asterisks, multiplication signs, and spaces. Unit characters are identified according to a preset unit character table. This embodiment counts the number of each type of character and divides the number of each type of character by the total number of characters in the field value to obtain the corresponding character percentage. Combining the total number of characters in the field value, the number of each type of character, and the percentage of each type of character yields the character structure features.
[0021] It should be noted that the preset unit character table is generated using a database construction method. Specifically, this embodiment extracts fields whose names contain unit, specification, packaging, or measurement text from historical material master data records. Chinese unit fragments and English letter unit fragments in the corresponding field values are used as candidate unit texts. The frequency of each candidate unit text in the historical material master data records is counted. Candidate unit texts with a frequency lower than a preset frequency limit are deleted. The remaining candidate unit texts are written into the unit character table. The historical material master data records refer to all material master data records generated by the procurement system and warehousing system within the 12 months prior to the current processing time. The preset frequency limit is set to 50 times, and its value is determined by extracting all candidate unit texts from the historical material master data records, counting the frequency of each candidate unit text, deleting candidate unit texts with a frequency of 1, and selecting the frequency value corresponding to the 95th percentile from the frequency distribution of the remaining candidate unit texts. The historical material master data records have the same data structure as the material master data records in this step.
[0022] It is worth noting that when identifying the field hierarchy for the basic field information, this embodiment parses the source system identifier, data object name, record identifier, and field name in the field source path to determine the data hierarchy to which the field belongs. If the data object name in the field source path corresponds to a purchase order object, this embodiment marks the business hierarchy in the field path feature as a purchasing-side field. If the data object name in the field source path corresponds to an inbound acceptance object or an inventory ledger object, this embodiment marks the business hierarchy in the field path feature as a warehousing-side field. The field path feature includes the source system identifier, data object name, record identifier, field name, and business hierarchy marker.
[0023] In this embodiment, the field source path is composed of the source system identifier, data object name, record identifier, and field name in hierarchical order, which is used to uniquely identify the complete position of the field in the source system; the slice positioning information consists of the slice start character number and the slice end character number, which is used to indicate the character position range of the attribute fragment in the material composite field value; the original field name is the field name of the material composite field to which the slice belongs in the source system.
[0024] In step S12, heterogeneous fields are filtered according to the characteristics of the material fields to obtain a set of candidate heterogeneous fields. Composite attributes are identified in the set of candidate heterogeneous fields to obtain composite material fields.
[0025] Specifically, heterogeneous fields are filtered based on the characteristics of the material fields to obtain a candidate set of heterogeneous fields, including: Based on the material field characteristics and the preset field synonym list, the fields are semantically categorized to obtain the field categorization results; Based on the field classification results, fields from different source systems but with the same semantic category are filtered to obtain heterogeneous fields of the same type; Based on the aforementioned heterogeneous fields, a set of candidate heterogeneous fields is obtained by performing field integrity filtering.
[0026] In one implementation, the material field features are the field-level data description results obtained in step S11. These material field features include character structure features and field path features. The character structure features reflect the distribution of Chinese characters, English letters, Arabic numerals, date separators, specification separators, unit characters, and other visible symbols in the field value. The field path features reflect the source system identifier, data object name, record identifier, field name, and business level marker to which the field belongs. This embodiment simultaneously reads the character structure features and field path features to obtain the content format and business location of each field.
[0027] It should be noted that the preset field synonym table is generated using a database construction method. Specifically, this embodiment extracts field names and field values from historical material master data records. First, it performs character conversion (full-width to half-width), space deletion, connector deletion, and case unification on the field names to obtain standardized field names. Then, it establishes a target field directory based on the field definitions in the standard material master data table. The field definitions include field codes, Chinese field names, and field value formats. The target field directory includes seven semantic categories: material name, specifications, batch number, expiration date, unit, supplier, and compound description. This embodiment categorizes standardized field names into these seven semantic categories. If the number of times the same standardized field name appears in historical material master data records reaches the minimum writing limit, and the character structure characteristics of the corresponding field value are consistent with the value format of the same semantic category, then this embodiment writes the standardized field name, semantic category, and value format into the preset field synonym table. The write limit is set to 100 times. Its value is determined by extracting all normalized field names from the historical material master data record, counting the occurrences of each normalized field name, deleting the normalized field names that occur 1 time, and selecting the 90th percentile value from the distribution of the occurrences of the remaining normalized field names.
[0028] It is worth noting that when performing semantic classification of fields, this embodiment reads the field names in the material field features, performs the same normalization operation on the field names as the preset field synonym list, and obtains the field names to be classified; then, the field names to be classified are precisely matched with the normalized field names in the preset field synonym list to obtain the corresponding semantic category. If the field name to be categorized does not match the preset field synonym list, this embodiment reads the character structure features in the material field features and performs auxiliary categorization based on the proportion of date separators, specification separators, unit characters, and field value length in the field value. If the proportion of date separators is greater than 0.15 and the field value length is between 6 and 20, it is assisted in categorizing into the date category; if the proportion of specification separators is greater than 0.1 or the proportion of unit characters is greater than 0.2, it is assisted in categorizing into the specification category or the unit category. The values 0.15, 0.1, 0.2, and 6 to 20 are determined by statistically analyzing the distribution of corresponding feature values of each semantic category field in historical material master data records, selecting the interval boundary from the 90th percentile to the 99th percentile of each feature value distribution. The auxiliary categorization result and the field source path together constitute the field semantic record, and all field semantic records constitute the field categorization result.
[0029] It should be noted that the field classification results include the field source path, source system identifier, field name, field value, and semantic category. When filtering similar heterogeneous fields based on the field classification results, this embodiment first groups the field semantic records according to semantic category, and then reads the source system identifier within each group. If a semantic record from both the procurement system and the warehousing system exists within the same semantic category group, this embodiment identifies the field semantic records in that semantic category group as similar heterogeneous fields. If a semantic category group contains only a single source system identifier, that semantic category group will not be included in subsequent field integrity filtering.
[0030] It is worth noting that the field integrity screening process handles null fields, fields with missing paths, and missing fields that do not participate in subsequent composite attribute identification. Specifically, this embodiment reads the field value, field source path, and character structure features of the aforementioned heterogeneous fields of the same type. If the field value is null, or the field source path lacks any one of the following: source system identifier, data object name, record identifier, or field name, this embodiment deletes the corresponding field. If the field value is not null and the field source path is complete, this embodiment continues to determine whether the field value has a medical supply attribute expression form. The medical supply attribute expression form includes any one of the following: containing unit characters, containing specification separators, containing date separators, containing consecutive Arabic numeral fragments, or containing Chinese supply name fragments. This embodiment groups the heterogeneous fields of the same type that meet the above conditions to obtain a candidate heterogeneous field set.
[0031] Among them, composite attribute identification is performed on the candidate heterogeneous field set to obtain material composite fields, including: Based on a preset medical supplies attribute dictionary, attribute matching is performed on the field names in the candidate heterogeneous field set to obtain name matching results; Based on a preset medical supplies attribute dictionary, attribute matching is performed on the field values in the candidate heterogeneous field set to obtain the field value matching result; Based on the name matching results and the field value matching results, fields containing at least two medical supply attribute categories are filtered to obtain composite supply fields.
[0032] In one implementation, the candidate heterogeneous field set is the field set obtained in this step after field semantic classification and field integrity filtering. Each candidate field in the candidate heterogeneous field set includes a field name, field value, field source path, source system identifier, and semantic category. The composite attribute identification refers to determining whether a candidate field simultaneously carries multiple medical supply attribute categories. The medical supply attribute categories include supply name, specifications, batch number, expiration date, unit, and supplier.
[0033] It should be noted that the preset medical supply attribute dictionary is generated using a database construction method. Specifically, in this embodiment, the fields of material name, specification model, batch number, expiration date, unit, and supplier are extracted from historical material master data records and standard material master data tables. High-frequency text fragments in each field are then assigned to their corresponding medical supply attribute categories. The extracted text fragments are processed by removing spaces, converting full-width characters to half-width characters, unifying capitalization, and deleting duplicates to obtain candidate attribute entries. The frequency of each candidate attribute entry in the historical material master data records is counted, and candidate attribute entries whose frequency reaches the minimum threshold are written into the preset medical supply attribute dictionary. The minimum threshold is set to 30 times, and its value is determined by extracting all candidate attribute entries from the historical material master data records, counting the frequency of each candidate attribute entry, deleting candidate attribute entries with a frequency of 1, and then selecting the frequency value corresponding to the 80th percentile from the frequency distribution of the remaining candidate attribute entries. The preset medical supplies attribute dictionary includes attribute entries, attribute categories, and matching types. The matching type is either field name matching or field value matching. Specifically, if the attribute entry comes from a historical field name or a standard field name, the matching type is field name matching; if the attribute entry comes from a historical field value or a standard field value, the matching type is field value matching. The same attribute entry can have both matching types simultaneously.
[0034] It is worth noting that when performing attribute matching identification on the field names in the candidate heterogeneous field set, this embodiment reads the field names of the candidate fields and performs space removal, connector removal, full-width character to half-width character conversion, and case unification on the field names to obtain unified field names. Subsequently, this embodiment compares the unified field names with attribute entries in the preset medical supplies attribute dictionary whose matching type is field name matching. If the unified field name contains a certain attribute entry, or if the unified field name is completely consistent with a certain attribute entry, then the corresponding field name matching category is generated. The field source path, field name, and all field name matching categories of the candidate fields are combined to obtain the name matching result.
[0035] It should be noted that when performing attribute matching identification on the field values in the candidate heterogeneous field set, this embodiment reads the field values of the candidate fields and establishes a character position sequence according to the original character order. Subsequently, this embodiment arranges the attribute entries with the matching type of field value matching in the preset medical supplies attribute dictionary from largest to smallest according to the entry length, and matches them item by item in the field values. If there are consecutive character segments in the field value that are the same as the attribute entry, this embodiment records the attribute category, the start position of the hit, and the end position of the hit corresponding to the attribute entry. For specification and unit attributes, this embodiment further reads the Arabic numerals, English letters, Chinese unit characters, and specification separators in the field values, and takes the consecutively appearing numerical unit segments as the hit segments of the specification or unit. The field source path, field value, field value hit category, and hit position of the candidate field are combined to obtain the field value hit result.
[0036] It is worth noting that when filtering composite material fields based on the name matching results and the field value matching results, this embodiment merges the field name matching categories and field value matching categories corresponding to the same field source path, and deletes duplicate medical material attribute categories to obtain an attribute category set. If the attribute category set contains at least two different medical material attribute categories, this embodiment determines the corresponding candidate field as a composite material field. If the attribute category set contains only one medical material attribute category, the corresponding candidate field remains a single attribute field and is not proceeded to the subsequent segmentation and positioning steps.
[0037] In step S13, the composite material field is segmented and located to obtain field segmentation points, and the composite material field is sliced according to the field segmentation points to obtain material attribute fragments with field source paths and slice location information.
[0038] Specifically, segmenting and locating the composite field of the material to obtain the field segmentation point includes: The delimiter is identified in the composite field of the material to obtain the explicit segmentation position; Based on the preset date format template, batch number prefix template, and specification unit template, attribute boundary identification is performed on the composite field of the material to obtain the implicit segmentation position; Field split points are generated based on the explicit split positions and the implicit split positions.
[0039] Furthermore, the material composite field is sliced according to the field segmentation point, and the field source path corresponding to the material composite field and the slice positioning information of each slice in the material composite field are retained to obtain material attribute fragments with field source path and slice positioning information.
[0040] In one implementation, the composite material field is the field obtained in step S12 that simultaneously contains at least two medical supply attribute categories. The composite material field includes a field name, field value, field source path, source system identifier, name hit result, and field value hit result. The segmentation and positioning refers to determining the boundary positions between attribute fragments within the field value. The field segmentation point is the character position between adjacent attribute fragments within the field value. The slice positioning information includes a slice start character number and a slice end character number, counted in units of characters. Each Chinese character, English letter, Arabic numeral, punctuation mark, and special symbol is counted as one character; byte counting is not used. The first character in the field value is marked as 1, and the slice end character number is the character number corresponding to the last character of the slice.
[0041] It should be noted that, when identifying delimiters for the composite material field, this embodiment reads the field value of the composite material field character by character, and uses hyphens, forward slashes, decimal points, asterisks, multiplication signs, spaces, commas, semicolons, parentheses, and colons in the field value as candidate delimiters. Subsequently, this embodiment confirms the candidate delimiter based on the character types on both sides of it. If there are Chinese characters and Arabic numerals on both sides of the candidate delimiter, or if there are specification unit segments and packaging unit segments on both sides of the candidate delimiter, then the character sequence number of the candidate delimiter is determined as the explicit delimiter position; if the candidate delimiter is located inside a date segment or inside a decimal value, then the candidate delimiter is not determined as the explicit delimiter position.
[0042] It should be noted that the preset date format template, batch number prefix template, and specification unit template are all generated using a database construction method. Specifically, this embodiment extracts fields whose names contain production date, expiration date, expiry date, validity period, and sterilization date from historical material master data records. It reads the date text from the corresponding field values and validates the year digits, month value range, and date value range. Date text that passes the validation is categorized into date format templates according to its character structure. The date format templates include a template with four-digit year, two-digit month, and two-digit date consecutively; a template with four-digit year followed by two-digit month followed by two-digit date; and a template with two-digit year followed by two-digit month followed by two-digit date. The connecting characters are hyphens, slashes, or decimal points.
[0043] It is worth noting that this embodiment extracts fields whose names contain batch number, batch number, and batch number expiration date from historical material master data records. It reads the prefix text preceding consecutive alphanumeric strings in the corresponding field values and writes the prefix text that appears at least 20 times into the batch number prefix template. The batch number prefix template includes batch number, batch number, LOT, Lot, lot, and B. This embodiment also extracts specification and unit of measurement fields from historical material master data records and standard material master data tables. It identifies text fragments formed by consecutive Arabic numerals, decimal points, English letter units, Chinese units, and specification connectors. After deleting duplicates and illegal characters, it counts the occurrences of each text fragment and writes text fragments that appear at least 40 times into the specification unit template. The specification unit template includes combinations of numerical values plus English letter units, numerical values plus Chinese units, and numerical units plus specification connectors plus numerical units.
[0044] It should be noted that when identifying the attribute boundaries of the composite material field, this embodiment first matches date segments in the field value according to a preset date format template and records the slice positioning information of the date segments. Then, this embodiment identifies the batch number prefix according to a preset batch number prefix template and reads the continuous alphanumeric string after the batch number prefix, determining the batch number prefix and the continuous alphanumeric string together as the batch number segment. Next, it identifies the specification segment and unit segment according to a preset specification unit template and records the slice positioning information of the corresponding segments. The slice positioning information of the above segments is used as the implicit segmentation position.
[0045] It is worth noting that when generating field split points based on the explicit and implicit split positions, this embodiment first sorts all explicit and implicit split positions in ascending order of character number and removes duplicate positions. If an explicit split position is located inside an implicit fragment, the boundary position corresponding to the implicit fragment is retained, and the explicit split position located inside the implicit fragment is deleted. If an implicit split position is located inside a fragment separated by an explicit delimiter, and the explicit fragment spans multiple attributes, the explicit split position containing the implicit split position is deleted, using the implicit split position as the boundary. If both explicit and implicit split positions exist at the same location, they are merged into one split point. If two split positions are adjacent and there are no visible characters in between, the split position closest to the boundary of the valid attribute fragment is retained. After the above processing, the field split points are obtained.
[0046] It should be noted that when performing attribute slicing on the composite material field according to the field segmentation points, this embodiment divides the field value into multiple continuous text segments according to the field segmentation points, and deletes text segments consisting only of delimiters. Segments without business meaning that do not contain valid characters or belong to any implicit template after segmentation are deleted. For each retained text segment, the segment content, field source path, source system identifier, and slice location information are recorded to form a material attribute segment with field source path and slice location information.
[0047] In step S14, attribute attribution is marked according to the fragment content of the material attribute fragment, the slice positioning information and the relationship between adjacent fragments to obtain structural attribute fragments, and cross-system format alignment is performed on the structural attribute fragments to obtain standard attribute features.
[0048] Specifically, attribute attribution is performed based on the fragment content of the material attribute fragment, the slice positioning information, and the relationship between adjacent fragments to obtain structural attribute fragments, including: Tag matching is performed based on the fragment content and preset entity tag rules to obtain tag matching results; When there are multiple candidate labels in the label matching result, label disambiguation is performed based on the slice positioning information and the relationship between adjacent segments to obtain the target attribute label; The target attribute label, the material attribute fragment, the field source path, and the slice location information are bound together to obtain the structural attribute fragment.
[0049] In one implementation, the material attribute fragment is a continuous text fragment with slice location information obtained in step S13. Each material attribute fragment includes fragment content, field source path, source system identifier, and slice location information. Attribute attribution labeling refers to classifying each material attribute fragment into one of the following attribute categories: material name, date, unit, specification, supplier, or batch number.
[0050] It should be noted that the preset entity label rules are generated using a database construction method. Specifically, in this embodiment, the already split field content is read from historical material master data records and standard material master data tables. The field values, defined as material name, production date, expiration date, validity period, unit of measurement, packaging unit, specifications, supplier name, supplier code, batch number, or batch number, are written into the corresponding label samples. Character structure statistics and text fragment extraction are performed on various label samples to obtain various entity label rules. The preset entity label rules include label name, matching conditions, and disambiguation order. The label name includes material name, date type, unit type, specification type, supplier, and batch number. The matching conditions include field name hit conditions, field value character structure conditions, and fragment position conditions. The field name hit conditions are derived from the preset medical material attribute dictionary in step S12. The field value character structure conditions are derived from the character type distribution in the corresponding label samples. The fragment position conditions are calculated by statistically analyzing the ratio of the starting position of each attribute label fragment in the corresponding material composite field to the total length of the field in the historical material master data records. The 25th percentile of the ratio distribution is taken. The interval from the 12th percentile to the 75th percentile is used as the position determination interval for the attribute label. When the ratio of the starting position of the segment to be labeled falls into this interval, the attribute label is used as one of the candidate labels. The disambiguation order is used as follows: when the label matching result is multiple candidate labels, the candidate labels are matched with the segment content in descending order of the F1 score of the label sample in the manually verified sample set. The first candidate label that passes all matching conditions is used as the target attribute label. If all candidate labels fail the matching conditions, the label category already labeled in the adjacent segment is selected as a reference based on the slice positioning information. The intersection of the reference label and the candidate label is used as the target attribute label.
[0051] It is worth noting that during tag matching, this embodiment reads the fragment content and slice positioning information of the material attribute fragment, and matches the fragment content with the preset entity tag rules one by one. If the fragment content conforms to the date character structure and satisfies the date validity check, a date candidate tag is generated; if the fragment content contains consecutive Arabic numerals and unit characters, a specification candidate tag or a unit candidate tag is generated; if the fragment content matches the supplier tag sample or supplier code structure, a supplier candidate tag is generated; if the fragment content matches the material name tag sample, or if the fragment content consists of consecutive Chinese characters and is located at the beginning of the material composite field, a material name candidate tag is generated. All candidate tags and corresponding fragment identifiers together form the tag matching result. The date validity check includes year digit check, month value range check, and date value range check, where the month value is 1 to 12, and the date value is limited to the range of 1 to 28, 1 to 29, 1 to 30, or 1 to 31 depending on the month; if a date fragment fails any check, no date candidate tag is generated.
[0052] It should be noted that when multiple candidate tags exist in the tag matching result, this embodiment performs tag disambiguation based on the slice positioning information and the relationship between adjacent segments. If a segment simultaneously matches both the specification category candidate tag and the unit category candidate tag, and the segment contains Arabic numerals and unit characters, then the segment is labeled as a specification category; if a segment only contains unit characters, and the adjacent preceding segment has already been labeled as a specification category, then the segment is labeled as a unit category; if a Chinese segment is located in the starting area of the material composite field, and the adjacent following segment is a specification category, then the Chinese segment is labeled as a material name; if a mixed alphanumeric segment has a batch number prefix, then the segment is labeled as a batch number; if a segment simultaneously matches both the material name and the supplier candidate tag, and the segment content matches the preset supplier name dictionary, then it is labeled as a supplier; otherwise, it is labeled as a material name. After disambiguation, the target attribute tag is obtained; if none of the above preset rules can disambiguate, then the candidate tag with the earliest position is selected as the target attribute tag based on the slice's starting position. Subsequently, in this embodiment, the target attribute label, the material attribute fragment, the field source path, and the slice location information are bound together to generate a structural attribute fragment. The structural attribute fragment includes the target attribute label, fragment content, field source path, source system identifier, slice location information, and original field name.
[0053] Specifically, cross-system format alignment is performed on the structural attribute fragments to obtain standard attribute features, including: The material name fragments in the structural attribute fragments are normalized to obtain standard name fragments; The date fragments in the structural attribute fragments are formatted to obtain standard date fragments; Unit-class segments in the structural attribute segments are normalized to obtain standard unit segments; The specification-type segments in the structural attribute segments are split into numerical units to obtain standard specification segments; Supplier normalization is performed on the supplier segments in the structural attribute segments to obtain standard supplier segments; By combining the standard name fragment, the standard date fragment, the standard unit fragment, the standard specification fragment, and the standard supplier fragment, standard attribute features are obtained.
[0054] In one implementation, cross-system format alignment refers to converting structural attribute fragments written differently in the procurement and warehousing systems into a unified data structure. Standard attribute features are key-value structures, where keys include standard name fragments, standard date fragments, standard unit fragments, standard specification fragments, and standard supplier fragments, and values are the standard text of the corresponding fragments after normalization, unification, or splitting, while retaining the source path and slice location information of the corresponding fields.
[0055] It should be noted that name normalization is performed according to preset name normalization rules. These rules are constructed from the material name field in the standard material master data table. Specifically, the system reads the standard material name and performs the following operations: space removal, bracket removal, connector removal, conversion of full-width characters to half-width characters, and case unification, resulting in a standard name index. After reading a material name fragment, the same processing is performed, and the processed fragment is matched against the standard name index. If a match is found, the corresponding standard material name is output, resulting in a standard name fragment. If a match is not found, the processed fragment is retained, and a "missed" flag is added. Name fragments with a "missed" flag are considered as having an empty standard name fragment when extracting records to be scored later. The system also matches these fragments based on standard specification fragments and standard unit fragments, and outputs the "missed" flag along with the mapped record for manual review.
[0056] It should be noted that the format of date fragments is uniformly executed according to the preset date format rules. The preset date format rules are generated by the preset date format template in step S13. Specifically, the year, month, and day in the date fragment are read, the two-digit year is padded to a four-digit year according to the year distribution of the date field in the historical material master data record, zeros are added before the numbers of months and days that are less than two digits, and a standard date fragment is generated in the order of four-digit year, hyphen, two-digit month, hyphen, and two-digit day.
[0057] It is worth noting that unit normalization is performed according to preset unit conversion rules. The preset unit conversion rules are constructed from the measurement unit field and packaging unit field in the standard material master data table. Specifically, it reads the smallest measurement unit, packaging unit, and packaging conversion quantity under the same material code to form a unit conversion record including the source unit, target unit, and conversion quantity. When normalizing unit fragments, the unit fragment is matched with the source unit. If a match is found, the corresponding target unit is read to obtain the standard unit fragment. If the unit fragment is already equal to the target unit, the unit fragment is directly determined as the standard unit fragment.
[0058] It is worth noting that the numerical unit splitting of specification fragments follows a preset specification normalization rule. This rule is constructed from the specification model field in the standard material master data table. Specifically, it reads consecutive numerical values, unit characters, and specification connectors from the specification fragment, treating consecutive numerical values as specification values, unit characters as specification units, and specification connectors as hierarchical connection identifiers. If a specification fragment contains multiple combinations of numerical units, they are recorded sequentially according to their original order of appearance, and the specification values, specification units, and hierarchical connection identifiers are combined into a standard specification fragment. The hierarchical connection identifier is the character connecting different combinations of numerical units in the specification fragment, including "×", "*", and " / ". The hierarchical connection identifier in the standard specification fragment is uniformly represented by "×". When splitting numerical units, they are extracted sequentially according to the alternating order of numerical values and units. The extraction rule is that consecutive digit strings (including decimal points) are treated as specification values, the unit character immediately adjacent to that value is treated as the specification unit, and the connector between adjacent combinations of numerical units is treated as the hierarchical connection identifier.
[0059] It should be noted that supplier normalization is performed according to preset supplier normalization rules. These rules are constructed from the supplier name and supplier code fields in the standard material master data table. Specifically, the standard supplier name and standard supplier code are read, spaces, parentheses, and connectors are removed from the supplier name, and a supplier name index is created. The same processing is then performed on supplier fragments, which are then matched against the supplier name index. If a match is found, the corresponding standard supplier name and standard supplier code are output to obtain the standard supplier fragment. If the supplier fragment is a supplier code, the standard supplier name and standard supplier code are retrieved based on the supplier code.
[0060] It is worth noting that when combining the standard name fragment, the standard date fragment, the standard unit fragment, the standard specification fragment, and the standard supplier fragment, this embodiment writes them into the same standard attribute record using a fixed key name. If a certain type of fragment does not exist in the current structural attribute fragment, a missing tag is written under the corresponding key name. The missing tag does not delete the field source path and slice location information. The resulting standard attribute feature includes the standardized attribute content and the original source location content.
[0061] In step S15, a standard mapping association is performed based on the standard attribute features to obtain a mapping relationship, including: Based on the standard attribute features, target standard material records that meet the preset similarity threshold are selected from the pre-constructed standard material master data feature library; Determine the target field path corresponding to the structural attribute fragment based on the target standard material record; The source path of the field corresponding to the structural attribute fragment, the slice location information, and the target field path are bound together to obtain a mapping relationship.
[0062] In one implementation, the standard attribute features are the standardized attribute records obtained in step S14, including standard name fragments, standard date fragments, standard unit fragments, standard specification fragments, standard supplier fragments, field source paths, and slice location information. The standard mapping association refers to the process in this embodiment of matching the standard attribute features with standard material records in a pre-built standard material master data feature library and mapping structural attribute fragments to target field paths. The mapping relationship includes field source paths, slice location information, target field paths, target attribute tags, and target standard material record identifiers.
[0063] It should be noted that the standard material master data table is a standard field table jointly accessed by the procurement system and the warehousing system. Each record in the table stores the standard material code, standard material name, standard specification, standard unit, standard supplier, and standard field code in separate columns. The pre-built standard material master data feature library is generated using a database construction method. Specifically, this embodiment reads the standard material code, standard material name, standard specification, standard unit, standard supplier, and standard field code from the standard material master data table; performs space deletion, bracket deletion, connector deletion, full-width character conversion to half-width character, and case unification on the standard material name to obtain the standard name index; performs numerical unit splitting on the standard specification to obtain the standard specification index; converts the standard unit to a standard unit index according to the unit normalization method in step S14; converts the standard supplier to a standard supplier index according to the supplier normalization method in step S14; combines the standard material record object name and standard field code in hierarchical order to obtain the target field path; and writes the standard material code, standard name index, standard specification index, standard unit index, standard supplier index, and target field path into the same index record to obtain the standard material master data feature library.
[0064] It is worth noting that when filtering target standard material records from the pre-built standard material master data feature library, the records to be scored are first extracted based on the standard attribute features. Specifically, in this embodiment, the standard name fragment in the standard attribute features is read, and the standard name fragment is matched with the standard name index using character binary segmentation to obtain the first record to be scored. Specifically, character binary segmentation matching involves dividing the standard name fragment and the standard material name into sets of adjacent two-character fragments, calculating the Jaccard similarity coefficient between the two sets, and writing the standard material record into the first record to be scored when the Jaccard similarity coefficient is greater than or equal to 0.4. The 0.4 is determined by statistically analyzing the distribution of Jaccard similarity coefficients of historical positive and negative sample records and selecting the critical value corresponding to the largest difference between the true positive rate and the false positive rate. Character binary segmentation refers to splitting text into a set of fragments formed by two adjacent characters according to character order. When any character fragment in the standard name fragment appears in the standard name index, the corresponding standard material record is written into the first record to be scored. If the standard name fragment is empty or contains a missing flag, the standard specification fragment and standard unit fragment are read and matched according to the standard specification index and standard unit index to obtain the second record to be scored. If the standard name fragment is not empty and the first record to be scored is empty, the second record to be scored is generated according to the standard specification fragment and standard unit fragment. The first and second records to be scored are merged, and duplicate records are deleted to obtain the set of material records to be scored.
[0065] It should be noted that in this embodiment, each material record in the set of material records to be evaluated is scored based on name similarity, specification consistency, unit convertibility, and supplier consistency to obtain the record similarity score. Name similarity is determined by character binary segmentation results. Specifically, the standard name fragment and the standard material name are each segmented by character binary, the number of identical character fragments in the two fragment sets is counted, and the number of duplicate character fragments after merging the two fragment sets is counted. The two counts are then divided to obtain the name similarity score. If the standard name fragment is empty or has a missing marker, the name similarity score is recorded as 0. Specification consistency is determined by standard specification fragments and standard specification indexes. If the specification values are the same and the specification units are the same, the specification consistency score is recorded as 1. If the specification values are the same and the specification units are the same after conversion using a preset unit conversion rule, the specification consistency score is recorded as 0.8. The remaining cases are recorded as 0. Unit convertibility is determined by standard unit fragments and standard unit indexes. If the units are the same, the unit convertibility is recorded as 1; if the units can be converted to the same target unit using preset unit conversion rules, the unit convertibility is recorded as 0.8; otherwise, it is recorded as 0. Supplier consistency is determined by standard supplier fragments and standard supplier indexes. If the standard supplier codes are the same, the supplier consistency is recorded as 1; if the standard supplier code is missing but the standard supplier name is the same, the supplier consistency is recorded as 0.9; if the standard supplier fragment has a missing tag or the standard supplier name is different, the supplier consistency is recorded as 0.
[0066] It is worth noting that the record similarity is obtained by weighted summation of name similarity, specification consistency, unit convertibility, and supplier consistency. The weight of name similarity is set to 0.45, specification consistency to 0.30, unit convertibility to 0.15, and supplier consistency to 0.10. The above weights are determined through historical mapping confirmation records. Specifically, in this embodiment, combinations that were successfully written back and did not undergo correction are extracted from the historical mapping confirmation records to form positive samples, and combinations that were returned or had their standard material codes changed after writing back are extracted to form negative samples. Under the condition that the sum of the four weights is 1, the weight combinations are traversed with a step size of 0.05, and the F1 value of each combination on the verification sample set is calculated. The weighted combination with the largest F1 value is selected as a candidate. If there are multiple weighted combinations with the largest F1 value, the weighted combination with the fewest false matches is selected. False matches are false positive records that predict negative samples as positive samples. The F1 value is the harmonic mean of precision and recall. The preset similarity threshold is set to 0.72. This threshold is determined by statistically analyzing the similarity between positive and negative samples. The threshold is traversed within the range of 0.50 to 0.95 with a step size of 0.01. The difference between the true positive rate and the false positive rate corresponding to each threshold is calculated, and the threshold corresponding to the largest difference is selected as the preset similarity threshold. Positive samples are records in the historical mapping confirmation records that have been successfully written back and have not been corrected. Negative samples are records that have been returned or had their standard material codes changed after being written back. The ratio of positive to negative samples is 1:1.
[0067] It should be noted that when selecting target standard material records that meet the preset similarity threshold, this embodiment determines material records in the set of material records to be scored with a similarity of not less than 0.72 as material records that meet the preset similarity threshold. If only one material record meets the preset similarity threshold, then that material record is determined as the target standard material record. If multiple material records meet the preset similarity threshold, then the material record with the highest similarity is selected; if the similarity is the same, then the material record whose standard material update time is closer to the current processing time is selected; if the update times are still the same, then the material record whose standard material code is arranged first when sorted by character code from smallest to largest is selected. If the set of material records to be scored does not contain any material records with a similarity of not less than 0.72, then no automatic confirmation mapping record will be generated, and a mapping record to be reviewed will be generated instead. The mapping record to be reviewed includes standard attribute features, field source paths, slice location information, and the record similarity of each material record to be scored. The mapping record to be reviewed is written into the review queue, and the mapping relationship is confirmed or corrected by manual verification. The manually confirmed mapping record is written back to the standard material master data feature library for the expansion of candidate records for subsequent mapping matching. Records in the review queue do not participate in the automatic mapping output before manual processing.
[0068] It is worth noting that the target field path originates from the standard field codes in the target standard material record. The target standard material record stores the standard material record object name and standard field codes. In this embodiment, the target attribute tag is read from the structural attribute fragment, and the target field path is determined based on the mapping table between the target attribute tag and the standard field code. This mapping table is constructed from the field codes and Chinese field names in the standard material master data table. When the Chinese field name contains the material name, it corresponds to the material name tag; when it contains specifications or model, it corresponds to the specification category tag; when it contains the unit or packaging unit, it corresponds to the unit category tag; when it contains the supplier name or supplier code, it corresponds to the supplier tag; when it contains the production date, expiration date, or validity period, it corresponds to the date category tag; and when it contains the batch number or batch number, it corresponds to the batch number tag. The target field path is obtained by combining the standard material record object name and the corresponding standard field code in hierarchical order.
[0069] It should be noted that when generating the mapping relationship, this embodiment reads the field source path, slice start character number, and slice end character number corresponding to the structural attribute fragment, and reads the target field path, writing the above content into the same mapping record. The mapping record also includes the target standard material code, target attribute label, and standardized fragment content. After processing all structural attribute fragments, all mapping records form a mapping relationship.
[0070] It should be noted that the standard material master data feature library, field thesaurus, medical material attribute dictionary, and entity label rules are incrementally reconstructed based on newly added material master data records and manual review and confirmation records at a preset cycle (every 30 days). During incremental reconstruction, the newly added historical records are merged with the original historical records and the construction steps of each feature library are re-executed.
[0071] In summary, this invention performs a series of steps on the master data records of the procurement and warehousing systems. These steps include extracting field structures, filtering heterogeneous fields, identifying composite attributes, segmenting and locating data, slicing attributes, labeling attribute affiliation, aligning cross-system formats, and associating with standard mappings. This forms a complete processing flow from identifying multi-source heterogeneous fields to generating mapping relationships. The mapping relationships carry the field source path, slice location information, and target field path, thus achieving accurate mapping of multi-source heterogeneous master data of medical supplies.
[0072] It should be noted that the multi-source heterogeneous master data mapping system for medical supply chain provided in this embodiment of the invention is used to execute all the process steps of the multi-source heterogeneous master data mapping method for medical supply chain in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.
[0073] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0074] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for mapping multi-source heterogeneous master data in the medical supplies supply chain, characterized in that, include: Obtain the master data records of materials from the procurement system and warehousing system in the medical supplies supply chain, extract the field structure of the master data records of materials, and obtain the material field features; Heterogeneous fields are filtered based on the characteristics of the material fields to obtain a set of candidate heterogeneous fields. Composite attributes are then identified from the set of candidate heterogeneous fields to obtain composite material fields. The composite material field is segmented and located to obtain field segmentation points, and the composite material field is sliced according to the field segmentation points to obtain material attribute fragments with field source paths and slice location information. Based on the fragment content of the material attribute fragment, the slice positioning information, and the relationship between adjacent fragments, attribute attribution is marked to obtain structural attribute fragments, and cross-system format alignment is performed on the structural attribute fragments to obtain standard attribute features; Based on the standard attribute features, a standard mapping association is performed to obtain a mapping relationship, which includes the field source path, slice location information, and target field path.
2. The multi-source heterogeneous master data mapping method for the medical supplies supply chain according to claim 1, characterized in that, The process of extracting the field structure from the material master data records to obtain material field features includes: Parse the field names, field values, field source paths, and source system identifiers in the material master data records to obtain basic field information; Character type statistics are performed on the basic information of the fields to obtain character structure features; Field path features are obtained by performing field-level identification on the basic information of the fields. By combining the character structure features and the field path features, the material field features are obtained.
3. The multi-source heterogeneous master data mapping method for the medical supplies supply chain according to claim 1, characterized in that, The step of filtering heterogeneous fields based on the characteristics of the material fields to obtain a candidate heterogeneous field set includes: Based on the material field characteristics and the preset field synonym list, the fields are semantically categorized to obtain the field categorization results; Based on the field classification results, fields from different source systems but with the same semantic category are filtered to obtain homogeneous fields; Based on the aforementioned heterogeneous fields, a set of candidate heterogeneous fields is obtained by performing field integrity filtering.
4. The multi-source heterogeneous master data mapping method for the medical supplies supply chain according to claim 1, characterized in that, The step of performing composite attribute identification on the candidate heterogeneous field set to obtain material composite fields includes: Based on a preset medical supplies attribute dictionary, attribute matching is performed on the field names in the candidate heterogeneous field set to obtain name matching results; Based on a preset medical supplies attribute dictionary, attribute matching is performed on the field values in the candidate heterogeneous field set to obtain the field value matching result; Based on the name matching results and the field value matching results, fields containing at least two medical supply attribute categories are filtered to obtain composite supply fields.
5. The multi-source heterogeneous master data mapping method for the medical supplies supply chain according to claim 1, characterized in that, The step of segmenting and locating the composite field of the material to obtain the field segmentation point includes: The delimiter of the composite material field is identified to obtain the explicit segmentation position; Based on the preset date format template, batch number prefix template, and specification unit template, attribute boundary identification is performed on the composite field of the material to obtain the implicit segmentation position; Field split points are generated based on the explicit split positions and the implicit split positions.
6. The multi-source heterogeneous master data mapping method for the medical supplies supply chain according to claim 1, characterized in that, The step of assigning attributes based on the content of the material attribute fragment, the slice positioning information, and the relationship between adjacent fragments to obtain structural attribute fragments includes: Tag matching is performed based on the fragment content and preset entity tag rules to obtain tag matching results; When there are multiple candidate labels in the label matching result, label disambiguation is performed based on the slice positioning information and the relationship between adjacent segments to obtain the target attribute label; The target attribute label, the material attribute fragment, the field source path, and the slice location information are bound together to obtain the structural attribute fragment.
7. The multi-source heterogeneous master data mapping method for the medical supplies supply chain according to claim 1, characterized in that, The cross-system format alignment of the structural attribute fragments to obtain standard attribute features includes: The material name fragments in the structural attribute fragments are normalized to obtain standard name fragments; The date fragments in the structural attribute fragments are formatted to obtain standard date fragments; Unit-class segments in the structural attribute segments are normalized to obtain standard unit segments; The specification-type segments in the structural attribute segments are split into numerical units to obtain standard specification segments; Supplier normalization is performed on the supplier segments in the structural attribute segments to obtain standard supplier segments; By combining the standard name fragment, the standard date fragment, the standard unit fragment, the standard specification fragment, and the standard supplier fragment, standard attribute features are obtained.
8. The multi-source heterogeneous master data mapping method for the medical supplies supply chain according to claim 1, characterized in that, The step of performing standard mapping association based on the standard attribute features to obtain the mapping relationship includes: Based on the standard attribute features, target standard material records that meet the preset similarity threshold are selected from the pre-constructed standard material master data feature library; Determine the target field path corresponding to the structural attribute fragment based on the target standard material record; The source path of the field corresponding to the structural attribute fragment, the slice location information, and the target field path are bound together to obtain a mapping relationship.
9. A multi-source heterogeneous master data mapping system for the medical supplies supply chain, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the method described in any one of claims 1 to 8.