Data migration method and device, equipment, medium and product
The mapping relationship of data storage architecture description information is determined through natural language processing technology, which solves the problems of inefficiency and error-prone traditional manual data migration methods, and achieves the efficiency and accuracy of data migration.
Patent Information
- Application Number
- CN202510199180.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
After the asset management system of complex organizations such as medical institutions is changed, when the data of the original system is migrated to the new system, the traditional manual migration method is inefficient and prone to errors, especially when processing large amounts of data and complex department correspondence.
Through natural language processing technology, the mapping relationship of data storage architecture description information of different two systems is determined, and the semantic feature information is extracted using the pre-trained semantic recognition model, and the mapping relationship of data storage architecture description information is established based on these feature information, thereby realizing data migration.
Improve the accuracy and efficiency of data migration, ensure that data remains complete, accurate and consistent during the migration process, and reduce the occurrence of human errors.
Smart Images

Figure CN119988353A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computer technology, and in particular, to a data migration method, device, equipment, medium and product. Background Art
[0002] When complex organizations such as medical institutions make changes to their asset management systems, it is necessary to migrate data from the original system to the new system. Migrating historical asset files from the original asset management system is a complex and time-consuming process that requires ensuring data integrity, accuracy, and consistency.
[0003] However, the traditional manual migration method is not only inefficient but also prone to errors, especially when dealing with large amounts of data and complex department correspondence. Summary of the invention
[0004] Embodiments of the present invention provide a data migration method, apparatus, device, medium and product, which can determine the mapping relationship between data storage architecture description information of two different systems through natural language processing technology, thereby improving the accuracy and efficiency of data migration.
[0005] In a first aspect, an embodiment of the present invention provides a data migration method, the method comprising:
[0006] Obtaining first data storage architecture description information in the first data management system and second data storage architecture description information in the second data management system; wherein the first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organizational architectures of the corresponding data management systems;
[0007] Extracting semantic feature information from the first data storage architecture description information and the second data storage architecture description information through a pre-trained semantic recognition model, and establishing a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information based on the semantic feature information;
[0008] According to the mapping relationship, the system data in the first data management system is migrated to the data directory of the corresponding description information in the second data management system.
[0009] In a second aspect, an embodiment of the present invention further provides a data migration device, the device comprising:
[0010] A data acquisition module, used to acquire first data storage architecture description information in the first data management system and second data storage architecture description information in the second data management system; wherein the first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organizational architectures of the corresponding data management systems;
[0011] A mapping relationship determination module, used to extract semantic feature information in the first data storage architecture description information and the second data storage architecture description information through a pre-trained semantic recognition model, and establish a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information based on the semantic feature information;
[0012] The data migration module is used to migrate the system data in the first data management system to the data directory of the corresponding description information in the second data management system according to the mapping relationship.
[0013] In a third aspect, an embodiment of the present invention further provides a computer device, the computer device comprising:
[0014] one or more processors;
[0015] A memory for storing one or more programs;
[0016] When one or more programs are executed by one or more processors, the one or more processors implement the data migration method provided by any embodiment of the present invention.
[0017] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a data migration method as provided in any embodiment of the present invention.
[0018] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, including a computer program, which, when executed by a processor, implements a data migration method as provided in any embodiment of the present invention.
[0019] The embodiments of the above invention have the following advantages or beneficial effects:
[0020] According to an embodiment of the present invention, first data storage architecture description information in a first data management system and second data storage architecture description information in a second data management system are obtained; wherein the first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organizational architecture of the corresponding data management system; semantic feature information in the first data storage architecture description information and the second data storage architecture description information are extracted through a pre-trained semantic recognition model, and a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information is established based on the semantic feature information; according to the mapping relationship, the system data in the first data management system is migrated to the data directory of the corresponding description information in the second data management system. The technical solution of the embodiment of the present invention solves the problem of difficulty in data migration of complex data systems at present, and the mapping relationship between the data storage architecture description information of two different systems can be determined through natural language processing technology, thereby improving the accuracy and efficiency of data migration. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flow chart of a data migration method provided by an embodiment of the present invention;
[0022] Figure 2 is a flow chart of a data migration method provided by an embodiment of the present invention;
[0023] Figure 3 is a flow chart of a data migration method provided by an embodiment of the present invention;
[0024] Figure 4 is a structural schematic diagram of a data migration device provided by an embodiment of the present invention;
[0025] Figure 5 It is a structural schematic diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0027] Figure 1 This is a flow chart of a data migration method provided by an embodiment of the present invention. This embodiment can be applied to data migration scenarios, especially data migration of medical institutions. The method can be executed by a data migration device, which can be implemented by software and / or hardware and integrated into a computer device with application development function.
[0028] like Figure 1 As shown, the data migration method of this embodiment includes the following steps:
[0029] S110, obtaining first data storage architecture description information in the first data management system and second data storage architecture description information in the second data management system.
[0030] The first data management system and the second data management system are different data management systems, and may be data management systems of institutions with complicated data management requirements, for example, systems used by medical institutions to manage medical asset data.
[0031] The first data management system may be the data source system when performing data migration, and the second data management system may be the target migration system when performing data migration. For example, due to the iterative upgrade of the business of an organization, a new data management system needs to be replaced. In this embodiment, the first data management system may be the data management system used historically, and the second data management system may be the new data management system that needs to be replaced.
[0032] The first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organizational architecture of the corresponding data management system. The system organizational architecture may be an organizational architecture corresponding to the business system management system of the application data management system. For example, the system organizational architecture may be an organizational architecture corresponding to the medical institution used to associate medical asset data, including different departments.
[0033] Since the old and new systems may come from different manufacturers, or different versions from the same manufacturer, they cannot be directly imported through the relevant migration program. Instead, the data storage architecture description information needs to be analyzed and matched to achieve data migration based on the matching results.
[0034] The data storage architecture description information may be information describing various parts of the data storage, such as the classified storage category information and file directory information of the institution. If the institution is a medical institution, the data storage architecture description information may be organizational structure description information, i.e., department name information, such as "Department of Cardiology", "Department of Radiology - Imaging Center", etc.
[0035] S120. Extract semantic feature information from the first data storage architecture description information and the second data storage architecture description information through a pre-trained semantic recognition model, and establish a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information based on the semantic feature information.
[0036] The semantic recognition model can be obtained by training based on a natural language processing (NLP) model, and the natural language processing model can include a BERT (Bidirectional Encoder Representations from Transformers) model and a Word2Vec model, etc., which is not limited in this embodiment. The semantic recognition model performs word segmentation, keyword extraction and vector encoding through the data storage architecture description information to obtain semantic feature information of the first data storage architecture description information and the second data storage architecture description information.
[0037] After determining the semantic feature information, the mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information is determined by similarity calculation matching or classification model prediction matching. By determining the mapping relationship, it can ensure that the subsequently migrated data is consistent with the current organizational structure and department planning of the medical institution, thereby improving management efficiency.
[0038] S130. Migrate system data in the first data management system to a data directory of corresponding description information in the second data management system according to the mapping relationship.
[0039] According to the mapping relationship, the system data in the first data management system is migrated to the data directory of the corresponding description information in the second data management system through a data migration tool or a data migration program written in a programming language. For example, the system data of "Radiology Department - Imaging Center" in the first data management system is migrated to the data directory of "Medical Imaging Department" in the second data management system. The system data can be the archive data of the stock fixed assets, which can include the basic information of the assets (such as name, model, purchase date, etc.), the using department and the management department, etc.
[0040] After completing the data migration, the system data in the migrated first data management system can be automatically archived, and the database can be divided when the data volume exceeds the preset data volume threshold. Archiving can include query display and statistical analysis of the data, and a migration report can be generated based on the archived information. By establishing a data archiving mechanism, the traceability of historical data, i.e., migrated data, can be ensured, which is convenient for subsequent auditing and querying.
[0041] The technical solution of this embodiment is to obtain the first data storage architecture description information in the first data management system and the second data storage architecture description information in the second data management system; wherein the first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organizational structure of the corresponding data management system; extract the semantic feature information in the first data storage architecture description information and the second data storage architecture description information through a pre-trained semantic recognition model, and establish a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information based on the semantic feature information; according to the mapping relationship, migrate the system data in the first data management system to the data directory of the corresponding description information in the second data management system. The technical solution of the embodiment of the present invention solves the problem of difficulty in data migration of complex data systems at present, and can determine the mapping relationship between the data storage architecture description information between the two systems for data migration through natural language processing technology, thereby improving the accuracy and efficiency of data migration.
[0042] Figure 2 This is a flow chart of a data migration method provided in an embodiment of the present invention. This embodiment and the data migration method in the above embodiment belong to the same inventive concept, and further describes the process of determining the mapping relationship of storage architecture description information. The method can be executed by a data migration device, which can be implemented by software and / or hardware and integrated in a computer device with application development function.
[0043] like Figure 2 As shown, the data migration method of this embodiment includes the following steps:
[0044] S210: Obtain first data storage architecture description information in the first data management system and second data storage architecture description information in the second data management system.
[0045] The first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organization architecture of the corresponding data management system.
[0046] S220 , analyzing and processing the first data storage architecture description information and the second data storage architecture description information through the word segmentation module of the semantic recognition model to obtain information segmentation results.
[0047] The word segmentation module of the semantic recognition model can be a word segmentation tool, such as Jieba word segmentation, HanLP (Han Language Processing) and PKUSeg (Peking University segmentation), etc. The word segmentation tool is used to split the data storage architecture description information, such as the department name, to obtain the information segmentation result.
[0048] S230. Extract keywords from the information segmentation result through the keyword extraction module of the semantic recognition model to obtain the information keyword extraction result.
[0049] The keyword extraction module can be a keyword extraction tool such as rake-nltk, THUCKE and glossary, which extracts keywords from the information segmentation results, such as "cardiovascular" and "imaging".
[0050] S240. Perform vector encoding and semantic extraction on the information keyword extraction results through the semantic extraction module of the semantic recognition model to obtain corresponding semantic feature information.
[0051] The semantic extraction module can be a pre-trained word embedding model, such as BERT and Word2Vec, which converts the information keyword extraction results into word vectors. The semantic extraction module extracts semantic feature information based on the word vectors. Specifically, it can calculate the vector similarity between words to extract semantic feature information. For example, the semantic feature information of "cardiology" is extracted based on the short vector distance between "cardiology" and "cardiovascular medicine".
[0052] S250 , respectively calculating the semantic similarity between the semantic feature information of each description information in the first data storage architecture description information and the semantic feature information of each description information in the second data storage architecture description information.
[0053] For example, Word2Vec generates word vectors through context prediction, and calculates the semantic similarity between the semantic feature information of each description information in the first data storage architecture description information and the semantic feature information of each description information in the second data storage architecture description information by the proximity of similar words in the vector space.
[0054] In an optional embodiment, after calculating the semantic similarity between the semantic feature information of each description information in the first data storage architecture description information and the semantic feature information of each description information in the second data storage architecture description information, the corresponding semantic similarity is corrected according to the organizational structure hierarchy information corresponding to the description information where each semantic feature information is located to obtain a corrected semantic similarity.
[0055] For example, in the BERT model, some words in the input sequence are randomly masked during the training phase, forcing the model to predict the masked words through the context, thereby learning local and global semantic associations. In this embodiment, the local and global semantic associations can be organizational structure hierarchy information, that is, department hierarchy relationships (such as "surgery → general surgery → hepatobiliary surgery") or functional labels (such as "diagnosis category" and "treatment category"). After pre-training is completed, the BERT model corrects the corresponding semantic similarity according to the features learned in the training phase and the organizational structure hierarchy information corresponding to the descriptive information where each semantic feature information is located, and obtains the corrected semantic similarity, thereby improving the accuracy of matching the data storage architecture description information.
[0056] S260: Determine semantic feature information according to the value of semantic similarity to establish a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information.
[0057] According to the value of semantic similarity, for each description information in the first data storage architecture description information, the description information with the highest semantic similarity value, that is, the most similar description information in the second data storage architecture description information is used as matching description information, and a mapping relationship between the description information and the matching description information is established.
[0058] S270: Migrate the system data in the first data management system to the data directory of the corresponding description information in the second data management system according to the mapping relationship.
[0059] The technical solution of this embodiment is to obtain the first data storage architecture description information in the first data management system and the second data storage architecture description information in the second data management system; wherein the first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organizational architecture of the corresponding data management system; the first data storage architecture description information and the second data storage architecture description information are analyzed and processed by the word segmentation module of the semantic recognition model to obtain the information word segmentation result; the keywords in the information word segmentation result are extracted by the keyword extraction module of the semantic recognition model to obtain the information keyword extraction result; the information keyword extraction result is vector encoded and semantically extracted by the semantic extraction module of the semantic recognition model to obtain the corresponding semantic feature information; the semantic similarity of the semantic feature information of each description information in the first data storage architecture description information and the semantic feature information of each description information in the second data storage architecture description information is calculated respectively; the semantic feature information is determined according to the numerical value of the semantic similarity to establish a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information; according to the mapping relationship, the system data in the first data management system is migrated to the data directory of the corresponding description information in the second data management system. The technical solution of the embodiment of the present invention solves the problem of difficulty in data migration in complex data systems. Semantic feature information can be determined through a semantic recognition model, and the mapping relationship between the data storage architecture description information between the two systems to be migrated can be determined based on the similarity calculated according to the semantic feature information, thereby improving the accuracy and efficiency of data migration.
[0060] Figure 3 This is a flow chart of a data migration method provided in an embodiment of the present invention. This embodiment and the data migration method in the above embodiment belong to the same inventive concept, and further illustrates the process of determining the mapping relationship of storage architecture description information. The method can be executed by a data migration device, which can be implemented by software and / or hardware and integrated into a computer device with application development function.
[0061] like Figure 3 As shown, the data migration method of this embodiment includes the following steps:
[0062] S310: Obtain first data storage architecture description information in the first data management system and second data storage architecture description information in the second data management system.
[0063] The first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organization architecture of the corresponding data management system.
[0064] The first data storage architecture description information, i.e., the historical department name information (such as "Department of Cardiology", "Department of Radiology - Imaging Center"), is exported from the old system, i.e., the first data management system, and the second data storage architecture description information, i.e., the current department name information (such as "Department of Cardiology", "Department of Medical Imaging"), is obtained from the new system, i.e., the second data management system.
[0065] In an optional implementation, after obtaining the first data storage architecture description information and the second data storage architecture description information, the first data storage architecture description information and the second data storage architecture description information are standardized, wherein the standardized processing includes at least one of the following steps A1-A4:
[0066] Step A1: unify the characters in the first data storage architecture description information and the second data storage architecture description information into uppercase characters or lowercase characters.
[0067] For example, in programming languages such as Python, the str.upper function converts a string to uppercase, and the str.lower function converts it to lowercase.
[0068] Step A2: remove special characters included in the first data storage architecture description information and the second data storage architecture description information, wherein the special characters are characters in a preset special character dictionary.
[0069] A preset special character dictionary is pre-set, and characters included in the preset special character dictionary are removed from the first data storage architecture description information and the second data storage architecture description information by means of a regular expression module or the like.
[0070] Step A3: Identify and correct misspelled contents in the first data storage architecture description information and the second data storage architecture description information.
[0071] Misspellings can be identified and corrected using a programming language or online spell-checking tool.
[0072] Step A4: modify the abbreviation information in the first data storage architecture description information and the second data storage architecture description information into corresponding non-abbreviation information content.
[0073] You can use toolkits such as NLTK (Natural Language Toolkit) in combination with corpora and statistical models to identify abbreviation information and expand the abbreviation information to obtain the corresponding non-abbreviated information content, such as expanding the abbreviation information "heart" into the non-abbreviated information content "cardiovascular medicine".
[0074] This embodiment uses the above-mentioned data standardization processing operation to improve the accuracy of subsequent model predictions and ensure the effect of data migration.
[0075] S320 . Extracting semantic feature information from the first data storage architecture description information and the second data storage architecture description information through a pre-trained semantic recognition model.
[0076] S330. Input the semantic feature information corresponding to each description information in the first data storage architecture description information into a pre-trained architecture description matching model to obtain a corresponding model output result.
[0077] The pre-trained architecture description matching model can be a classification model and a clustering algorithm, etc. The classification model and the neural network are supervised learning. In the pre-training stage, the labeled data representing the mapping relationship between the first data storage architecture description information and the second data storage architecture description information, such as the correspondence between the new and old departments, can be obtained. The classification model, such as a random forest, XGBoost or a neural network model, can be trained to directly predict the model output result, that is, the text information matched by the first data storage architecture description information predicted by the model, such as the name of the new system department.
[0078] Clustering algorithms, such as density-based spatial clustering algorithm DBSCAN (Density-Based Spatial Clustering of Applications with Noise) or text similarity algorithms, such as cosine similarity and Levenshtein distance, are used to classify similar first data storage architecture description information and text information determined based on the second data storage architecture description information into one category to obtain a data storage architecture description information group. It can also be combined with a rule engine, such as regular expression matching of specific patterns and model output, to improve coverage and accuracy.
[0079] During the model training phase, the cross-validation method can be used to divide the training set and the test set, evaluate the model's accuracy, recall rate and other indicators, and determine the optimal model as the pre-trained architecture description matching model based on the indicator evaluation values.
[0080] S340. Match information in the second data storage architecture description information according to the model output result, and determine a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information according to the matching result.
[0081] When the model output result is predicted text information, the predicted text information is directly matched with the second data storage architecture description information to determine that the matched second data storage architecture description information is description information corresponding to the description information in the first data storage architecture description information.
[0082] When the model output result is a data storage architecture description information group obtained by classification, determine the data storage architecture description information group where the description information in the first data storage architecture description information is located, match the text information in the group that has the highest similarity to the description information in the first data storage architecture description information with the second data storage architecture description information, and determine that the matched second data storage architecture description information is the description information corresponding to the description information in the first data storage architecture description information.
[0083] In an optional implementation, when the confidence of the model output result is less than a preset confidence lower limit threshold, the model output result is displayed on a preset mapping relationship editing operation page to determine the corresponding mapping relationship in response to the acquired mapping relationship editing operation.
[0084] For each description information of the first data storage architecture description information, that is, each old department name, the model outputs the first preset number of candidate matching results and calculates the confidence score. For example, "Radiology Department-Imaging Center" may match "Medical Imaging Department" (confidence 95%) or "Radiotherapy Department" (confidence 60%). High-confidence results are automatically mapped, and low-confidence results are transferred to the manual review queue. The model output results with confidence less than the preset confidence lower limit threshold are displayed on the preset mapping relationship editing operation page, and the corresponding mapping relationship is determined according to the manual mapping relationship editing operation. For example, manually select the correct new department or add a new mapping rule. The manually corrected data can be fed back to the semantic recognition model as a training set to continuously optimize the algorithm (such as online learning or incremental training). The mapping rule can be a rule to ensure that the matching result conforms to the hospital organizational structure (such as "Pediatrics" should not be mapped to "Geriatrics"), and unreasonable matches can be filtered through the rule base.
[0085] Adjust model parameters according to the actual migration effect, or introduce more advanced NLP models (such as BERT fine-tuned based on domain knowledge).
[0086] S350: Migrate the system data in the first data management system to the data directory of the corresponding description information in the second data management system according to the mapping relationship.
[0087] In this embodiment, the system data in the migrated first data management system can also be standardized using the standardized processing method in step S310 to improve the data quality of the migrated system data and reduce data errors.
[0088] The technical solution of this embodiment is to obtain the first data storage architecture description information in the first data management system and the second data storage architecture description information in the second data management system; wherein the first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organization architecture of the corresponding data management system; extract the semantic feature information in the first data storage architecture description information and the second data storage architecture description information through a pre-trained semantic recognition model, input the semantic feature information corresponding to each description information in the first data storage architecture description information into the pre-trained architecture description matching model, and obtain the corresponding model output result; perform information matching in the second data storage architecture description information according to the model output result, and determine the mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information according to the matching result. According to the mapping relationship, the system data in the first data management system is migrated to the data directory of the corresponding description information in the second data management system. The technical solution of the embodiment of the present invention solves the problem of the difficulty of data migration in complex data systems at present, and can predict the semantic feature information through a pre-trained architecture description matching model, and determine the mapping relationship of the data storage architecture description information between the two systems for data migration according to the prediction result, thereby improving the accuracy and efficiency of data migration.
[0089] Figure 4 A structural schematic diagram of a data migration device provided in an embodiment of the present invention. This embodiment can be applied to data migration scenarios, especially data migration for medical institutions. The data migration device can be implemented by software and / or hardware and integrated into a computer terminal device with application development function.
[0090] like Figure 4 The data migration device shown includes: a data acquisition module 410 , a mapping relationship determination module 420 and a data migration module 430 .
[0091] Among them, the data acquisition module 410 is used to obtain the first data storage architecture description information in the first data management system and the second data storage architecture description information in the second data management system; wherein the first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organizational structure of the corresponding data management system; the mapping relationship determination module 420 is used to extract the semantic feature information in the first data storage architecture description information and the second data storage architecture description information through a pre-trained semantic recognition model, and establish a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information based on the semantic feature information; the data migration module 430 is used to migrate the system data in the first data management system to the data directory of the corresponding description information in the second data management system according to the mapping relationship.
[0092] The technical solution of this embodiment is to obtain the first data storage architecture description information in the first data management system and the second data storage architecture description information in the second data management system; wherein the first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organizational structure of the corresponding data management system; extract the semantic feature information in the first data storage architecture description information and the second data storage architecture description information through a pre-trained semantic recognition model, and establish a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information based on the semantic feature information; according to the mapping relationship, migrate the system data in the first data management system to the data directory of the corresponding description information in the second data management system. The technical solution of the embodiment of the present invention solves the problem of difficulty in data migration of complex data systems at present, and can determine the mapping relationship between the data storage architecture description information of two different systems through natural language processing technology, thereby improving the accuracy and efficiency of data migration.
[0093] In an optional implementation, the mapping relationship determination module 420 is specifically configured to:
[0094] The first data storage architecture description information and the second data storage architecture description information are analyzed and processed by the word segmentation module of the semantic recognition model to obtain information segmentation results; keywords in the information segmentation results are extracted by the keyword extraction module of the semantic recognition model to obtain information keyword extraction results; vector encoding and semantic extraction are performed on the information keyword extraction results by the semantic extraction module of the semantic recognition model to obtain corresponding semantic feature information.
[0095] In an optional implementation, the mapping relationship determination module 420 is further configured to:
[0096] Calculate the semantic similarity between the semantic feature information of each description information in the first data storage architecture description information and the semantic feature information of each description information in the second data storage architecture description information respectively; determine the semantic feature information based on the numerical value of the semantic similarity to establish a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information.
[0097] In an optional implementation, the mapping relationship determination module 420 is further configured to:
[0098] According to the organizational structure level information corresponding to the description information where each semantic feature information is located, the corresponding semantic similarity is corrected to obtain the corrected semantic similarity.
[0099] In an optional implementation, the mapping relationship determination module 420 is further configured to:
[0100] The semantic feature information corresponding to each description information in the first data storage architecture description information is input into a pre-trained architecture description matching model to obtain a corresponding model output result; information matching is performed in the second data storage architecture description information according to the model output result, and a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information is determined according to the matching result.
[0101] In an optional implementation, the mapping relationship determination module 420 is further configured to:
[0102] When the confidence of the model output result is less than the preset confidence lower limit threshold, the model output result is displayed on a preset mapping relationship editing operation page to determine the corresponding mapping relationship in response to the acquired mapping relationship editing operation.
[0103] In an optional embodiment, the device further comprises:
[0104] A description information standardization module is used to standardize the first data storage architecture description information and the second data storage architecture description information; wherein the standardization processing includes at least one of the following: unifying the characters in the first data storage architecture description information and the second data storage architecture description information into uppercase characters or lowercase characters; removing special characters contained in the first data storage architecture description information and the second data storage architecture description information, wherein the special characters are characters in a preset special character dictionary; identifying and correcting misspelled content in the first data storage architecture description information and the second data storage architecture description information; modifying the abbreviation information in the first data storage architecture description information and the second data storage architecture description information into corresponding non-abbreviated information content.
[0105] The data migration device provided in the embodiment of the present invention can execute the data migration method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0106] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 5 A block diagram of an exemplary computer device 12 suitable for use in implementing embodiments of the present invention is shown. Figure 5 The computer device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention. The computer device 12 can be any terminal device with computing capabilities, such as an intelligent controller and server, a mobile phone and other terminal devices.
[0107] like Figure 5 As shown, the computer device 12 is in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including the system memory 28 and the processing unit 16).
[0108] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. By way of example, these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0109] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0110] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 5 not shown, usually called a "hard drive"). Although Figure 5Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The system memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.
[0111] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods of the embodiments described herein.
[0112] The computer device 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, etc.), one or more devices that enable a user to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the computer device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with the other modules of the computer device 12 via the bus 18. It should be understood that although Figure 5 Not shown, other hardware and / or software modules may be used in conjunction with computer device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0113] The processing unit 16 executes various functional applications and data processing by running the program stored in the system memory 28, for example, implementing the data migration method provided in the embodiment of the present invention, which includes:
[0114] Obtaining first data storage architecture description information in the first data management system and second data storage architecture description information in the second data management system; wherein the first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organizational architectures of the corresponding data management systems;
[0115] Extracting semantic feature information from the first data storage architecture description information and the second data storage architecture description information through a pre-trained semantic recognition model, and establishing a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information based on the semantic feature information;
[0116] According to the mapping relationship, the system data in the first data management system is migrated to the data directory of the corresponding description information in the second data management system.
[0117] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the data migration method provided in any embodiment of the present invention is implemented. The method includes:
[0118] Obtaining first data storage architecture description information in the first data management system and second data storage architecture description information in the second data management system; wherein the first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organizational architectures of the corresponding data management systems;
[0119] Extracting semantic feature information from the first data storage architecture description information and the second data storage architecture description information through a pre-trained semantic recognition model, and establishing a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information based on the semantic feature information;
[0120] According to the mapping relationship, the system data in the first data management system is migrated to the data directory of the corresponding description information in the second data management system.
[0121] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.
[0122] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0123] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0124] Computer program code for performing the operation of the present invention may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect through the Internet).
[0125] It should be understood by those skilled in the art that the modules or steps of the present invention described above can be implemented by a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, optionally, they can be implemented by a program code executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0126] An embodiment of the present disclosure further provides a computer program product, including a computer program, which, when executed by a processor, implements a data migration method as provided in any one of the embodiments of the present disclosure.
[0127] In the process of implementation, the computer program product can be written in one or more programming languages or a combination thereof to perform the computer program code for the disclosed operation, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect through the Internet).
[0128] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A data migration method, characterized in that: include: Obtaining first data storage architecture description information in a first data management system and second data storage architecture description information in a second data management system; wherein the first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organizational architectures of the corresponding data management systems; Extracting semantic feature information in the first data storage architecture description information and the second data storage architecture description information through a pre-trained semantic recognition model, and establishing a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information based on the semantic feature information; According to the mapping relationship, the system data in the first data management system is migrated to the data directory of the corresponding description information in the second data management system.
2. The method according to claim 1, characterized in that The extracting semantic feature information from the first data storage architecture description information and the second data storage architecture description information by using a pre-trained semantic recognition model includes: The first data storage architecture description information and the second data storage architecture description information are analyzed and processed by the word segmentation module of the semantic recognition model to obtain an information word segmentation result; Extracting keywords from the information segmentation result by using the keyword extraction module of the semantic recognition model to obtain an information keyword extraction result; The semantic extraction module of the semantic recognition model performs vector encoding and semantic extraction on the information keyword extraction result to obtain corresponding semantic feature information.
3. The method according to claim 2, characterized in that The establishing, based on the semantic feature information, a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information includes: respectively calculating the semantic similarity between the semantic feature information of each piece of description information in the first data storage architecture description information and the semantic feature information of each piece of description information in the second data storage architecture description information; The semantic feature information is determined according to the value of the semantic similarity to establish a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information.
4. The method according to claim 3, characterized in that After calculating the semantic similarity between the semantic feature information of each description information in the first data storage architecture description information and the semantic feature information of each description information in the second data storage architecture description information, the method further includes: According to the organizational structure level information corresponding to the description information where each semantic feature information is located, the corresponding semantic similarity is corrected to obtain a corrected semantic similarity.
5. The method according to claim 1, characterized in that The establishing, based on the semantic feature information, a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information includes: Inputting the semantic feature information corresponding to each description information in the first data storage architecture description information into a pre-trained architecture description matching model to obtain a corresponding model output result; Information matching is performed in the second data storage architecture description information according to the model output result, and a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information is determined according to the matching result.
6. The method according to claim 5, characterized in that The method further comprises: When the confidence of the model output result is less than a preset confidence lower limit threshold, the model output result is displayed on a preset mapping relationship editing operation page to determine the corresponding mapping relationship in response to the acquired mapping relationship editing operation.
7. The method according to claim 1, characterized in that After acquiring the first data storage architecture description information and the second data storage architecture description information, the method further includes: Performing standardization processing on the first data storage architecture description information and the second data storage architecture description information; The standardization process includes at least one of the following: Unifying the characters in the first data storage architecture description information and the second data storage architecture description information into uppercase characters or lowercase characters; removing special characters included in the first data storage architecture description information and the second data storage architecture description information, wherein the special characters are characters in a preset special character dictionary; Identify and correct misspelled contents in the first data storage architecture description information and the second data storage architecture description information; The abbreviation information in the first data storage architecture description information and the second data storage architecture description information is modified into corresponding non-abbreviation information content.
8. A data migration device, characterized in that: include: A data acquisition module, used to acquire first data storage architecture description information in a first data management system and second data storage architecture description information in a second data management system; wherein the first data storage architecture description information and the second data storage architecture description information respectively correspond to the system organizational architectures of the corresponding data management systems; A mapping relationship determination module, used to extract semantic feature information in the first data storage architecture description information and the second data storage architecture description information through a pre-trained semantic recognition model, and establish a mapping relationship between each description information in the first data storage architecture description information and the description information in the second data storage architecture description information based on the semantic feature information; The data migration module is used to migrate the system data in the first data management system to the data directory of the corresponding description information in the second data management system according to the mapping relationship.
9. A computer device, characterized in that: The computer device comprises: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data migration method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the data migration method as described in any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the data migration method according to any one of claims 1 to 7.