Metadata-based Model Reconstruction Method, Apparatus, Electronic Device, and Storage Medium

Through the metadata-based model reconstruction method, the company's existing business models are automatically screened and reconstructed, and the problems of repeated development and high maintenance costs in the enterprise's digital transformation are solved, achieving the effect of cost reduction and extension of the result period.

CN114416174BActive Publication Date: 2025-06-17PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210078236.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-22
Publication Date
2025-06-17
Estimated Expiration
2042-01-22

AI Technical Summary

Technical Problem

In the prior art, due to the lack of standardized business model development in the process of digital transformation, enterprises have resulted in repeated development, waste of computing resources, high development and maintenance costs, and low data assets, which in turn has difficulties and risks of giving up halfway through digital transformation.

Method used

The metadata-based model reconstruction method is adopted to extract feature fields from the metadata of the model to be reconstructed, determine the business domain and reconstruction template, give priority to target fields with high frequency, and determine their standard processing logic to realize automatic screening and reconstruction of existing business models.

Benefits of technology

Model reconstruction can be achieved without a large number of professional model designers, reduce reconstruction costs, extend the model achievement period, and extend the reconstructed model achievement period through regular inspections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114416174B_ABST
    Figure CN114416174B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence technology, and specifically discloses a model reconstruction method, device, electronic device and storage medium based on metadata. Among them, the method includes: extracting feature fields from the metadata corresponding to the model to be reconstructed to obtain at least one feature field, where each feature field in the at least one feature field is used to identify the characteristics of the business corresponding to the model to be reconstructed; determining the business domain of the model to be reconstructed according to the at least one feature field; determining the reconstruction template of the model to be reconstructed according to the business domain; determining at least one target field in the at least one feature field, where the occurrence frequency of each target field in the at least one target field is greater than a first threshold; determining the standard processing logic of each target field; reconstructing the model to be reconstructed according to the reconstruction template and the standard processing logic of each target field to obtain a reconstructed model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a model reconstruction method, device, electronic device and storage medium based on metadata. Background Art

[0002] With the advent of the digital society, data, as an asset of an enterprise, has become increasingly important for the survival and development of the enterprise. In this context, the digital transformation of enterprises has become an essential step in the development of modern enterprises. However, in the initial stage of enterprise development, due to the rapid development needs of business, the development of business models usually has no standards at this time. Generally, developers develop business models according to their own habits. As a result, there are a large number of duplicate developments (chimney developments) in the enterprise system. The consequences caused by this chimney development are as follows: waste of computing resources, that is, the same indicator exists in multiple data processing tasks; high development and maintenance costs, that is, the same indicator requires developers to develop repeatedly, and a large number of tasks also result in high personnel maintenance costs; low value of data assets, that is, due to different statistical calibers and logics for the same indicator name, it causes troubles to data users and is not conducive to the precipitation of data assets. Therefore, the current non-standard business models of historical stocks make it very difficult for enterprises to carry out digital transformation now, and it is easy to cause the digital transformation of enterprises to fall by the wayside.

[0003] In view of the above situation, the existing processing method is to arrange personnel to re-comb the models, analyze the existing indicators, and then reconstruct the models through corresponding modeling methods. However, the defects and deficiencies of this solution are that a considerable number of professional model designers are required, resulting in relatively high model reconstruction costs. At the same time, due to the lack of corresponding measures in the later stage, the newly added data models soon need to be re-regulated and reconstructed again, resulting in a short achievement period for the reconstructed models. Summary of the Invention

[0004] In order to solve the above problems existing in the prior art, the embodiments of the present application provide a model reconstruction method, device, electronic device and storage medium based on metadata, which can realize the automatic screening and reconstruction of the existing business models in the system, without a large number of professional model designers, reduce the reconstruction cost, and at the same time can realize the regular inspection of the quality of existing models and extend the achievement period of the reconstructed models.

[0005] In a first aspect, an embodiment of the present application provides a model reconstruction method based on metadata, including:

[0006] Performing feature field extraction on the metadata corresponding to the model to be reconstructed to obtain at least one feature field, where each feature field in the at least one feature field is used to identify the features of the business corresponding to the model to be reconstructed;

[0007] Determine the business domain of the model to be reconstructed based on at least one feature field;

[0008] Determine the reconstruction template of the model to be reconstructed according to the business domain;

[0009] Determine at least one target field from the at least one feature field, where the occurrence frequency of each target field in the at least one target field is greater than the first threshold;

[0010] Determine the standard processing logic for each target field;

[0011] Reconstruct the model to be reconstructed according to the reconstruction template and the standard processing logic of each target field to obtain a reconstructed model.

[0012] In a second aspect, an embodiment of the present application provides a model reconstruction device based on metadata, including:

[0013] An extraction module, configured to extract feature fields from the metadata corresponding to the model to be reconstructed to obtain at least one feature field, where each feature field in the at least one feature field is used to identify the features of the business corresponding to the model to be reconstructed;

[0014] A processing module, configured to determine the business domain of the model to be reconstructed according to the at least one feature field, determine the reconstruction template of the model to be reconstructed according to the business domain, determine at least one target field from the at least one feature field, where the occurrence frequency of each target field in the at least one target field is greater than the first threshold, and determine the standard processing logic for each target field;

[0015] A reconstruction module, configured to reconstruct the model to be reconstructed according to the reconstruction template and the standard processing logic of each target field to obtain a reconstructed model.

[0016] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, the processor is connected to a memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device executes the method as in the first aspect.

[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and the computer program causes the computer to execute the method as in the first aspect.

[0018] In a fifth aspect, an embodiment of the present application provides a computer program product, the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer is operable to cause the computer to execute the method as in the first aspect.

[0019] Implementing the embodiments of the present application has the following beneficial effects:

[0020] In the embodiments of the present application, by obtaining the metadata of the model to be reconstructed, at least one feature field identifying the characteristics of the business corresponding to the model to be reconstructed is determined. Subsequently, according to the at least one feature field, the business domain corresponding to the reconstructed model is determined, and further the standard model template of the reconstructed model, that is, the reconstructed model, is determined, realizing the standardization of the reconstructed model. While ensuring the achievement period of the reconstructed model, after the standard changes later, it is also possible to complete the standardized changes of all models of the same type by modifying the reconstructed model once. Then, according to the occurrence frequency of each feature field in the at least one feature field, the feature fields with occurrence frequencies greater than the first threshold are used as the target fields to be preferentially considered during model reconstruction, thereby ensuring that the reconstructed model can meet the basic operation requirements of the business corresponding to the model. Finally, the standard processing logic of each target field is determined, and thus, according to the reconstruction template and the standard processing logic of each target field, the model to be reconstructed is reconstructed to obtain the reconstructed model. Thereby, the automatic reconstruction of the existing business models in the system is realized, without the need for a large number of professional model designers, reducing the cost of reconstruction. At the same time, through the method provided by this embodiment, the reconstructed models can also be regularly inspected to extend the achievement period of the reconstructed models. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 Schematic diagram of the hardware structure of a model reconstruction device based on metadata provided by the embodiments of the present application;

[0023] Figure 2 Schematic diagram of the flow of a model reconstruction method based on metadata provided by the embodiments of the present application;

[0024] Figure 3 Schematic diagram of the flow of a method for extracting feature fields from the metadata corresponding to the model to be reconstructed to obtain at least one feature field provided by the embodiments of the present application;

[0025] Figure 4 Schematic diagram of the flow of a method for determining the field group corresponding to each string among at least one second candidate field provided by the embodiments of the present application;

[0026] Figure 5 Schematic diagram of the flow of a method for determining the business domain of the model to be reconstructed according to the at least one feature field provided by the embodiments of the present application;

[0027] Figure 6 A functional block diagram of the functional modules of a model reconstruction device based on metadata provided by an embodiment of the present application;

[0028] Figure 7 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0030] The terms "first", "second", "third", and "fourth", etc. in the specification and claims of the present application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having", and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0031] Referring to "embodiments" herein means that the specific features, results, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0032] First, refer to Figure 1 , Figure 1 A schematic hardware structure diagram of a model reconstruction device based on metadata provided by an embodiment of the present application. The model reconstruction device 100 based on metadata includes at least one processor 101, a communication line 102, a memory 103, and at least one communication interface 104.

[0033] In this embodiment, the processor 101 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of this application.

[0034] The communication line 102 can include a path for transmitting information between the above components.

[0035] The communication interface 104 can be any transceiver-like device (such as an antenna, etc.) for communicating with other devices or communication networks, such as Ethernet, RAN, wireless local area networks (WLAN), etc.

[0036] The memory 103 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0037] In this embodiment, the memory 103 can exist independently and be connected to the processor 101 through the communication line 102. The memory 103 can also be integrated with the processor 101. The memory 103 provided by the embodiment of this application generally can have non-volatility. Among them, the memory 103 is used to store the computer execution instructions for executing the solution of this application and is controlled by the processor 101 to execute. The processor 101 is used to execute the computer execution instructions stored in the memory 103, so as to implement the method provided in the following embodiments of this application.

[0038] In an alternative embodiment, the computer execution instructions can also be referred to as application code, and this application does not make specific limitations thereto.

[0039] In an alternative embodiment, the processor 101 may include one or more CPUs, such as Figure 1 CPU0 and CPU1 in

[0040] In an alternative embodiment, the metadata-based model reconstruction apparatus 100 may include multiple processors, such as Figure 1 processor 101 and processor 107 in

[0041] Each of these processors may be a single-CPU processor or a multi-CPU processor. The processors herein may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0042] The above-described metadata-based model reconstruction apparatus 100 may be a general-purpose device or a special-purpose device. The embodiments of the present application do not limit the type of the metadata-based model reconstruction apparatus 100.

[0043] Secondly, it should be noted that the disclosed embodiments of the present application may acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0044] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0045] Hereinafter, a model reconstruction method based on metadata disclosed in the present application will be described:

[0046] Refer to Figure 2 , Figure 2 FIG. is a schematic flowchart of a model reconstruction method based on metadata provided for an implementation manner of the present application. The model reconstruction method based on metadata includes the following steps:

[0047] 201: Extract feature fields from the metadata corresponding to the model to be reconstructed to obtain at least one feature field.

[0048] In this embodiment, since a large number of system models have been accumulated in the system, and some of these system models are added for the current operation plan, while some are deprecated system models corresponding to historical operation plans. In other words, these deprecated system models are also stored in the model library of the system together with the system models that are still in use. However, since the operation plans corresponding to these deprecated system models have been completed and will not be called in the short term, reconstructing them will instead increase the workload and cause waste of reconstruction costs. Based on this, in this embodiment, before model reconstruction, it is also necessary to select system models with reconstruction value from the large number of system models accumulated in the system as the models to be reconstructed. Specifically, system models that are still used by the system within a time period based on the current time can be regarded as system models with reconstruction value.

[0049] Based on this, in this embodiment, it is possible to determine whether an existing model is in use by obtaining the occurrence time of the last data query or data usage of each system model in at least one system model accumulated in the system, and then determine the model to be reconstructed from at least one system model according to the occurrence time of the last data query or data usage of each system model. Exemplarily, a system model whose interval between the occurrence time of the last data query or data usage and the current time is less than or equal to a second threshold can be regarded as the model to be reconstructed. For example, if it is found that a certain system model in the system has not been accessed within 3 months, the model can be recorded in the model offline list, and at the same time, the model can be removed from the model reconstruction list because the model has no reconstruction value.

[0050] Meanwhile, in this embodiment, the metadata may refer to the business metadata in the system, which is used to store relevant information of the corresponding system model, such as: table name, field name, Chinese name of the field, table data update time, and other relevant information. Therefore, the dimensions of the system model, the usage frequency of the metrics, and which relevant business scenario dimension metrics there are can be intuitively seen from the business metadata. Based on this, at least one feature field for identifying the characteristics of the business corresponding to the model to be reconstructed can be obtained by extracting the feature fields from the business metadata.

[0051] Based on this, this embodiment provides a method for extracting feature fields from the metadata corresponding to the model to be reconstructed to obtain at least one feature field, as Figure 3 shown, the method includes:

[0052] 301: Determine at least one field name according to the model structure of the model to be reconstructed.

[0053] In this embodiment, each of the at least one field names is used to identify the naming information of the corresponding field in the model to be reconstructed. Specifically, for different models, the storage locations of the important fields in the model can be determined according to their corresponding business attributes, operation logics, model structures, etc., and then the names of these locations are obtained as field names.

[0054] 302: Determine at least one string corresponding one-to-one to the at least one field name in the metadata according to the at least one field name.

[0055] In this embodiment, since the business metadata is used to store relevant information of the corresponding system model, such as: table name, field name, Chinese name of the field, table data update time, and other relevant information. Therefore, at least one string corresponding to each field name in the at least one field name can be obtained in the business metadata by means of field name matching, and at least one string is obtained.

[0056] 303: Perform text segmentation processing on each of the at least one strings to obtain at least one field group corresponding one-to-one to the at least one strings.

[0057] In this embodiment, first, each string can be subjected to text segmentation processing to obtain at least one first candidate field corresponding to each string. Exemplarily, a delimiter set can be preset first, which includes some delimiter characters commonly used in Chinese, such as punctuation marks, special symbols, diagrams, conjunctions, stop words, etc. Then, each string is matched with the delimiter set, and the delimiters existing in each string are replaced with spaces, and each string is subjected to text segmentation processing to obtain at least one substring. Then, each substring in the at least one substring is respectively subjected to forward maximum matching with a general word segmentation dictionary. When a substring matches a word in the dictionary successfully, the successfully matched word in the substring is extracted to obtain at least one first candidate field corresponding to the string.

[0058] Then, the part-of-speech information of each first candidate field in the at least one first candidate field can be determined, and thus, at least one second candidate field can be determined from the at least one first candidate field according to the part-of-speech information of each first candidate field. In this embodiment, the part-of-speech information may refer to information describing the nature of a field, such as a verb, a noun, a named entity, etc., which can be determined by analyzing the sentence pattern and semantics of each first candidate field, and then, the first candidate fields with a certain or multiple part-of-speech information are screened out from the at least one first candidate field as the at least one second candidate field. Exemplarily, the candidate condition can be set to that the part-of-speech information is a named entity. Thus, the first candidate fields with the part-of-speech information of named entity in the at least one first candidate field are screened out as the second candidate fields.

[0059] Finally, in the at least one second candidate field, a field group corresponding to each string is determined to obtain at least one field group. In this embodiment, since the fields screened by the word segmentation method may include long fields that are split into multiple parts during word segmentation. And the overall meaning of the long field may not be the same as, or even conflict with, the meanings of the multiple fields split out. Therefore, in order to obtain a field group that can represent the accurate meaning of each character segment, these split long fields need to be found again and used to replace the multiple fields split out. Based on this, this embodiment provides a method for determining a field group corresponding to each string in the at least one second candidate field, as Figure 4 shown, the method includes:

[0060] 401: Combine the first adjacent field and the second adjacent field in the at least one second candidate field to obtain at least one third candidate field.

[0061] In this embodiment, the first adjacent field and the second adjacent field are any two different second candidate fields, and the field interval between the first adjacent field and the second adjacent field is less than the first threshold. Specifically, the first adjacent field and the second adjacent field are two adjacent fields among the second candidate fields with a field interval less than the first threshold, and this field interval can be understood as the number of characters between the corresponding positions of the first adjacent field and the second adjacent field in their corresponding strings. Exemplarily, for the string "I graduated from Fudan University in Shanghai in 2021", after word segmentation and screening, the second candidate fields can be obtained: "in 2021", "Shanghai", "Fudan", and "University". At this time, the number of characters between the corresponding positions of the second candidate fields "in 2021" and "Shanghai" in the original string is 3, so the character distance between the second candidate fields "in 2021" and "Shanghai" is 3. And the number of characters between the corresponding positions of the second candidate fields "Fudan" and "University" in the original string is 0, so the character distance between the second candidate fields "Fudan" and "University" is 0.

[0062] In this embodiment, the first threshold can be set to 2. Thus, taking the above string "I graduated from Fudan University in Shanghai in 2021" as an example, the second candidate fields that meet the requirements are: "Shanghai" and "Fudan", and "Fudan" and "University". Thus, the third candidate fields "Shanghai Fudan" and "Fudan University" can be obtained.

[0063] 402: Perform semantic extraction on each of the at least one third candidate field to obtain at least one semantic vector corresponding one-to-one to the at least one third candidate field.

[0064] 403: Determine at least one fourth candidate field from the at least one third candidate field according to the at least one semantic vector.

[0065] In this embodiment, the similarity between the semantic vector and the standard vector corresponding to its semantics can be calculated, and when the calculated similarity is greater than the preset threshold, the third candidate field of the semantic vector corresponding to this similarity is used as the fourth candidate field.

[0066] 404: In the at least one second candidate field, delete the second candidate fields that make up each of the at least one fourth candidate field to obtain at least one fifth candidate field.

[0067] In this embodiment, the fifth candidate field is the remaining second candidate field after removing the second candidate fields that make up each of the at least one fourth candidate field. Exemplarily, continuing with the above example of the string "I graduated from Fudan University in Shanghai in 2021", assuming that the fourth candidate field determined by semantic similarity calculation is "Fudan University", since the fourth candidate field "Fudan University" is composed of the second candidate fields "Fudan" and "University", therefore, the second candidate fields "Fudan" and "University" are removed from the originally obtained several second candidate fields: "2021", "Shanghai", "Fudan", and "University", and the remaining second candidate fields "2021" and "Shanghai" are the fifth candidate fields.

[0068] 405: Combine at least one fourth candidate field and at least one fifth candidate field to obtain a field group corresponding to each string.

[0069] Exemplarily, continuing with the above example of the string "I graduated from Fudan University in Shanghai in 2021", combine the fourth candidate field "Fudan University" with the fifth candidate fields "2021" and "Shanghai" to obtain the field group corresponding to the string "I graduated from Fudan University in Shanghai in 2021": "2021", "Shanghai", and "Fudan University".

[0070] 304: Determine at least one characteristic field in at least one field group.

[0071] In this embodiment, all the fields in at least one field group can be summarized and de-duplicated, and then at least one characteristic field is obtained.

[0072] 202: Determine the business domain of the model to be reconstructed according to at least one characteristic field.

[0073] In this embodiment, due to the different processing requirements and purposes of different services, there are also certain differences in their processing logics. Therefore, each service in the system can be classified, and some services with similar processing logics can be grouped under the same business domain. Thus, by determining the business domain of the model to be reconstructed, the general processing logic of the model to be reconstructed can be determined.

[0074] At the same time, in this embodiment, since the at least one characteristic field is obtained by extracting the business metadata storing the relevant information of the corresponding system model. Therefore, the at least one characteristic field can comprehensively represent the characteristics of the service corresponding to the model to be reconstructed. Therefore, this embodiment provides a method for determining the business domain of the model to be reconstructed according to the at least one characteristic field. As Figure 5 shown, the method includes:

[0075] 501: Determine the business domain groups corresponding to each feature field in the at least one feature field, and obtain at least one business domain group that corresponds one-to-one with the at least one feature field.

[0076] In this embodiment, the business metadata of the historical model can be analyzed to obtain a business domain table, which records one or more business domains corresponding to each feature field. For example, the feature field "AUTO" can correspond to business domains related to vehicles, such as: auto loan, auto insurance, vehicle claim settlement, vehicle mortgage, etc. Thus, by querying this business domain table, the business domain group corresponding to each feature field can be obtained.

[0077] 502: Statistically analyze the at least one business domain group to determine the scores of the business domains included in each business domain group in the at least one business domain group.

[0078] Exemplarily, the at least one business domain group can be scanned in sequence. Each time a business domain is scanned, the corresponding score is incremented by 1 until all the business domain groups are scanned, and the scores of the business domains included in each business domain group are obtained. Specifically, assume that there are now 3 business domain groups, namely: Business Domain Group 1 [auto insurance, claim settlement, mortgage], Business Domain Group 2 [claim settlement, loan], and Business Domain Group 3 [claim settlement, auto insurance, purchase]. After statistically analyzing the above 3 business domain groups, it is obtained that auto insurance appears 2 times and is recorded as 2 points; claim settlement appears 3 times and is recorded as 3 points; mortgage appears 1 time and is recorded as 1 point; loan appears 1 time and is recorded as 1 point; purchase appears 1 time and is recorded as 1 point.

[0079] 503: Use the business domain with the highest score as the business domain of the model to be reconstructed.

[0080] Specifically, continuing with the above example, since the score of claim settlement is the highest, at 3 points, therefore, the business domain of the model to be reconstructed is the claim settlement business domain.

[0081] 203: Determine the reconstruction template of the model to be reconstructed according to the business domain.

[0082] In this embodiment, since there are certain similarities in the processing logics of the businesses under the same business domain, therefore, the common processing logic of each business domain can be extracted, and combined with some pain points of this business domain to generate the corresponding reconstruction template. Thus, when performing model reconstruction, by invoking this common reconstruction template, some common operations can be quickly generated, reducing the model reconstruction time and cost, and improving the model reconstruction efficiency.

[0083] 204: Determine at least one target field in the at least one feature field.

[0084] In this embodiment, the occurrence frequency of each target field in the at least one target field is greater than a first threshold. Exemplarily, the target field is a field that appears more frequently in the business metadata of the corresponding system model, which indicates that this field is a field that is frequently invoked or used by the system model and has a relatively high importance to the system model. Therefore, these fields can be preferentially considered as the basic fields for reconstructing the model.

[0085] 205: Determine the standard processing logic for each target field.

[0086] In this embodiment, the metadata database can be retrieved according to each target field to obtain at least one piece of business metadata. Then, based on the at least one piece of business metadata, the corresponding system model library for each piece of business metadata can be found to obtain at least one system model. Next, in each of the at least one system models, the processing logic corresponding to each target field can be determined to obtain at least one candidate processing logic that corresponds one-to-one with the at least one system model. Finally, by determining the proportion of each candidate processing logic in the at least one candidate processing logic, the candidate processing logic with the highest proportion is used as the standard processing logic for each target field.

[0087] Specifically, assume that n candidate processing logics are finally obtained. When n = 1, it means that for this target field, the processing logics of all system models are the same. Therefore, this candidate processing logic can be directly used as the standard processing logic for this field. When n > 1, it means that there are multiple processing logics for this target field in the system. At this time, by counting the proportion of each candidate processing logic in the n candidate processing logics, the candidate processing logic with the highest proportion is selected as the standard processing logic for this target field.

[0088] 206: Reconstruct the model to be reconstructed according to the reconstruction template and the standard processing logic of each target field to obtain a reconstructed model.

[0089] In this embodiment, the processing order of the standard processing logic corresponding to each target field can be determined through the reconstruction template, and then the corresponding process chain can be obtained. Then, according to this process chain, the standard processing logic corresponding to each target field is filled into the corresponding position in the reconstruction template to obtain a reconstructed model.

[0090] In summary, in the method for model reconstruction based on metadata provided by the present invention, at least one feature field for identifying the characteristics of the business corresponding to the model to be reconstructed is determined by obtaining the metadata of the model to be reconstructed. Then, according to the at least one feature field, the business domain corresponding to the reconstructed model is determined, and further the standard model template for the reconstructed model, that is, the reconstructed model, is determined, realizing the standardization of the reconstructed model. While ensuring the achievement period of the reconstructed model, after the standard changes later, it is also possible to complete the standardized change of all models of the same type by making a single modification to the reconstructed model. Then, according to the occurrence frequency of each feature field in the at least one feature field, the feature fields with an occurrence frequency greater than the first threshold are used as the target fields to be preferentially considered during model reconstruction, thereby ensuring that the reconstructed model can meet the basic operation requirements of the business corresponding to the model. Finally, the standard processing logic for each target field is determined, so that the model to be reconstructed is reconstructed according to the reconstruction template and the standard processing logic of each target field to obtain the reconstructed model. Thus, the automatic reconstruction of the existing business models in the system is realized, without the need for a large number of professional model designers, reducing the cost of reconstruction. At the same time, the method provided by this embodiment can also regularly inspect the reconstructed models to extend the achievement period of the reconstructed models.

[0091] Refer to Figure 6 , Figure 6 which is a functional module composition block diagram of a model reconstruction device based on metadata provided by an embodiment of the present application. As Figure 6 shown, the model reconstruction device 600 based on metadata includes:

[0092] An extraction module 601, configured to extract feature fields from the metadata corresponding to the model to be reconstructed to obtain at least one feature field, where each feature field in the at least one feature field is used to identify the characteristics of the business corresponding to the model to be reconstructed;

[0093] A processing module 602, configured to determine the business domain of the model to be reconstructed according to the at least one feature field, determine the reconstruction template of the model to be reconstructed according to the business domain, determine at least one target field in the at least one feature field, where the occurrence frequency of each target field in the at least one target field is greater than the first threshold, and determine the standard processing logic for each target field;

[0094] A reconstruction module 603, configured to reconstruct the model to be reconstructed according to the reconstruction template and the standard processing logic of each target field to obtain the reconstructed model.

[0095] In an embodiment of the present invention, in terms of extracting feature fields from the metadata corresponding to the model to be reconstructed to obtain at least one feature field, the extraction module 601 is specifically configured to:

[0096] Determine at least one field name according to the model structure of the model to be reconstructed, where each field name in the at least one field name is used to identify the naming information of the corresponding field in the model to be reconstructed;

[0097] Determine at least one string in the metadata according to the at least one field name, where the at least one string corresponds to the at least one field name one by one;

[0098] Perform text segmentation processing on each string in the at least one string to obtain at least one field group, where the at least one field group corresponds to the at least one string one by one;

[0099] Determine at least one characteristic field in the at least one field group.

[0100] In an embodiment of the present invention, in terms of performing text segmentation processing on each string in the at least one string to obtain at least one field group, the extraction module 601 is specifically configured to:

[0101] Perform text segmentation processing on each string to obtain at least one first candidate field corresponding to each string;

[0102] Determine the part-of-speech information of each first candidate field in the at least one first candidate field;

[0103] Determine at least one second candidate field in the at least one first candidate field according to the part-of-speech information of each first candidate field;

[0104] In the at least one second candidate field, determine the field group corresponding to each string to obtain at least one field group.

[0105] In an embodiment of the present invention, in terms of determining the field group corresponding to each string in the at least one second candidate field, the extraction module 601 is specifically configured to:

[0106] Combine the first adjacent field and the second adjacent field in the at least one second candidate field to obtain at least one third candidate field, where the first adjacent field and the second adjacent field are any two different second candidate fields, and the field interval between the first adjacent field and the second adjacent field is less than the first threshold;

[0107] Perform semantic extraction on each third candidate field in the at least one third candidate field to obtain at least one semantic vector, where the at least one semantic vector corresponds to the at least one third candidate field one by one;

[0108] Determine at least one fourth candidate field in the at least one third candidate field according to the at least one semantic vector;

[0109] In at least one second candidate field, delete the second candidate fields that make up each fourth candidate field in at least one fourth candidate field to obtain at least one fifth candidate field;

[0110] Combine at least one fourth candidate field and at least one fifth candidate field to obtain a field group corresponding to each string.

[0111] In an embodiment of the present invention, the metadata-based model reconstruction device 600 may further include: a screening module (not shown), which is used to:

[0112] Obtain the occurrence time of the last data query or data usage of each system model in at least one system model;

[0113] According to the occurrence time of the last data query or data usage of each system model, determine the model to be reconstructed in at least one system model, where the interval between the occurrence time of the last data query or data usage corresponding to the model to be reconstructed and the current time is less than or equal to a second threshold.

[0114] In an embodiment of the present invention, in terms of determining the business domain of the model to be reconstructed according to at least one feature field, the processing module 602 is specifically used to:

[0115] Determine the business domain group corresponding to each feature field in at least one feature field to obtain at least one business domain group, where at least one business domain group and at least one feature field are in one-to-one correspondence;

[0116] Statistically analyze at least one business domain group to determine the score of the business domain included in each business domain group in at least one business domain group;

[0117] Use the business domain with the highest score as the business domain of the model to be reconstructed.

[0118] In an embodiment of the present invention, in terms of determining the standard processing logic of each target field, the processing module 602 is specifically used to:

[0119] Retrieve the metadata database according to each target field to obtain at least one business metadata;

[0120] Search the system model library according to at least one business metadata to obtain at least one system model, where at least one system model and at least one business metadata are in one-to-one correspondence;

[0121] In each system model in at least one system model, determine the processing logic corresponding to each target field to obtain at least one candidate processing logic, where at least one candidate processing logic and at least one system model are in one-to-one correspondence;

[0122] Determine the proportion of each candidate processing logic in at least one candidate processing logic, and use the candidate processing logic with the highest proportion as the standard processing logic for each target field.

[0123] See Figure 7 , Figure 7 FIG. 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 7 shown, the electronic device 700 includes a transceiver 701, a processor 702, and a memory 703. They are connected through a bus 704. The memory 703 is used to store computer programs and data, and can transmit the data stored in the memory 703 to the processor 702.

[0124] The processor 702 is configured to read the computer program in the memory 703 and perform the following operations:

[0125] Extract feature fields from the metadata corresponding to the model to be reconstructed to obtain at least one feature field, where each feature field in the at least one feature field is used to identify the features of the service corresponding to the model to be reconstructed;

[0126] Determine the business domain of the model to be reconstructed according to the at least one feature field;

[0127] Determine the reconstruction template of the model to be reconstructed according to the business domain;

[0128] Determine at least one target field in the at least one feature field, where the occurrence frequency of each target field in the at least one target field is greater than a first threshold;

[0129] Determine the standard processing logic for each target field;

[0130] Reconstruct the model to be reconstructed according to the reconstruction template and the standard processing logic of each target field to obtain a reconstructed model.

[0131] In an embodiment of the present invention, in terms of extracting feature fields from the metadata corresponding to the model to be reconstructed to obtain at least one feature field, the processor 702 is specifically configured to perform the following operations:

[0132] Determine at least one field name according to the model structure of the model to be reconstructed, where each field name in the at least one field name is used to identify the naming information of the corresponding field in the model to be reconstructed;

[0133] Determine at least one string in the metadata according to the at least one field name, where the at least one string corresponds to the at least one field name one by one;

[0134] Perform text segmentation processing on each of at least one string to obtain at least one field group, where the at least one field group corresponds one-to-one with the at least one string;

[0135] Determine at least one characteristic field in the at least one field group.

[0136] In an embodiment of the present invention, in performing text segmentation processing on each of at least one string to obtain at least one field group, the processor 702 is specifically configured to perform the following operations:

[0137] Perform text segmentation processing on each string to obtain at least one first candidate field corresponding to each string;

[0138] Determine the part-of-speech information of each first candidate field in the at least one first candidate field;

[0139] According to the part-of-speech information of each first candidate field, determine at least one second candidate field in the at least one first candidate field;

[0140] In the at least one second candidate field, determine the field group corresponding to each string to obtain at least one field group.

[0141] In an embodiment of the present invention, in determining the field group corresponding to each string in the at least one second candidate field, the processor 702 is specifically configured to perform the following operations:

[0142] Combine the first adjacent field and the second adjacent field in the at least one second candidate field to obtain at least one third candidate field, where the first adjacent field and the second adjacent field are any two different second candidate fields, and the field interval between the first adjacent field and the second adjacent field is less than the first threshold;

[0143] Perform semantic extraction on each third candidate field in the at least one third candidate field to obtain at least one semantic vector, where the at least one semantic vector corresponds one-to-one with the at least one third candidate field;

[0144] According to the at least one semantic vector, determine at least one fourth candidate field in the at least one third candidate field;

[0145] In the at least one second candidate field, delete the second candidate fields that make up each fourth candidate field in the at least one fourth candidate field to obtain at least one fifth candidate field;

[0146] Combine the at least one fourth candidate field and the at least one fifth candidate field to obtain the field group corresponding to each string.

[0147] In an embodiment of the present invention, before obtaining the metadata corresponding to the model to be reconstructed, the processor 702 is further configured to perform the following operations:

[0148] Obtain the occurrence time of the last data query or data usage of each system model in at least one system model;

[0149] According to the occurrence time of the last data query or data usage of each system model, determine the model to be reconstructed in at least one system model, wherein the interval between the occurrence time of the last data query or data usage corresponding to the model to be reconstructed and the current time is less than or equal to a second threshold.

[0150] In an embodiment of the present invention, in terms of determining the business domain of the model to be reconstructed according to at least one feature field, the processor 702 is specifically configured to perform the following operations:

[0151] Determine the business domain group corresponding to each feature field in at least one feature field to obtain at least one business domain group, wherein the at least one business domain group and the at least one feature field are in one-to-one correspondence;

[0152] Perform statistics on at least one business domain group to determine the score of each business domain included in each business domain group in at least one business domain group;

[0153] Use the business domain with the highest score as the business domain of the model to be reconstructed.

[0154] In an embodiment of the present invention, in terms of determining the standard processing logic of each target field, the processor 702 is specifically configured to perform the following operations:

[0155] Retrieve the metadata database according to each target field to obtain at least one business metadata;

[0156] Search the system model library according to at least one business metadata to obtain at least one system model, wherein the at least one system model and the at least one business metadata are in one-to-one correspondence;

[0157] In each system model in at least one system model, determine the processing logic corresponding to each target field to obtain at least one candidate processing logic, wherein the at least one candidate processing logic and the at least one system model are in one-to-one correspondence;

[0158] Determine the proportion of each candidate processing logic in at least one candidate processing logic, and use the candidate processing logic with the highest proportion as the standard processing logic of each target field.

[0159] It should be understood that the metadata-based model reconstruction device in the present application may include smartphones (such as Android phones, iOS phones, Windows Phone phones, etc.), tablet computers, handheld computers, laptop computers, mobile Internet devices MID (Mobile Internet Devices), robots, or wearable devices, etc. The above-mentioned metadata-based model reconstruction devices are only examples, not an exhaustive list, including but not limited to the above-mentioned metadata-based model reconstruction devices. In practical applications, the above-mentioned metadata-based model reconstruction device may also include: intelligent vehicle terminals, computer devices, and so on.

[0160] Through the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software combined with a hardware platform. Based on such an understanding, all or part of the technical solution of the present invention that contributes to the background art can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0161] Therefore, the embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement some or all of the steps of any one of the metadata-based model reconstruction methods described in the above method embodiments. For example, the storage medium may include a hard disk, a floppy disk, an optical disk, a magnetic tape, a magnetic disk, a USB flash drive, a flash memory, etc.

[0162] The embodiments of the present application also provide a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps of any one of the metadata-based model reconstruction methods described in the above method embodiments.

[0163] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0164] In the above embodiments, the descriptions of the various embodiments each have their own emphasis. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0165] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical or other forms.

[0166] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0167] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software program modules.

[0168] If the above-mentioned integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. And the aforementioned memory includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0169] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable memory, which may include: a flash drive, a read-only memory (abbreviation: ROM), a random access memory (abbreviation: RAM), a magnetic disk or an optical disc, etc.

[0170] The above has introduced the embodiments of the present application in detail. Specific examples are used herein to elaborate on the principle and embodiments of the present application. The description of the above embodiments is only for helping to understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific embodiments and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for model reconstruction based on metadata, characterized in that, The method includes: Determine at least one field name according to the model structure of the model to be reconstructed, where each field name in the at least one field name is used to identify the naming information of the corresponding field in the model to be reconstructed; Determine at least one string in the metadata corresponding to the model to be reconstructed according to the at least one field name, where the at least one string corresponds to the at least one field name one by one; Perform text segmentation processing on each string in the at least one string to obtain at least one first candidate field corresponding to each string; Determine the part-of-speech information of each first candidate field in the at least one first candidate field; Determine at least one second candidate field from the at least one first candidate field according to the part-of-speech information of each first candidate field; In the at least one second candidate field, determine the field group corresponding to each string to obtain at least one field group; Determine at least one feature field in the at least one field group, where each feature field in the at least one feature field is used to identify the features of the business corresponding to the model to be reconstructed; Determine the business domain of the model to be reconstructed according to the at least one feature field; Determine the reconstruction template of the model to be reconstructed according to the business domain; Determine at least one target field from the at least one feature field, where the occurrence frequency of each target field in the at least one target field is greater than a first threshold; Determine the standard processing logic of each target field; Reconstruct the model to be reconstructed according to the reconstruction template and the standard processing logic of each target field to obtain a reconstructed model.

2. The method according to claim 1, characterized in that, The determining, in the at least one second candidate field, the field group corresponding to each string includes: Combine a first adjacent field and a second adjacent field in the at least one second candidate field to obtain at least one third candidate field, where the first adjacent field and the second adjacent field are any two different second candidate fields, and the field interval between the first adjacent field and the second adjacent field is less than a first threshold; Perform semantic extraction on each third candidate field in the at least one third candidate field to obtain at least one semantic vector, where the at least one semantic vector corresponds to the at least one third candidate field one by one; Determine at least one fourth candidate field from the at least one third candidate field according to the at least one semantic vector; In the at least one second candidate field, delete the second candidate fields that make up each fourth candidate field in the at least one fourth candidate field to obtain at least one fifth candidate field; Combine the at least one fourth candidate field and the at least one fifth candidate field to obtain the field group corresponding to each string.

3. The method according to claim 1, characterized in that, Before determining at least one field name according to the model structure of the model to be reconstructed, the method further includes: Obtain the occurrence time of the last data query or data usage of each system model in at least one system model; Based on the occurrence time of the last data query or data usage of each system model, determine the model to be reconstructed in the at least one system model, where the interval between the occurrence time of the last data query or data usage corresponding to the model to be reconstructed and the current time is less than or equal to a second threshold.

4. The method according to claim 1, characterized in that, The determining the business domain of the model to be reconstructed according to the at least one feature field includes: Determine the business domain group corresponding to each feature field in the at least one feature field to obtain at least one business domain group, where the at least one business domain group and the at least one feature field are in one-to-one correspondence; Perform statistics on the at least one business domain group to determine the score of each business domain included in each business domain group in the at least one business domain group; Use the business domain with the highest score as the business domain of the model to be reconstructed.

5. The method according to claim 1, characterized in that, The determining the standard processing logic of each target field includes: Retrieve the meta database according to each target field to obtain at least one business metadata; Search the system model library according to the at least one business metadata to obtain at least one system model, where the at least one system model and the at least one business metadata are in one-to-one correspondence; In each system model of the at least one system model, determine the processing logic corresponding to each target field to obtain at least one candidate processing logic, where the at least one candidate processing logic and the at least one system model are in one-to-one correspondence; Determine the proportion of each candidate processing logic in the at least one candidate processing logic, and use the candidate processing logic with the highest proportion as the standard processing logic of each target field.

6. A device for model reconstruction based on metadata, characterized in that, The device includes: An extraction module, configured to determine at least one field name according to the model structure of the model to be reconstructed, where each field name in the at least one field name is used to identify the naming information of the corresponding field in the model to be reconstructed; determine at least one string in the metadata corresponding to the model to be reconstructed according to the at least one field name, where the at least one string and the at least one field name are in one-to-one correspondence; perform text segmentation processing on each string in the at least one string to obtain at least one first candidate field corresponding to each string; determine the part-of-speech information of each first candidate field in the at least one first candidate field; determine at least one second candidate field in the at least one first candidate field according to the part-of-speech information of each first candidate field; in the at least one second candidate field, determine the field group corresponding to each string to obtain at least one field group; determine at least one feature field in the at least one field group, where each feature field in the at least one feature field is used to identify the feature of the business corresponding to the model to be reconstructed; A processing module, configured to determine the business domain of the model to be reconstructed according to the at least one feature field, determine a reconstruction template of the model to be reconstructed according to the business domain, determine at least one target field from the at least one feature field, where the occurrence frequency of each target field in the at least one target field is greater than a first threshold, and determine the standard processing logic of each target field; A reconstruction module, configured to reconstruct the model to be reconstructed according to the reconstruction template and the standard processing logic of each target field, to obtain a reconstructed model.

7. An electronic device, characterized in that, It includes a processor, a memory, a communication interface, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the processor, and the one or more programs include instructions for executing the steps in the method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Metadata management method, system and storage medium

    CN111858584A

  • Source code processing method and device

    CN112379915A