Enterprise data evaluation and analysis system and method based on big data
By establishing a scene ontology grammar rule base and deploying a grammar perceptron to generate grammar fingerprints, and combining this with a bootstrap synthesizer to adapt to the new system, the problem of dynamic adaptation of grammar rules for multi-source heterogeneous enterprise data was solved. This achieved data structure normalization and semantic alignment, improving the quality and stability of data evaluation and analysis.
Patent Information
- Application Number
- CN202511489653.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-01-23
AI Technical Summary
Existing technologies struggle to dynamically adapt syntax rules to multi-source heterogeneous enterprise data, making it difficult for traditional static syntax parsing rules to adapt in real time to new systems or data format changes. This prevents the automatic syntax verification and structure normalization of multi-source heterogeneous data.
Establish a scenario ontology grammar rule library, decompose ERP, CRM, and supply chain management fields into grammar atoms, annotate entity relationship semantics according to business scenarios, deploy a grammar perceptron at the data collection end to generate grammar fingerprints, generate field naming, format mapping and association mapping based on grammar fingerprints and grammar atom templates, build a structure normalization pipeline, deploy a grammar change listener to perceive changes in the source system and update the rules, and introduce a bootstrap synthesizer to adapt to the new system.
It achieves a continuous closed loop from syntax parsing to structure normalization of multi-source heterogeneous data, reduces the frequency of manual maintenance, reduces naming conflicts and format inconsistencies, improves the structural consistency and semantic integrity of data evaluation and analysis, and supports the stable operation of newly added systems.
Smart Images

Figure CN121391015A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to an enterprise data evaluation and analysis system and method based on big data. BACKGROUND
[0002] In the prior art, the enterprise data evaluation and analysis method relies on a big data platform and a data warehouse or a data lake, adopts an "extraction-transformation-loading" process to converge business system data, realizes caliber unification through metadata management, data blood relationship, master data management and data standards, and completes field naming, data format, value range and primary-foreign key integrity checking by a rule engine; a subject domain and an index system are constructed at the semantic layer to support report and model scoring. A common practice is to take a manually maintained syntax analysis rule and a mapping table as the core, maintain a field comparison, format conversion and association relationship according to system coding, and generate a unified detail table and a wide table by combining a scheduling job and an incremental change capture to provide input for evaluation and analysis. Some schemes introduce ontology or knowledge graph to describe business entities and relationships to enhance cross-system semantic alignment.
[0003] However, the prior art has the problem of dynamic adaptation of syntax rules of multi-source heterogeneous enterprise data: enterprise data comes from multiple systems such as ERP, CRM, supply chain management, and the syntax structures (such as field naming rules, data formats, and association logic) of data from different systems are significantly different. Traditional static syntax analysis rules are difficult to adapt to newly added systems or data format changes in real time. How to construct a dynamic syntax rule library based on scene perception to realize automatic syntax checking and structure normalization of multi-source heterogeneous data has become a key bottleneck to improve the quality of evaluation and analysis basic data. SUMMARY
[0004] The purpose of the present application is to provide an enterprise data evaluation and analysis system and method based on big data to solve the problems in the background art: The purpose of the present application can be achieved by the following technical solutions: An enterprise data evaluation and analysis method based on big data, comprising: S1: Establish a scene ontology syntax rule library, disassemble ERP, CRM, supply chain management fields into syntax atoms, label entity relationship semantics according to business scenes, and summarize the results into reusable syntax atom templates; S2: Deploy a syntax perceiver at the data collection end to capture field naming, data format and association logic, generate a syntax fingerprint, and write it into the scene ontology syntax rule library for backup; S3: According to the syntax fingerprint and the syntax atom template, generate field naming mapping, format mapping and association mapping, form an alignment mapping table as input for structure normalization; S4: according to the scene label and the alignment mapping table, a structure normalization pipeline is constructed, field renaming, format conversion and correlation assertion verification are performed, and a unified structure intermediate layer table is generated; S5: a syntax change listener is deployed, the table structure and field change of the source system are perceived, a regression verification task is triggered, the syntax fingerprint, the alignment mapping table and the structure normalization pipeline are updated; S6: a bootstrap synthesizer is introduced, boundary sample records are synthesized and the structure normalization pipeline is run when a new system is accessed without annotation, and an initial mapping and a correlation assertion are derived according to a difference track.
[0005] As a further scheme of the application: in step S1, the process of establishing a scene ontology syntax rule library, disassembling ERP, CRM and supply chain management fields into syntax atoms, annotating entity relationship semantics according to business scenes, and summarizing the results into reusable syntax atom templates is: A scene vocabulary and an entity vocabulary are established, field naming rules, data format rules and correlation rules are defined, and an initial list of the scene ontology syntax rule library is generated; A syntax splitter is deployed to disassemble ERP, CRM and supply chain management fields according to naming fragments, format fragments and correlation fragments, and output syntax atom records with source labels; A scene behavior script set is constructed to define the field appearance order and necessary constraints for each business process, and the syntax atoms are annotated with entity relationship semantics and scene labels according to the script; A bidirectional mapping graph is established with syntax atoms as nodes and entity relationships as edges, and syntax atoms with consistent structures are merged into reusable syntax atom templates and written into the scene ontology syntax rule library.
[0006] As a further scheme of the application: in step S2, the process of deploying a syntax recognizer at the data acquisition end to capture field naming, data format and correlation logic, generating a syntax fingerprint, and writing it into the scene ontology syntax rule library is: A syntax recognizer is embedded at the data acquisition end, a data extraction adapter is accessed and the source system identifier is registered, field naming rules, data format rules and correlation rules are loaded, and the basis for subsequent source system labeling is laid; The syntax recognizer is run to scan the table structure and sample records, analyze the naming fragments, format fragments and correlation fragments, and generate the syntax fingerprint, and the source system label and time label are attached; The syntax fingerprint is written into the scene ontology syntax rule library with the source system + data table + version as the unique identification key, and the field naming index and the correlation index are established.
[0007] As a further scheme of the present application: in the step S3, the process of generating the field naming mapping, the format mapping and the association mapping according to the syntax fingerprint and the syntax atom template, forming the alignment mapping table as the input of the structure normalization, is: Enabling the naming comparison module, extracting the naming fragment from the syntax fingerprint, comparing with the field naming unit of the syntax atom template, and generating the field naming mapping candidate set; Building a format conversion procedure, generating type mapping, date format mapping and enumeration value order mapping according to the format fragment of the syntax fingerprint and the format constraint of the syntax atom template; Generating association constraint graph, combining the association fragment of the syntax fingerprint and the entity relationship semantics, forming the primary-foreign key mapping, the reference path mapping and the cascade direction mapping, and executing the reverse conflict resolution process; Calling the alignment assembler, assembling the alignment mapping table according to the field naming mapping, the format mapping and the association mapping, outputting the field-level alignment record as the input of the structure normalization pipeline.
[0008] As a further scheme of the present application: the process of generating the association constraint graph, combining the association fragment of the syntax fingerprint and the entity relationship semantics, forming the primary-foreign key mapping, the reference path mapping and the cascade direction mapping, and executing the reverse conflict resolution process, is: Extracting the association fragment from the syntax fingerprint, parsing the source entity, the target entity, the association field and the matching rule, and building the initial association constraint graph with the field as the endpoint; Labeling the edge type on the association constraint graph, forming the primary-foreign key mapping according to one-to-many and one-to-one, arranging the reference path mapping according to the sequence of multiple paths, and determining the cascade direction mapping; The target entity executes the reverse conflict resolution process to the source entity, backtracks the reference path mapping, checks the consistency of the cascade direction mapping and the primary-foreign key mapping, records the correction items and updates the association constraint graph.
[0009] As a further scheme of the present application: in the step S4, the process of constructing the structure normalization pipeline according to the scene label and the alignment mapping table, executing the field renaming, the format conversion and the association relationship assertion verification, and generating the uniform structure intermediate layer table, is: Loading the scene label and the alignment mapping table, enabling the scene gate control arrangement module, and generating the structure normalization pipeline task sequence according to the scene label; Starting the double-track conversion unit in the task sequence, executing the field renaming and the format conversion according to the field naming mapping and the format mapping, and forming the uniform structure draft table; Executing the association relationship assertion verification on the uniform structure draft table, sequentially executing the assertion according to the reference path mapping and the cascade direction mapping, and recording the violation position and the correction rule; Call the alignment assembler to write the records that pass the assertion check into the unified schema intermediate layer table and preserve the field naming mapping and trace markers of the association mapping.
[0010] As a further aspect of the present application: in step S5, the deployment of the syntax change listener to perceive the table structure and field changes of the source system, trigger the regression verification task, and update the syntax fingerprint, alignment mapping table, and structure normalization pipeline is as follows: Deploy the syntax change listener in the source system, interface the metadata log and data dictionary, capture the table structure and field changes, and generate normalized change event records; Introduce a point-in-time snapshot recorder to generate snapshots before and after the change event and compare them with the existing syntax fingerprint to generate a fingerprint difference and a list of affected objects; Start the regression verification task according to the list of affected objects, replay the key links of the structure normalization pipeline, and verify the consistency of the field naming mapping, format mapping, and association assertion rules; According to the regression results, revise the syntax fingerprint and alignment mapping table, generate a pipeline revision package, and load it into the structure normalization pipeline through a gray strategy.
[0011] As a further aspect of the present application: in step S6, the deployment of the syntax change listener to perceive the table structure and field changes of the source system, trigger the regression verification task, and update the syntax fingerprint, alignment mapping table, and structure normalization pipeline is as follows: Start the bootstrap synthesizer, read the table structure and field value domain of the new system, and synthesize boundary sample records covering naming boundaries, format boundaries, and association boundaries according to the syntax atomic template; Send the boundary sample records and the original sample records to the structure normalization pipeline, enable full-link logging, generate field difference tracks, and association difference tracks; According to the difference tracks, generate mapping rules according to the naming dimension, format dimension, and association path, derive the initial mapping and association assertion, and output the trial run configuration and trace markers.
[0012] An enterprise data evaluation and analysis system based on big data, comprising: A scene ontology syntax rule library construction module for disassembling ERP, CRM, and supply chain management fields into syntax atoms, annotating entity relationship semantics according to business scenarios, and inducing reusable syntax atomic templates; A syntax perception and fingerprint generation module for capturing field naming, data format, and association logic at the data collection end, generating a syntax fingerprint, and writing it into the scene ontology syntax rule library for backup; An alignment mapping generation module for generating field naming mapping, format mapping, and association mapping according to the syntax fingerprint and syntax atomic template, forming an alignment mapping table as input for structure normalization. A structure normalization pipeline module is configured to perform field renaming, format conversion and association assertion verification according to the scene label and the alignment mapping table, and generate a unified structure intermediate table; A syntax change listener and regression verification module is configured to sense source system table structure and field changes, trigger a regression verification task, and update the syntax fingerprint, the alignment mapping table and the structure normalization pipeline; A synthesis and difference derivation module is configured to synthesize boundary sample records when a new system without annotation is accessed, run the structure normalization pipeline, and derive initial mapping and association assertions according to difference trajectories.
[0013] The present application has the following advantages: The scene ontology syntax rule library and the pre-built syntax atom template are combined with the syntax fingerprint generated by the syntax sensor to complete field naming mapping, format mapping and association mapping under the driving of the alignment mapping table, and field renaming, format conversion and association assertion verification are implemented in the structure normalization pipeline to form a unified structure intermediate table. The above link realizes continuous closed loop from syntax analysis to structure normalization of multi-source heterogeneous data, reduces the frequency of manual maintenance of static rules, reduces the basic data deviation caused by naming conflicts, format inconsistencies and incomplete associations, and provides consistent, traceable and reusable basic structure and semantic alignment results for enterprise data evaluation and analysis.
[0014] Further advantages are that the syntax change listener continuously senses source system table structure and field changes, triggers a regression verification task, and synchronously revises the syntax fingerprint, the alignment mapping table and the structure normalization pipeline to ensure that the rules and data forms remain consistent; in the scene of accessing a new system without annotation, a bootstrap synthesizer is introduced to synthesize boundary sample records and run the structure normalization pipeline, initial mapping and association assertions are derived according to difference trajectories to realize bootstrap adaptation and rapid convergence of unknown syntax. This mechanism enables the scene ontology syntax rule library to have dynamic evolution capability, takes into account the stable operation of new and existing systems, improves the structural consistency and semantic integrity of evaluation and analysis input data, and provides stable and reliable basic support for subsequent model scoring and index calculation. BRIEF DESCRIPTION OF DRAWINGS
[0015] The present application will be further described below with reference to the accompanying drawings.
[0016] Figure 1 is a flowchart of an enterprise data evaluation and analysis method based on big data according to the present application.
[0017] Figure 2 is a structural diagram of an enterprise data evaluation and analysis system based on big data according to the present application. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0019] Please refer to Figure 1 The present application is a kind of enterprise data evaluation analysis method based on big data, comprising: S1: Establish a scene ontology grammar rule library, disassemble ERP, CRM, supply chain management fields into grammar atoms, mark entity relationship semantics according to business scenarios, and summarize the results into reusable grammar atom templates; S2: Deploy a grammar sensor at the data collection end to capture field naming, data format and associated logic, generate a grammar fingerprint, and write it into the scene ontology grammar rule library for backup; S3: According to the grammar fingerprint and the grammar atom template, generate field naming mapping, format mapping and association mapping, form an alignment mapping table as input for structure normalization; S4: According to the scene label and the alignment mapping table, construct a structure normalization pipeline, perform field renaming, format conversion and association relationship assertion verification, and generate a unified structure intermediate layer table; S5: Deploy a grammar change listener to sense the table structure and field changes of the source system, trigger a regression verification task, and update the grammar fingerprint, alignment mapping table and structure normalization pipeline; S6: Introduce a bootstrap synthesizer to synthesize boundary sample records and run the structure normalization pipeline when a new system without annotation is accessed, and deduce the initial mapping and association relationship assertion according to the difference trajectory.
[0020] In step S1, the process of establishing a scene ontology grammar rule library, disassembling ERP, CRM, supply chain management fields into grammar atoms, marking entity relationship semantics according to business scenarios, and summarizing the results into reusable grammar atom templates is as follows: In this embodiment, first, the scene vocabulary and entity vocabulary are established, and for the deterministic business scenarios such as sales, procurement, warehousing, finance, etc., the entity names and field lists of customers, orders, order details, materials, suppliers, etc. are given. Then the field naming rules, data format rules and association relationship rules are defined to form unified semantic and structural constraints. For example, "customer number" and "customer code" are classified into the same naming category, "order date" is unified into the date format of "year-month-day", and "order-customer" is set as a one-to-many association relationship. Based on the above rules, an "initial list of scene ontology grammar rule library" is generated, which records the standard name, format requirement and association constraint of each entity and field.
[0021] After completing the rule baseline, the grammar splitter is deployed to split the fields in the enterprise resource planning system, customer relationship management system and supply chain management system according to the naming fragments, format fragments and association fragments, and output the grammar atom records with source labels. The grammar splitter splits "customer number" into the naming fragments of "customer" and "number", annotates the format fragment as "string", and annotates the association fragment as "subordinate to the customer entity". For another example, "order date" is split into the naming fragments of "order" and "date", the format fragment is annotated as "date", and the association fragment is annotated as "subordinate to the order entity". Each grammar atom record is attached with a source label to indicate that it comes from the enterprise resource planning system, the customer relationship management system or the supply chain management system.
[0022] Then, the scene behavior script set is constructed to determine the field appearance order and mandatory constraints driven by business processes. Taking "sales order creation" as an example, the script specifies that "customer identification" appears first, then "order identification", followed by "order details", "amount" and "order date", and requires that "customer identification" is a valid customer, "amount" is a non-negative number, and "order date" is earlier than "shipping date". According to the scene behavior script set, the aforementioned grammar atoms are annotated with entity relationship semantics and scene labels one by one: the grammar atoms related to customers are annotated as "customer-order" reference relationship, the grammar atoms related to details are annotated as "order-order details" inclusion relationship, and the scene label of "sales order creation" is assigned. In this way, the grammar atoms obtain clear business positioning and context boundaries, and the subsequent mapping and normalization steps establish consistent relationship constraints accordingly.
[0023] After the syntax atom and the scene semantics are established, a bidirectional mapping diagram is established, taking the syntax atom as the node and the entity relationship as the edge, and performing bidirectional alignment and conflict resolution in three dimensions of naming, format and association. Specifically, the "customer code" in the enterprise resource planning system and the "customer number" in the customer relationship management system are mapped to the same standard field "customer identification", the "order date" and "order date" are mapped to the same standard field "order date", and the edges of "order-customer" and "order-order detail" are unified according to the entity relationship. For syntax atoms with consistent structure, merge into reusable syntax atom templates, such as forming "customer identification (string, unique, reference to customer entity)" and "order date (date, mandatory, subordinate to order entity)". After merging, reusable syntax atom templates are written into the scene ontology syntax rule library, realizing consistent naming, consistent format and consistent relationship across systems.
[0024] In the step S2, a syntax perceiver is deployed at the data collection end to capture field naming, data format and association logic, generate syntax fingerprints, and write into the scene ontology syntax rule library for backup. The process is as follows: The embodiment embeds a syntax perceiver at the data collection end and realizes stable connection with each source system through a data extraction adapter. When deployed, the source system identifier is registered first to distinguish the source channels of the enterprise resource planning system, the customer relationship management system and the supply chain management system. Then, the field naming rules, data format rules and association relationship rules are loaded into the syntax perceiver as the basis for analysis and the boundary for verification. For example, the field naming rule requires that fields with the same meaning use a unified root combination, and "customer number" and "customer code" are constrained to the naming structure of "customer + number"; the data format rule requires that date fields use the "year-month-day" format, and numerical fields specify the decimal place and thousandth place writing method; the association relationship rule clearly indicates the connection direction and the primary-foreign key constraint of one-to-many and one-to-one. Through the above loading operation, the source system tag can be accurately pointed to the data source when generated subsequently, and the determination standard of syntax analysis is ensured to be consistent.
[0025] After loading, run the syntax recognizer, and scan the table structure and sample records of the source system table by table. The syntax recognizer splits each field into named fragments, format fragments, and association fragments according to the aforementioned rules, and generates a syntax fingerprint based on the combination thereof. The named fragments are used to express semantic roots and modifiers, the format fragments are used to express data types and value specifications, and the association fragments are used to express entity relationships and foreign key paths. For example, in the order table, the "customer number" is split into the named fragment "customer / number", the format fragment is marked as "string", and the association fragment is marked as "reference to the primary key of the customer entity"; the "order date" is split into the named fragment "order / date", the format fragment is marked as "date", and the association fragment is marked as "belonging to the order entity". The syntax recognizer combines the split results of the fields in the same table according to the field order and constraint relationship to form the syntax fingerprint of the table, and adds the source system label and the time label to the fingerprint to indicate the source system and the generation time point. With the fingerprint, the meaning, format, and relationship path of any subsequent field can be traced back to the specific parsing basis and extraction scene, avoiding ambiguity caused by homonymy or synonymy.
[0026] In order to realize unified management of retrievability and traceability, the embodiment takes "source system + data table + version" as a unique identification key, writes the generated syntax fingerprint into the scene ontology syntax rule library, and synchronously establishes a field naming index and an association relationship index. The unique identification key is used to distinguish different evolution versions of the table structure in the same source system, such as "a certain ERP system + order table + third edition", so as to keep historical traces when migrating versions or adjusting structures. The field naming index is used for retrieval according to standard roots and named fragments, for example, inputting "customer / number" can locate all fields with the same semantics of "customer number" and their format requirements in the source system. The association relationship index is used for reverse and forward queries along the primary-foreign key path and the reference path, for example, starting from the path of "order-customer", the field where the foreign key is located, the connection direction, and the cascade strategy can be found. Through the above writing and indexing mechanisms, the storage, positioning, and comparison of the syntax fingerprint form a closed loop, and when new data is accessed or the table structure is changed, the affected objects can be quickly located based on the unique identification key, and the related fields and relationships can be quickly linked based on the index, thereby providing stable, clear, and reviewable basis for continuous structure normalization and rule updating.
[0027] In the step S3, the field naming mapping, the format mapping, and the association mapping are generated according to the syntax fingerprint and the syntax atom template, and an alignment mapping table is formed, which is used as the input of the structure normalization process: The embodiment first enables the naming matching module, extracts the naming fragment from the syntax fingerprint field by field, and compares it with the field naming unit in the syntax atom template one by one to form a field naming mapping candidate set. The naming fragment is used to express the semantic root and modifier of the field, and the field naming unit is used to express the standardized target naming. In order to facilitate understanding, taking "customer number" as an example, the naming fragment given by the syntax fingerprint is "customer / number", and the naming matching module retrieves the "customer identification" field naming unit in the syntax atom template, which is semantically equivalent, and generates the naming mapping candidate "customer number→customer identification". Taking "material code" as an example, the naming fragment is "material / code", and there are two candidates "commodity identification" and "material identification" on the template side. The naming matching module determines according to the scene label and entity attribution, selects "material identification" consistent with the business scene of the source field, and writes the selection basis and determination path in the candidate record.
[0028] After obtaining the field naming mapping candidate set, the format conversion procedure is constructed, and the type mapping, date format mapping and enumeration value order mapping are generated according to the format fragment of the syntax fingerprint and the format constraint of the syntax atom template. The type mapping is used to determine the data type consistency of the source field and the target field, the date format mapping is used to determine the unified style of date display and storage, and the enumeration value order mapping is used to unify the arrangement and coding of limited values. For example: the source system "order date" is stored in a compact form of "year-year-year-month-month-day-day", and the format fragment is marked as "date", and the template side requires "year-month-day", and the format conversion procedure generates "compact date→hyphenated date" date format mapping accordingly; the source system "order status" takes the value "new / audited / completed", the template side requires "created / audited / completed", and the natural order is "created→audited→completed", then the procedure generates "new→created" "audited→audited" "completed→completed" enumeration value mapping, and writes the order constraint to ensure the consistency of subsequent ordering and comparison.
[0029] Subsequently, the association constraint graph is generated, and the main foreign key mapping, reference path mapping and cascade direction mapping are formed by combining the association fragment of the syntax fingerprint and the entity relationship semantics. The main foreign key mapping is used to clarify the reference relationship between entities, the reference path mapping is used to describe the path order of cross-table access, and the cascade direction mapping is used to specify the propagation direction when inserting, updating and deleting. In order to avoid conflicts caused by differences between source systems, after generating the above three types of mapping, the reverse conflict resolution process is executed, which traces back to the existing mapping, identifies the situation where the direction is opposite or the path is inconsistent, and generates a correction item. For example: if the deletion propagation of "order details→order" is marked as "cascading up" at one place, and the template side requires "blocking down", the reverse conflict resolution process will record the difference, unify it to the requirement of the template side, and write the source, location and reason in the correction item.
[0030] After the naming mapping, format mapping and association mapping are clear, the alignment assembler is called to assemble according to the three types of mapping, generate the alignment mapping table, and output the field-level alignment record as the input of the structure normalization pipeline. The alignment mapping table takes "source system / data table / field" as the primary key, and pairs "standard field name" "type mapping strategy" "date format mapping strategy" "enumeration value order mapping strategy" and "association mapping summary". For example, "customer number" from the customer relationship management system is assembled as "customer identification", the type mapping is "string retention", the date mapping is empty, the enumeration mapping is empty, and the association mapping is marked as "reference to customer entity". This record can directly drive the field renaming and association relationship assertion verification in the pipeline. Through this assembly process, the input data obtains a unified naming, format and relationship baseline before entering the structure normalization pipeline.
[0031] In the further process of the embodiment, to more finely illustrate the internal process of "generating association constraint graph", the embodiment first extracts the association segment from the syntax fingerprint, parses the source entity, target entity, association field and matching rule, and constructs an initial association constraint graph with the field as the endpoint. The endpoint is identified in the form of "table name-field name", and the edge is used to express the reference direction and matching condition. For example, in the order table, "customer_id" as the field name of the source system is uniformly "customer identification foreign key" in this specification, the endpoint is "order-customer identification foreign key", the target endpoint is "customer-customer identification primary key", and the matching rule is "equality matching". According to this, a directed edge from "order-customer identification foreign key" to "customer-customer identification primary key" is generated in the graph. The graph is used to carry out subsequent type checking, path arrangement and cascade direction determination.
[0032] On the basis of the initial association constraint graph, the edge type is marked for each edge, and the primary-foreign key mapping is formed according to the one-to-many and one-to-one relationship forms; at the same time, the reference path mapping is arranged according to the multi-path sequence, and the cascade direction mapping is determined. Taking "order→order detail→goods" as an example, "order→order detail" is marked as "one-to-many", and the foreign key endpoint is "order detail-order identification foreign key"; "order detail→goods" is marked as "many-to-one", and the foreign key endpoint is "order detail-goods identification foreign key". The reference path mapping of "order→goods" is generated along the multi-path, the intermediate nodes and the connection order are marked, and the cascade direction mapping is determined in combination with the template side business constraints, such as "downwardly deleting order details" when deleting "order" and "blocking the propagation to goods". The above marking results will be written back to the index views of the primary-foreign key mapping, the reference path mapping and the cascade direction mapping at the same time, so as to facilitate retrieval and review.
[0033] Finally, the reverse conflict resolution process is performed in the direction of "target entity to source entity", and the consistency of the cascade direction mapping and the primary foreign key mapping is checked along the reference path mapping, and the problems found are generated and the associated constraint graph is updated. For example: if the same path appears on different sources "order deletion cascades to order details" description is inconsistent, one is "cascading deletion", the other is "keep details", reverse backtracking will locate the conflicting edge, and according to the template side established rules confirm "cascading deletion". According to this record the correction item, including the conflict source, path position and final decision, and update the cascade direction annotation of the corresponding edge. Through this process, the association constraint graph reaches a unified state in semantics and behavior, providing reliable and traceable association basis for the alignment assembler and structure normalization pipeline mentioned above.
[0034] In the step S4, the process of constructing the structure normalization pipeline according to the scene tag and alignment mapping table, performing field renaming, format conversion and association relationship assertion verification, and generating a unified structure intermediate layer table is as follows: The embodiment first loads the scene tag and alignment mapping table, enables the scene gate control arrangement module, and generates a structure normalization pipeline task sequence according to the scene tag. The scene gate control arrangement module reads the business process identifier and field priority in the scene tag, combines the field naming mapping, format mapping and association mapping confirmed in the alignment mapping table, and outputs an executable task list in the order of "reading source data - naming processing - format processing - association verification - result summary". Taking the "sales order import" scene as an example, the task sequence first locks the customer related fields and order master table fields, then continues the order detail fields, and finally arranges the association verification task of the order and order details; in the "purchase receipt entry" scene, the task sequence is expanded in the order of supplier, arrival order and arrival detail. Through the above arrangement, each subsequent link is executed in a unified order and boundary, avoiding ambiguous processing or repeated processing of the same field in different scenes.
[0035] In the task sequence, the dual-track conversion unit (named track and format track) is started, field renaming and format conversion are performed according to the field name mapping and format mapping, and a unified structure draft table is formed. The named track renews the source field to the standard field according to the field name mapping. For example, "customer number" and "customer code" are uniformly renamed as "customer identification", and "order date" and "order date" are uniformly renamed as "order date". The format track completes the format conversion of date, value and enumeration according to the format mapping. For example, "September 29, 2024" or "20240929" is uniformly converted to "year-month-day" format; the thousandth position separation in the amount field is removed and the uniform decimal place is retained; "new", "audited" and "completed" are mapped to "created", "audited" and "completed", and the order is maintained. The processing results of the named track and the format track are merged and output at the record level to form a unified structure draft table.
[0036] The association relationship assertion verification is performed on the unified structure draft table, and the assertion is sequentially performed according to the reference path mapping and the cascade direction mapping, and the violation position and the correction rule are recorded. According to the reference path mapping, the path verification is performed on the cross-table field. For example, according to the path of "order->order detail->product", it is verified whether the "order identification foreign key" in the "order detail" can find the corresponding primary key in the "order", and whether the "product identification foreign key" in the "order detail" can find the corresponding primary key in the "product"; then according to the cascade direction mapping, it is verified whether the propagation direction of deletion and update is consistent with the established direction of the template side. If it is found that the processing of "order deletion still retains order detail" is inconsistent with the cascade direction mapping, the object position, field name, path position and correction rule are recorded in the violation list, for example, "the deletion propagation direction is adjusted to be cascaded downward, and the order detail is processed first, then the order". All violations are recorded in a unified format, which is convenient for playback, review and revision.
[0037] The alignment assembler is invoked to write the records that pass the assertion check into the unified structure intermediate layer table, and to preserve the field naming mapping and provenance tags of the association mapping. The alignment assembler populates the intermediate layer table with the records that pass the naming and formatting processing and the associated assertions according to the standard field order and field grouping in the alignment mapping table; the records that fail the assertions are kept isolated and marked with violation flags, and are not populated into the intermediate layer table. To support traceability and incremental update, each record in the intermediate layer table is tagged with provenance information, including the source system identifier, the source data table name, the original field name, the adopted field naming mapping identifier, the adopted format mapping identifier, and the association mapping summary. Taking the “sales order import” as an example, a record written into the intermediate layer explicitly indicates that the record comes from the order table of the customer relationship management system, the original “customer number” is named as “customer identifier” through naming mapping, the original “order date” is formatted as “order date” through format mapping, and the path information of “referencing customer entity” is marked in the association mapping summary. The above provenance tags enable quick positioning of the affected objects and processing links when new fields or rules are added or adjusted, ensuring that the content of the unified structure intermediate layer table is consistent with the alignment mapping table, and providing clear and reviewable basic data for subsequent structure normalization pipeline output and evaluation analysis calculation.
[0038] In the step S5, the deployment of the syntax change listener to perceive the table structure and field changes of the source system and trigger the regression check task to update the syntax fingerprint, the alignment mapping table, and the structure normalization pipeline is as follows: The syntax change listener is deployed in each source system to continuously capture table structure and field changes by interfacing with the metadata log and the data dictionary, and to generate normalized change event records. The syntax change listener subscribes to the metadata change stream of the source system, combines the table description and field description in the data dictionary, and identifies categories such as “new table”, “deleted table”, “new field”, “deleted field”, “field renaming”, “data type change”, “constraint change”, “index change”, etc. For ease of unified processing, the normalized change event record uses a fixed field set, including event number, source system identifier, data table name, field name, change type, original value summary, new value summary, time point of occurrence, and event source. For example, if the customer relationship management system renames “customer code” to “customer number”, the listener reads the field name change from the metadata log and pulls the field attributes before and after the change from the data dictionary to generate a “field renaming” event; for another example, if the enterprise resource planning system adjusts the data type of “order time” from “text” to “date”, the listener generates a “data type change” event and marks the original type and new type in the record. Through the above standardization processing, the subsequent processes can stably identify and replay the change impact without relying on specific system details.
[0039] After obtaining the normalized change event, a point-in-time snapshot recorder is introduced to generate pre-change and post-change snapshots at the point-in-time of the event occurrence, and compare them with the existing syntax fingerprints to form a fingerprint difference and an affected object list. The pre-change snapshot saves the table structure, field attributes and constraint descriptions involved in the event, and the post-change snapshot saves the corresponding information after the change; the syntax fingerprint, as a result of the previous collection and analysis, contains three parts: naming fragments, format fragments and association fragments. The fingerprint difference issues a comparison result in the dimensions of "naming difference, format difference and association difference", and locates the difference points at the field level and the constraint level; the affected object list enumerates the units that need to be reviewed, covering syntax fingerprint entries, alignment mapping table entries and tasks in the structure normalization pipeline that depend on the entries. For example, if the renaming event of "customer code -> customer number" is triggered, the fingerprint difference gives "naming fragment difference: code -> number", and the affected object list includes "customer identification related field naming mapping" and "customer master data loading task"; if the type adjustment of "order time -> date" is triggered, the fingerprint difference gives "format fragment difference: text -> date", and the affected object list includes "date format mapping" and "order loading format conversion task".
[0040] According to the affected object list, a regression verification task is started to perform playback verification for key links in the structure normalization pipeline, focusing on verifying the consistency of field naming mapping, format mapping and association relationship assertion rules. The playback order is consistent with the online arrangement, first verifying whether the renaming can still be aligned to the standard field in the naming processing link, then verifying whether the type and format conversion meet the template constraints in the format processing link, and finally verifying whether the primary foreign key reference and cascade direction remain unchanged in the association relationship assertion link. For example, for the renaming of "customer number", the regression verification task checks whether the naming mapping still points to "customer identification", and verifies that the position and meaning of the renamed field in the unified structure draft table do not change; for the type adjustment of "order time", the regression verification task verifies whether the "text -> date" format mapping exists, whether the date parsing rule covers the changed format, and whether the reference path of "order -> order details" is not affected in the association relationship assertion. If inconsistencies are found, such as some pipeline still parses "customer code" as the old name or the date parsing rule does not cover the new format, the regression verification task records the violation position, triggering data, expected rules and correction rules in the report, providing a basis for subsequent revision.
[0041] According to the regression results, the syntax fingerprint and the alignment mapping table are revised, a pipeline revision package is generated, and is loaded into the structure normalization pipeline through the gray-scale strategy. The revision action follows the order of "fingerprint first, mapping second, pipeline last": first, the differences between the named fragments, the format fragments and the associated fragments are written back to the syntax fingerprint, then the field naming mapping and the format mapping in the alignment mapping table are updated accordingly, and finally the pipeline revision package containing the arrangement adjustment and the rule version number is generated. In order to reduce the online risk, the new rules are put into use in groups according to the source system identifier, the data table and the version through the gray-scale strategy, for example, the new version is first enabled for the "order table" of the customer relationship management system, and the old version is maintained for the enterprise resource planning system for a period of observation. During the gray-scale period, the traceability markers and the comparison logs are retained at the same time, and the differences in the naming processing, the format processing and the association assertion between the new and old versions are monitored; if the monitoring results are stable, the gray-scale range is expanded until the full replacement; if an abnormality occurs, the modified rules in the regression report are rolled back or revised again for retry. Through the above closed loop, the syntax fingerprint, the alignment mapping table and the structure normalization pipeline are updated in a timely, traceable and low-risk manner after the change event is triggered.
[0042] In the step S6, the syntax change listener is deployed to perceive the table structure and field change of the source system, trigger the regression verification task, and update the syntax fingerprint, the alignment mapping table and the structure normalization pipeline. The process is as follows: In the new system access link, the bootstrap synthesizer of the present embodiment is started, the table structure and field value domain of the new system are read, and the field list, type specification and value constraint are formed. The bootstrap synthesizer automatically synthesizes sample records covering the naming boundary, the format boundary and the association boundary according to the syntax atomic template, which is used to verify whether the rules are complete and robust. The naming boundary sample is used to verify the synonym naming and abbreviation difference, for example, whether "customer number", "customer code" and "customer number" can be mapped to the standard "customer identifier"; the format boundary sample is used to verify the extreme form of date, amount and enumeration value, for example, whether "20250929", "2025 September 29" and "1,234.00" and "new / audited / completed" can be unified to "year-month-day", decimal place and standard enumeration value according to the template requirements; the association boundary sample is used to cover the absence, suspension and multiple paths of the primary foreign key, for example, to construct complete link samples of "order-order details-goods", and to construct abnormal samples of missing "order identifier foreign key". The correspondence between the above-mentioned samples and the source, field and template entry is written into the sample list, which is used as the baseline for subsequent comparison.
[0043] The boundary sample record and the original sample record are sent into a structure normalization pipeline, and full-link logging is enabled to obtain sufficient process evidence. A naming track in the pipeline performs renaming according to a field naming mapping, records a track of “original field name → standard field name”, for example, “customer ID → customer identifier”; a format track performs conversion according to a format mapping, records a track of “original format → target format”, for example, “20240929 → 2024-09-29” and “1,234.00 → 1234.00”; and association verification checks whether a primary foreign key is reachable according to a reference path, and records a judgment process of a cascade direction. “Field difference tracks” are generated at a record level and a field level, and are used to mark inconsistent points of naming and format; and “association difference tracks” are generated at a relationship level, and are used to mark specific positions of path blockage, direction inconsistency and constraint conflict. For example, when “order detail – order identifier foreign key” cannot match “order – order identifier primary key”, the association difference track marks an error endpoint, an expected path and a failure cause; and when “has been reviewed” is identified as “has been reviewed” but does not enter the standard “has been reviewed”, the field difference track records a mapping gap. All the tracks are provided with time marks and sample marks.
[0044] According to the field difference track and the association difference track, a bootstrapping synthesizer generates mapping rules according to naming dimensions, format dimensions and association paths, derives initial mapping and association relationship assertions, and outputs trial operation configurations and traceability marks. On the naming dimension side, synonyms and abbreviations are converged into standard names, such as “customer code / customer number → customer identifier”; on the format dimension side, a rule chain of date parsing and amount standardization is deposited, such as “compact date → hyphenated date” and “remove thousandth place → unify decimal place”; and on the association path side, a primary foreign key pair and multiple access sequences are fixed, and cascade direction assertions are given, such as “order deletion cascades order details downward, and blocks propagation to commodity entities”. The above initial mapping and association relationship assertions are packaged into trial operation configurations, and are attached with traceability marks, which indicate source samples, hit templates and difference bases, so as to facilitate review and playback. For example, the configurations include a naming rule of “customer ID → customer identifier”, an enumeration mapping of “new / reviewed / completed → created / reviewed / completed”, and a relationship assertion of “order detail order identifier foreign key → order – order identifier primary key, deletion cascades downward”. The trial operation configurations are only used for verification in a controlled environment, and do not change online strategies; and after verification passes, the configurations enter a subsequent release process. Through the above steps, a new system can obtain an initial alignment scheme that is explainable and traceable without annotation.
[0045] Referring to Figure 2 As shown in FIG. 1, A scene ontology grammar rule base construction module is configured to disassemble ERP, CRM, and supply chain management fields into grammar atoms, label entity relationship semantics according to business scenarios, and induce reusable grammar atom templates; A grammar awareness and fingerprint generation module is configured to capture field naming, data format, and associated logic at a data acquisition end, generate grammar fingerprints, and write the grammar fingerprints into a scene ontology grammar rule base for backup; An alignment mapping generation module is configured to generate field naming mapping, format mapping, and association mapping according to grammar fingerprints and grammar atom templates, and form an alignment mapping table as an input of structure normalization; A structure normalization pipeline module is configured to perform field renaming, format conversion, and association relationship assertion verification according to a scene label and the alignment mapping table, and generate a unified structure intermediate layer table; A grammar change monitoring and regression verification module is configured to sense source system table structure and field changes, trigger a regression verification task, and update grammar fingerprints, an alignment mapping table, and a structure normalization pipeline; A synthesis and difference derivation module is configured to synthesize boundary sample records when a new system is accessed without annotation, run a structure normalization pipeline, and derive initial mapping and association relationship assertions according to a difference track.
[0046] The above describes one embodiment of the present application in detail, but the content described is only a preferred embodiment of the present application and cannot be considered as limiting the implementation scope of the present application. Any equivalent changes and improvements made according to the scope of the present application should still belong to the patent coverage scope of the present application.
Claims
1. A method for enterprise data evaluation and analysis based on big data, characterized in that, Includes the following steps: S1: Establish a scenario ontology syntax rule library, decompose ERP, CRM and supply chain management fields into syntax atoms, annotate entity relationship semantics according to business scenarios, and summarize the results into reusable syntax atom templates. S2: Deploy a syntax awareness device at the data acquisition end to capture field naming, data format and association logic, generate syntax fingerprints, and write them into the scene ontology syntax rule library for later use. S3: Based on the syntax fingerprint and syntax atomic template, generate field naming mapping, format mapping and association mapping to form an alignment mapping table, which serves as the input for structure normalization; S4: Based on scene labels and alignment mapping tables, construct a structure normalization pipeline, perform field renaming, format conversion and relational assertion validation, and generate a unified structure intermediate layer table; S5: Deploy a syntax change listener to detect changes in the table structure and fields of the source system, trigger regression verification tasks, and update the syntax fingerprint, alignment mapping table, and structure normalization pipeline. S6: Introduces a bootstrap synthesizer to synthesize boundary sample records and run a structure normalization pipeline when a new unlabeled system is connected, and derives initial mapping and correlation assertions based on the difference trajectory.
2. The enterprise data evaluation and analysis method based on big data according to claim 1, characterized in that, In step S1, the process of establishing a scenario ontology syntax rule base, decomposing ERP, CRM, and supply chain management fields into syntax atoms, annotating entity relationship semantics according to business scenarios, and summarizing and organizing the results into reusable syntax atom templates is as follows: Establish scene vocabularies and entity vocabularies, define field naming rules, data format rules and association rules, and generate an initial list of scene ontology syntax rules. Deploy a syntax splitter to break down ERP, CRM, and supply chain management fields into named fragments, formatted fragments, and associated fragments, and output syntax atomic records with source tags; Construct a set of scenario behavior scripts, define the order of field appearance and necessary constraints for each business process, and annotate entity relationship semantics and scenario tags with syntactic atoms based on the scripts; Establish a bidirectional mapping graph with syntactic atoms as nodes and entity relationships as edges. Merge syntactic atoms with consistent structures into reusable syntactic atom templates and write them into the scene ontology syntactic rule library.
3. The enterprise data evaluation and analysis method based on big data according to claim 1, characterized in that, In step S2, the process of deploying a syntax-aware device at the data acquisition end, capturing field naming, data format, and association logic, generating a syntax fingerprint, and writing it into the scene ontology syntax rule library for later use is as follows: Embed a syntax awareness at the data acquisition end, connect to the data extraction adapter and register the source system identifier, load field naming rules, data format rules and association rules, and lay the foundation for the subsequent generation of source system tags; Run the syntax sensor, scan the table structure and sample records, parse the named fragments, format fragments and related fragments, combine them to generate a syntax fingerprint, and attach source system tags and time tags; Using the source system, data table, and version as unique identifiers, the syntax fingerprint is written into the scene ontology syntax rule library, and field naming indexes and relationship indexes are established.
4. The enterprise data evaluation and analysis method based on big data according to claim 1, characterized in that, In step S3, the process of generating field naming mappings, format mappings, and association mappings based on grammatical fingerprints and grammatical atomic templates to form an alignment mapping table, which serves as input for structure normalization, is as follows: Enable the naming comparison module to extract naming fragments from the syntax fingerprint, compare them with the field naming units of the syntax atomic template, and generate a candidate set of field naming maps; Construct a format conversion procedure, and generate type mapping, date format mapping and enumeration value order mapping based on the format fragments of the syntax fingerprint and the format constraints of the syntax atomic template; Generate an association constraint graph, combine the association fragments of the syntactic fingerprint with the semantics of entity relationships to form primary and foreign key mappings, reference path mappings and cascading direction mappings, and execute a reverse conflict resolution process; The alignment assembler is invoked to assemble an alignment mapping table based on field naming mapping, format mapping, and association mapping. The output field-level aligned records serve as input to the structure normalization pipeline.
5. The enterprise data evaluation and analysis method based on big data according to claim 1, characterized in that, The process of generating the association constraint graph, combining the association fragments of the syntactic fingerprint with the semantics of entity relationships, forming primary and foreign key mappings, reference path mappings, and cascading direction mappings, and then executing the reverse conflict resolution process is as follows: Extract related fragments from syntactic fingerprints, parse source entities, target entities, related fields and matching rules, and construct an initial association constraint graph with fields as endpoints; Label the edge types on the association constraint graph, form primary and foreign key mappings based on one-to-many and one-to-one relationships, organize the reference path mappings according to multiple path sequences, and determine the cascading direction mappings; The target entity executes a reverse conflict resolution process to the source entity, backtracks the reference path mapping, verifies the consistency between the cascade direction mapping and the primary-foreign key mapping, records the correction items, and updates the association constraint graph.
6. The enterprise data evaluation and analysis method based on big data according to claim 1, characterized in that, In step S4, the process of constructing a structure normalization pipeline based on scene labels and alignment mapping tables, performing field renaming, format conversion, and association assertion verification to generate a unified structure intermediate layer table is as follows: Load the scene label and alignment mapping table, enable the scene gating orchestration module, and generate a structure-normalized pipeline task sequence according to the scene label. Initiate a dual-track conversion unit in the task sequence, perform field renaming and format conversion based on field naming mapping and format mapping, and form a draft table with a unified structure. Perform relational assertion checks on the unified structure draft table, execute assertions sequentially according to the reference path mapping and cascading direction mapping, and record the violation location and correction rules; The alignment assembler is invoked to write records that pass the assertion verification into a unified structure intermediate layer table, while preserving the traceability markers of field naming mappings and association mappings.
7. The enterprise data evaluation and analysis method based on big data according to claim 1, characterized in that, In step S5, the process of deploying the syntax change listener, detecting changes in the table structure and fields of the source system, triggering a regression verification task, and updating the syntax fingerprint, alignment mapping table, and structure normalization pipeline is as follows: Deploy a syntax change listener in the source system, connect to metadata logs and data dictionary, capture table structure and field changes, and generate standardized change event records; Introduce a point-in-time snapshot recorder to generate before and after snapshots based on change events, and compare them with existing syntax fingerprints to generate fingerprint differences and a list of affected objects; Initiate a regression verification task based on the list of affected objects, replay key steps of the structure normalization pipeline, and verify the consistency of field naming mapping, format mapping, and association assertion rules. Based on the regression results, grammatical fingerprints and alignment mapping tables are revised, pipeline revision packages are generated, and loaded into the structure normalization pipeline using a grayscale strategy.
8. The enterprise data evaluation and analysis method based on big data according to claim 1, characterized in that, In step S6, the process of deploying the syntax change listener, detecting changes in the table structure and fields of the source system, triggering a regression verification task, and updating the syntax fingerprint, alignment mapping table, and structure normalization pipeline is as follows: Start the bootstrap synthesizer, read the new system table structure and field value range, and synthesize sample records that cover naming boundaries, format boundaries and association boundaries according to the syntax atomic template; Boundary sample records and original sample records are fed into the structure normalization pipeline, full-link logging is enabled, and field difference trajectories and associated difference trajectories are generated. Based on the difference trajectory, mapping rules are generated by naming dimension, format dimension and association path, the initial mapping and association assertion are derived, and the trial operation configuration and traceability mark are output.
9. A big data-based enterprise data evaluation and analysis system, implemented in any one of claims 1-8, characterized in that, include: The scenario ontology syntax rule library construction module is used to decompose ERP, CRM, and supply chain management fields into syntax atoms, annotate entity relationship semantics according to business scenarios, and summarize them into reusable syntax atom templates. The syntax awareness and fingerprint generation module is used to capture field names, data formats and association logic at the data acquisition end, generate syntax fingerprints and write them into the scene ontology syntax rule library for later use. The alignment mapping generation module is used to generate field naming mappings, format mappings, and association mappings based on syntax fingerprints and syntax atomic templates, forming an alignment mapping table as input for structure normalization; The structure normalization pipeline module is used to perform field renaming, format conversion and relational assertion verification according to scene labels and alignment mapping tables, and generate a unified structure intermediate layer table. The syntax change monitoring and regression verification module is used to detect changes in the source system table structure and fields, trigger regression verification tasks, and update the syntax fingerprint, alignment mapping table and structure normalization pipeline. The synthesis and difference derivation module is used to synthesize boundary sample records when a new, unlabeled system is connected, run the structure normalization pipeline, and derive the initial mapping and association assertions based on the difference trajectory.
Citation Information
Cited By
Deepseek-enabled enterprise unstructured business data intelligent analysis method
CN122241128A