Data processing method, device and equipment and computer readable storage medium

By generating unified verification rules and data quality control strategies for cross-organizational data interaction, the problem of data silos has been solved, automated data verification and standardized transformation have been achieved, and data sharing efficiency and business collaboration capabilities have been improved.

CN121835663APending Publication Date: 2026-04-10CHINA CONSTRUCTION BANK +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Due to independent system construction and inconsistent data standards, data flow between various service organizations is inefficient, resulting in data silos that cannot meet the needs of business collaboration and data sharing.

Method used

By generating unified data verification rules for cross-institutional data interaction and combining them with data quality control strategies, the integration and standardization of heterogeneous data can be achieved, ensuring data format consistency and interoperability.

Benefits of technology

It enables automated verification and standardized transformation of cross-organizational data, breaks down data silos, improves data sharing efficiency, and supports cross-system business applications and decision analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121835663A_ABST
    Figure CN121835663A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, device and equipment and a computer readable storage medium, and relates to the technical field of data processing.The method comprises the steps that data structure rules and service data of different mechanisms are obtained, a data verification rule is generated based on the data structure rules and a preset data exchange rule, and the data verification rule is sent to a server; the data exchange rule is used for defining a structure standard of data interaction between mechanisms; performing data structure verification on the business data of the different mechanisms by using a data verification rule, and screening out a target mechanism passing verification according to a verification result; and performing data fusion on the business data of the target mechanism to obtain shared business data. And a data verification rule is obtained through fusion of a private rule and a global exchange rule of each mechanism to verify the service data, so that the consistency and interoperability of a service data format are ensured. And on the basis, fusion of different-source data is realized through a data quality control strategy, so that the problem of data islands is effectively solved, and the sharing efficiency of the data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of data processing, and in particular, to a data processing method, device, equipment and computer readable storage medium. BACKGROUND

[0002] In the field of public services, realizing data sharing between various service agencies is a key foundation for improving service quality and promoting departmental collaborative governance and service innovation.

[0003] However, due to independent system construction of agencies and non-uniform data standards, data between agencies has low circulation efficiency, resulting in data islands. SUMMARY

[0004] Embodiments of the present application provide a data processing method, device, equipment and computer readable storage medium, which verify business data by fusing private rules of each agency and global exchange rules to ensure consistency and interoperability of business data format. On this basis, the fusion of heterogeneous data is realized by data quality control strategy, thereby effectively solving the problem of data islands and improving the sharing efficiency of data.

[0005] To achieve the above purpose, embodiments of the present application adopt the following technical solutions: In a first aspect, a data processing method is provided, comprising: First, the data structure rules of different agencies and business data are obtained, wherein the data structure rules are used to define the data organization form of the corresponding agency, and the business data is the business data in the field of public services. Second, based on the data structure rules and the preset data exchange rules, a target data verification rule is generated, wherein the data exchange rules are used to define the data structure standard for data interaction between agencies. Then, the data structure of the business data of different agencies is verified based on the target data verification rule, and target agencies with a data structure verification result of pass are selected from different agencies according to the data structure verification result. Finally, the business data of the target agencies is fused based on the preset data quality control strategy to obtain shared business data, which is used to support data access between agencies or external systems.

[0006] The data processing method provided in the embodiments of the present application first takes the global preset data rule as a reference, fuses the private data structure rules of different institutions, and generates a data verification rule that can meet the cross-institution data interaction standard and is compatible with the existing heterogeneous data characteristics. On this basis, the business data of different institutions is verified according to the data verification rule, and the data from different institutions and with different structures can be automatically integrated into a data set conforming to the unified data standard, thereby realizing standardization while preserving data value. Finally, the business data of the target institution is fused by executing the preset data quality control strategy, and shared business data that can reliably support cross-institution business collaboration and decision analysis is obtained. In summary, compared with the data processing method relying on manual consultation and customized development in the related art, the present application takes the preset data exchange rule as a reference, converts the private rules of each institution into verification logic that can be automatically executed, and then realizes automatic verification and standardized conversion of multi-source data, and finally forms shared data that can directly support cross-system business applications, effectively solves the data island problem, and improves the sharing efficiency of data.

[0007] In a possible implementation form of the first aspect, the target data verification rule is generated based on the data structure rules and the preset data exchange rule, including: performing feature analysis on each data structure rule, extracting common features among the rules, and generating a first data verification rule based on the common features. Cross-comparing each data structure rule, identifying conflict rule items in each data structure rule, and generating a conflict rule set. Based on the data exchange rule, the conflict rule set is conflict-resolved to generate a second data verification rule. The target data verification rule is generated based on the first data verification rule and the second data verification rule.

[0008] It should be understood that the implementation form can systematically compatible with the existing data characteristics and resolve rule conflicts through the two-stage rule generation mechanism of "common feature extraction" and "conflict resolution". This not only improves the comprehensiveness and efficiency of the verification rule generation, but also ensures that the generated rule has wide inclusiveness and strict consistency, provides accurate and reliable basis for subsequent data verification, and guarantees the feasibility of data integration from the source.

[0009] In a possible implementation form of the first aspect, the generating the target data check rule based on the first data check rule and the second data check rule comprises: generating a third data check rule based on the first data check rule and the second data check rule. Historical business data of different institutions is obtained as test data. The test data is verified based on the third data check rule, and rule verification feedback information is obtained, the rule verification feedback information comprising rule misjudgment records occurring when the test data is verified. Based on the rule verification feedback information, the check threshold or the logic condition in the third data check rule is optimized to obtain the target data check rule.

[0010] It should be understood that this implementation introduces a rule verification and optimization closed loop based on historical data, which tests and calibrates the initially generated rule by using real business data samples. By analyzing the false positives and false negatives, the discrimination threshold and logic condition of the rule can be optimized accordingly, thereby significantly improving the accuracy and practicality of the data check rule, reducing the misjudgment of compliant data and the omission of abnormal data in actual application, and further enhancing the automation level and reliability of the data sharing process.

[0011] In a possible implementation form of the first aspect, before the data fusion of the business data of the target institution based on the preset data quality control strategy to obtain the shared business data, the method further comprises: for at least one to-be-updated institution whose data structure verification result of the business data is not passed and which is screened out from different institutions, determining data structure update information of each to-be-updated institution based on the data structure verification result of each to-be-updated institution and the target data check rule. Sending the corresponding data structure update information of each to-be-updated institution to each to-be-updated institution. In response to the received business data of the to-be-updated institution, performing structure verification on the business data of the to-be-updated institution, and in the case that the structure verification result is passed, determining the to-be-updated institution as the target institution.

[0012] It should be understood that this implementation builds a dynamic and scalable data access closed loop. It not only passively screens compliant data sources, but also actively provides accurate data structure update guidance to institutions that fail to pass the verification, and re-verifies the data submitted subsequently. This mechanism effectively promotes the standardization of the data standards of participating institutions, gradually expands the scale of qualified data sources, and thus realizes self-improvement and sustainable development of the data sharing ecosystem.

[0013] In a possible implementation form of the first aspect, the data quality control strategy comprises a data integration strategy, a data security strategy, a data storage strategy, a data sharing strategy, and a data lifecycle management strategy. The data integration strategy comprises a method of cleaning, transforming, and fusing business data of the target institution. The data lifecycle management strategy comprises updating a lifecycle state of the shared business data based on a preset rule, the lifecycle state comprising a creation state, a sharing state, an archiving state, and a destruction state.

[0014] It should be understood that this implementation form ensures the quality, security, and availability of the shared data in the entire lifecycle by defining a complete and multi-dimensional data quality control strategy. The data integration strategy is responsible for transforming multi-source data into unified, clean, and usable data assets; the data security and sharing strategy regulates data access and control; and the data lifecycle management strategy realizes fine management and control of data cost and value. This combined strategy collectively ensures that the final output of the shared business data is reliable, secure, and sustainable.

[0015] In a possible implementation form of the first aspect, the method further comprises: obtaining state snapshot data of the shared business data at different time points, the state snapshot data comprising feature parameters of the shared business data aggregated according to preset dimensions. Trend fitting and comparative analysis are performed on the state snapshot data to determine the evolution mode and correlation of the feature parameters corresponding to different dimensions. Based on the evolution mode and correlation of the feature parameters corresponding to different dimensions, a business trend analysis report is generated, which is used to support business decision-making.

[0016] It should be understood that this implementation form deeply mines the potential value of the shared business data, enabling it to change from a simple information carrier to a strategic asset supporting decision-making. Through trend analysis and correlation mining of historical state snapshots, the internal laws and potential problems of business development can be revealed, thereby generating a forward-looking and instructive business trend analysis report, ultimately realizing data-driven decision-making, and improving the intelligent management and service level of the entire public service system.

[0017] In a second aspect, a data processing apparatus is provided, comprising: An obtaining module is configured to obtain data structure rules of different institutions and business data. The data structure rules are used to define the data organization form of the corresponding institution. The business data is business data in the field of public services.

[0018] The processing module is configured to generate a target data verification rule based on the data structure rule and a preset data exchange rule. The data exchange rule is configured to define a data structure standard for data interaction between institutions. The target data verification rule is used to verify the data structure of the business data of different institutions, and based on the data structure verification result, a target institution with a passed data structure verification result is selected from the different institutions. The business data of the target institution is fused based on a preset data quality control strategy to obtain shared business data. The shared business data is used to support data access between institutions or external systems.

[0019] In a third aspect, a data processing device is provided, which includes a memory and at least one processor. The memory is in communication connection with the processor. The memory is configured to store computer program code including computer instructions. When the processor executes the computer instructions, the data processing device performs the method according to the first aspect and any possible implementation manner thereof.

[0020] In a fourth aspect, a computer readable storage medium is provided, which stores computer instructions. When the computer instructions are executed by a processor, the method according to the first aspect and any possible implementation manner thereof is implemented.

[0021] In a fifth aspect, a computer program product is provided, which, when executed on a computer / processor of a computer, implements the method according to the first aspect and any possible implementation manner thereof. The computer can be the data processing device according to the third aspect and any possible implementation manner thereof.

[0022] It can be understood that the beneficial effects of the data processing apparatus according to the second aspect, the data processing device according to the third aspect, the computer readable storage medium according to the fourth aspect, and the computer program product according to the fifth aspect can refer to the beneficial effects of the first aspect and any possible implementation manner thereof, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 A schematic diagram of a computing device interaction scenario is provided for an embodiment of the present application. Figure 2 A flowchart of a data processing method is provided for an embodiment of the present application. Figure 3 A flowchart of a data verification rule generation method is provided for an embodiment of the present application. Figure 4 A flowchart of another data verification rule generation method is provided for an embodiment of the present application. Figure 5Another data verification rule generation method flow diagram provided by the embodiment of the application is shown in FIG. 6; Figure 6 Another data processing method flow diagram provided by the embodiment of the application is shown in FIG. 7; Figure 7 Another data processing method flow diagram provided by the embodiment of the application is shown in FIG. 8; Figure 8 A structure diagram of a data processing device provided by the embodiment of the application is shown in FIG. 9; Figure 9 A structure diagram of a data processing device provided by the embodiment of the application is shown in FIG. 10. DETAILED DESCRIPTION

[0024] Hereinafter, the terms "first" and "second" are only used for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features. In the description of the embodiments, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0025] The exemplary embodiments will be described in detail below, and examples thereof are shown in the accompanying drawings. In the following description of the embodiments, the same numbers refer to the same or similar elements unless otherwise specified. The embodiments described in the following exemplary embodiments do not represent all the embodiments consistent with the present application. Rather, they are merely examples of devices and methods consistent with some aspects of the present application, as detailed in the appended claims.

[0026] In the technical solutions provided by the embodiments of the present application, the collection, storage, use, processing, transmission, provision and disclosure of information such as financial data or user data, etc. comply with relevant laws and regulations and do not violate public order and good customs.

[0027] It should be noted that in the embodiments of the present application, some industry existing schemes such as software, components, models, etc. may be mentioned, which should be considered as exemplary, and the purpose is only to illustrate the feasibility of the implementation of the technical solutions of the present application, but it does not mean that the applicant has or will necessarily use the scheme.

[0028] In the field of public services, such as public fund management, the demand for cross-regional and cross-center business collaboration is increasing, but the data island phenomenon seriously hinders this process. The direct technical root of the formation of the data island is that the information systems independently constructed by each institution adopt incompatible data structure rules. When these heterogeneous data are needed for cross-institution business, the differences in data format and standards prevent direct data flow and batch processing between systems. This incompatibility is manifested in two points: first, a large amount of repetitive manual data conversion and comparison work is required before any data exchange, resulting in a long data processing process that cannot meet the timeliness requirements of business development; second, due to the lack of unified and automated control of data quality, the integrated data quality is uneven and inconsistent, which cannot reliably support upper-level business applications.

[0029] Therefore, how to break through the data island and achieve efficient and reliable data integration and sharing has become a key technical problem to improve the efficiency of public services.

[0030] In view of this, the embodiment of the present application provides a data processing method, which combines the preset data exchange rules with the private data structure rules of each institution to generate executable data verification rules; then uses the rules to verify the data structure of multi-source heterogeneous business data, and integrates the originally different data structures into a data set that meets the unified standard; finally, by executing the data quality control strategy covering data integration, security and life cycle, the business data of the target institution is fused to obtain shared business data that can directly support cross-institution business access. The data island problem caused by the heterogeneous data standards between different institutions is effectively solved, and the sharing efficiency of business data in cross-institution scenarios is improved.

[0031] The data processing method provided by the embodiment of the present application can be applied in a computing device. Specifically, the computing device can be a single server or a server cluster composed of multiple servers, or a computer, or a processor or processing chip in a server or computer, etc. The embodiment of the present application does not limit the specific device form of the computing device.

[0032] As shown in Figure 1 , it is an interactive scene diagram of a computing device provided by the embodiment of the present application, wherein the interactive objects in the interactive scene can include: an institution, a computing device, an external system and a supervision platform, and the computing device is in communication connection with the institution, the external system and the supervision platform.

[0033] Wherein, the institution refers to an entity providing business data, such as a provident fund center institution, a social security institution, or a real estate registration center. The institution layer includes multiple institutions, represents multiple institutions as data source parties, and has an architectural relationship of centralized interaction with the computing device. Embodiments of the present application do not limit the specific type of institution. The external system refers to a platform or service that has carried out its own business to demand business data, such as a big data analysis platform, an analysis tool of a third-party research institution, or a mobile application for public government services. The supervision platform refers to an institution that supervises institutions in the field, that is, the superior supervisory institution of the institution. The platform can issue globally unified data exchange rules and use shared business data for business decision-making.

[0034] First, the data processing inside the computing device includes: Data verification rule generation. Based on the obtained data structure rules of each institution and the preset data exchange rules, a unified data verification rule is generated.

[0035] Data verification and screening. Based on the data verification rule, the business data of each institution is verified. For institutions with a verification result of pass, data governance and fusion are directly performed. For institutions with a verification result of not pass, data structure update information is sent to the institution based on the data verification rule.

[0036] Data governance. Data governance can include data integration, data storage, data security, data sharing, and data lifecycle management. In some embodiments, data governance can also be understood as data quality control.

[0037] Among them, data integration includes data cleaning, format conversion, value mapping, and data fusion processing of target institution business data, eliminating data differences and ensuring data consistency.

[0038] Data storage is to persistently save the governed data. Specifically, data storage can be divided into the following dimensions according to data application dimensions: metadata storage, used to store the original business data of each institution; data warehouse storage, used to store shared business data obtained after data governance, which can directly provide data support for data access of the institution; and master data storage, business entity data that can be directly accessed by external systems. Embodiments of the present application do not limit the specific data storage mode.

[0039] Data security is to ensure the security of shared data, including data desensitization processing, access permission control, and data encryption storage.

[0040] Data sharing is used to provide data services to institutions or external systems, including data query interfaces, data subscription services, and batch data export functions.

[0041] In some embodiments, the computing device can also deploy a comprehensive service library for recording data sharing behaviors in view of data sharing. Specifically, the data sharing behavior records can include access time, access frequency, requested data type, and requested system identifier, etc.

[0042] By performing multi-dimensional data analysis on these access records, the computing device can identify data usage patterns and business demand characteristics of different external systems, and further optimize data service strategies, including but not limited to: adjusting the response priority of data interfaces, optimizing data distribution mechanisms, identifying potential data security risks, and providing differentiated data service support for different business scenarios, thereby improving the intelligent level and resource utilization efficiency of data services.

[0043] Data lifecycle management is based on preset rules to update the lifecycle state of shared business data in the data storage module, including the management of creation state, sharing state, archiving state and destruction state.

[0044] Business trend analysis report generation. Based on the characteristic parameters of shared business data, trend analysis and correlation mining are performed to generate a trend analysis report to support business decision-making.

[0045] The interaction between the computing device and the institutions includes: 1. Obtain the data of each institution. That is, obtain the data structure rules and business data of each institution.

[0046] 2. Send data structure update information to the institution. Send data structure update information to the institution that fails the data verification (i.e., the institution to be updated).

[0047] 3. Receive updated business data. The computing device receives the updated business data from the institution and performs secondary verification on the data. If it passes, the institution that failed the last data verification is determined as the target institution.

[0048] 4. Provide shared business data to the institution.

[0049] The interaction between the computing device and the external systems includes: 5. Provide shared business data to the external systems.

[0050] In some embodiments, the external systems obtain shared business data by displaying data through corresponding clients (such as web pages, mobile applications, or mini programs, etc.), meeting the needs of external businesses.

[0051] The interaction between the computing device and the regulatory platform includes: 6. Obtain data exchange rules. The computing device receives and loads the data exchange rules formulated and published by the regulatory platform.

[0052] 7. Providing a business situation analysis report to an external system. The computing device sends a business situation analysis report based on the shared business data analysis to a regulatory platform.

[0053] It should be noted that the above sequence numbering does not constitute the interactive steps of the present application.

[0054] In some embodiments, the computing device can also receive policy configuration and control instructions from the regulatory platform, such as adjustment of data access range, change of data security level, or setting of data retention period, so that the entire data sharing system can operate under controlled and auditable conditions.

[0055] As shown in Figure 2 The data processing method provided by the embodiments of the present application, when applied to the computing device, specifically includes the following steps S101-S104: S101, obtaining the data structure rules and business data of different institutions.

[0056] The data structure rules are used to define the data organization form of the corresponding institution. The data organization form is the organization framework and constraint standard of the data, specifically including the format constraint, hierarchical division, and association logic of the data.

[0057] Specifically, the data structure rules can include naming specification of data fields, type of field value, length limit, coding rules, data hierarchical division logic, association field mapping relationship, data statistical caliber standard, etc. The data structure rules depend on the server system architecture of the corresponding institution, and the embodiments of the present application do not limit the specific content of the data structure rules.

[0058] The business data is the business data of the public service field. The business data refers to the data (records) generated and accumulated by the public service institution in the process of handling specific business.

[0059] For example, the business data can be business data of public accumulation fund management scenarios, social security scenarios, real estate registration scenarios, medical insurance reimbursement scenarios, etc., and correspondingly, the institutions can include local public accumulation fund management centers, social insurance administration bureaus, real estate registration centers, medical insurance agencies, etc. The embodiments of the present application do not limit the specific types of business data and the specific types of institutions.

[0060] In some embodiments, the process of obtaining the data structure rules and business data of different institutions by the computing device can include active acquisition or passive acquisition. Different acquisition methods correspond to different technical conditions and data providing habits of different institutions.

[0061] In one possible implementation, the computing device can automatically trigger data acquisition based on a timer. The data can be automatically extracted from the data structure rule document and the business data of the institution by calling the API interface opened by the institution or by accessing the institution database based on the connection protocol.

[0062] For example, the computing device establishes a database communication connection with the information system of the social security center of B city, and sets a preset timer to acquire data at 3 a.m. every day or sets a trigger condition of detecting that the data increment of the social security center business database exceeds 100, and when the timer condition or the trigger condition is met, the computing device executes a corresponding acquisition request.

[0063] In another possible implementation, the computing device can passively receive data submitted by the institution. The computing device waits for the staff of the institution to manually submit data through a client (such as a web page, a mobile application, or a mini program) or a file transmission channel (for example, a file server) for uploading data.

[0064] For example, the computing device deploys a web page with an identity authentication function, and the staff of the real estate registration center of C city logs in to the web page through a registered account and password, selects and uploads an Excel format data structure rule file in a “rule file upload” module, and uploads a CSV format business data file containing real estate registration information and right holder information in a “business data upload” module. After the computing device receives the above files, the computing device feeds back a notification of successful uploading to the staff through an email after performing a legal verification on the data.

[0065] In some embodiments, the data of the institution can also be stored in a blockchain platform to ensure the security of the data. The computing device accesses the data of the blockchain platform through a prior agreement with the institution, and acquires the data according to a data chaining rule.

[0066] In some embodiments, after acquiring the data structure rule and the business data of each institution, the computing device can also perform a preliminary integrity verification on the received business data, for example, by comparing the actual received data file size with a checksum provided by the source institution, or verifying whether the number of records in the batch data file is consistent with the number declared in the message, to ensure that no data loss occurs in the transmission process.

[0067] S102, generating a target data verification rule based on the data structure rule and a preset data exchange rule.

[0068] The data exchange rule refers to a data structure standard uniformly formulated to guarantee data interoperability in a cross-institution data sharing scenario, which defines the format, semantics and constraint specifications that an institution needs to comply with when interacting with data. The target data verification rule is a rule for verifying the business data of an institution, which is generated by fusing the data structure rules of each institution and the preset data exchange rule.

[0069] Specifically, in the related data sharing method, the data problems faced by each institution in data sharing include the following: Data specifications are not unified. Each institution uses private field naming and format specifications. For example, for the business concept of "personal public accumulation account balance", different institutions may use "personal account balance", "account balance" and other field names, which leads to semantic recognition difficulties in data interaction and cannot be directly compatible.

[0070] Data code values are not unified. The coding system of the same business concept is different. For example, "depositing state" may be defined as {1, 2, 3}, {normal, frozen, closed} or {M, F} in different institutions, which forces the data user to perform tedious mapping and conversion, significantly increasing the data processing complexity and use cost.

[0071] Data statistical standards are not unified. The statistical standards of key business indicators are inconsistent. For example, "depositing base" may be calculated based on "employee's last year monthly average salary", "employee's monthly salary" or "fixed deposit amount" in different institutions, which leads to ambiguity in cross-institution comparison of statistical reports and decision-making data, seriously affecting the credibility of the data and the value of decision support.

[0072] The above problems collectively lead to the phenomenon of data silos, where each institution's data environment is isolated from each other, making it difficult to systematically improve data quality and slow down data governance work. Based on this, the present application builds a unified data exchange rule system and performs semantic fusion and logical resolution with private rules of each institution to form a target data verification rule that is globally consistent and individually compatible.

[0073] In some embodiments, the computing device can obtain the data exchange rule by reading a preset global configuration file. The data exchange rule is pre-established by the superior department of the institution according to the business requirements and data integration standards in the public service field. The specific source of the data exchange rule is not limited in the embodiments of the present application.

[0074] In some embodiments, the target data verification rules can include format verification rules (e.g., the ID number field is 18 characters), code value verification rules (e.g., the deposit status code value is only allowed to be 01, 02, 03), field mapping verification rules (e.g., the "account balance" field of the institution needs to correspond to the global field "accumulation fund account balance"), and consistency verification rules (e.g., the deposit base calculation method needs to comply with the global statistical caliber). The embodiments of the present application do not limit the specific content of the target data verification rules.

[0075] In some embodiments, the computing device can generate target data verification rules by using a maximum compatibility algorithm. Specifically, the fields in the data exchange rules are taken as reference fields, and through semantic matching and business relevance analysis, a set of comparison fields corresponding to each reference field in the data structure rules of each institution is determined. Then, based on the constraints of the reference fields in the data exchange rules, the most extensive minimum attribute constraints are extracted from the corresponding set of comparison fields, which are converted into target data verification rule items corresponding to the reference fields, to realize the compatibility of global standards and multi-institution data characteristics.

[0076] In one possible implementation, the computing device deploys a maximum compatibility algorithm engine, which includes a field matching module, a compatible attribute extraction module, and a verification rule generation module. The field matching module calculates the matching degree of the field name and description by using a semantic similarity algorithm, and filters the set of comparison fields corresponding to the reference fields in combination with a pre-set business logic association mapping table. The compatible attribute extraction module counts the number of institutions covered by each attribute requirement in the set of comparison fields, and selects the minimum attribute requirement (e.g., the minimum precision, the most basic data type, and the most core semantic range) covering the most institutions as the maximum compatible attribute constraint. The verification rule generation module converts the maximum compatible attribute constraint into a standardized target data verification rule item, to ensure that the rule complies with both the global exchange standard and the actual data of most institutions.

[0077] For example, the constraint of the global reference field "deposit base" in the data exchange rules is "numeric, greater than 0". The field matching module filters out the corresponding fields of 10 institutions as the set of comparison fields through semantic similarity calculation and business relevance analysis. The compatible attribute extraction module finds that the attribute requirements of the fields of 8 institutions are "integer, greater than 0", and the attribute requirements of the fields of 2 institutions are "float, greater than 0 and retaining two decimal places". Therefore, "integer, greater than 0" covering the most institutions (8 institutions) is selected as the maximum compatible attribute constraint. The verification rule generation module converts it into the verification rule item "field value is an integer greater than 0" accordingly, to ensure that the data of most institutions can pass the verification without additional adjustment, while meeting the constraint of the global exchange rules.

[0078] In addition, the above embodiments generate the check rules at the level of the fields, and the core inventive concept of the generation of the check rules of the hierarchy, the association, etc. is the same as the maximum compatibility algorithm, that is, the same idea of following the global data exchange rules as the benchmark, comparing and adapting with the rules of the private data structure of each institution, and finally extracting the maximum compatibility constraints to form the check rules.

[0079] Specifically, for the hierarchy of the data, the computing device can parse the global ideal data model defined in the data exchange rules, and compare it with the actual hierarchical path defined in the rules of the data structure of each institution. By analyzing the actual hierarchical paths of all institutions, the computing device can find and adopt the simplest and effective hierarchical path that can cover the most actual data of the institutions, and convert it into a hierarchical existence check rule.

[0080] For example, the global data model requires that the “unit information” contains the “employee list”. The computing device finds through comparison that about 70% of the institutions use the direct hierarchy of “unit information -> employee list”, and 30% of the institutions use the indirect hierarchy of “unit information -> department information -> employee list”. The computing device will select “unit information -> employee list” as the maximum compatible hierarchical path that covers the majority of institutions, and generate a rule to verify whether the effective path exists in the data.

[0081] Specifically, for the association of the fields, the computing device can identify the core business association that is required by the data exchange rules. Then, the computing device analyzes the specific technical logic used by the rules of the data structure of each institution to implement the association. By counting the distribution of these specific logics, the computing device can select the association logic that is used by the most institutions and meets the business consistency as the benchmark, and generate an association relationship check rule.

[0082] For example, the data exchange rules require that the “deposit record” must be associated with the “personal account”. The computing device analyzes and finds that 12 institutions are directly associated through the “personal account number”, 5 institutions are indirectly associated through the “business flow number”, and 3 institutions are associated through the “citizen ID number”. The computing device will select the “personal account number” that supports the most institutions as the maximum compatible association key, and generate a rule to verify the validity and success rate of parsing of the association key.

[0083] In some other embodiments, the computing device can also extract the common rules and conflict rules in the rules of the data structure of each institution, and then dissolve the conflict rules based on the preset data exchange rules, fuse the common rules and the adapted rules after the dissolution, and generate the target data check rules that meet the global standards and are suitable for the characteristics of the data of the institutions.

[0084] The process can refer to the following Figure 3and specific descriptions thereof are not described in detail here.

[0085] S103, based on the target data verification rule, verifying the data structure of the business data of different institutions, and according to the data structure verification result, screening out a target institution with a passed data structure verification result from the different institutions.

[0086] In some embodiments, the computing device can adopt a traversal verification mechanism combining record-by-record and field-by-field. Specifically, the computing device reads each business data record of each institution in turn, and verifies the fields, hierarchical structure and association in each record according to the loaded target data verification rule set.

[0087] In a possible implementation, the computing device implements the target data verification rule as a rule verifier. The rule verifier develops a corresponding verifier unit for each type of verification rule (such as format verification, code value verification, and dimension consistency verification). When verifying a business data record, the rule verifier performs semantic retrieval based on the target data verification rule, acquires the target data verification rule item corresponding to the business data record according to the semantic retrieval result, and calls the corresponding verifier unit to form a verification sequence. Each verifier unit runs independently, and its verification result (pass or fail) is collected into the verification result set of the record.

[0088] For example, when verifying a “personal deposit record” of the A city public accumulation fund center, the computing device will call the ID number format verifier unit, the deposit status code value verifier unit, and the deposit amount logical relationship verifier unit in turn. Only when all the called verifiers return “pass”, the record is marked as verified.

[0089] In some embodiments, for the case of excessive business data volume, the computing device can use distributed parallel processing technology to improve the throughput of business data verification. Specifically, the computing device divides the business data to be verified by institution or by data shard, and allocates independent computing nodes for each data shard to verify.

[0090] For example, the computing device needs to verify the batch settlement data of 30 medical insurance centers. The computing device divides the total data of about 5 TB into 30 partitions according to the medical insurance centers, and submits to a cluster composed of 50 computing nodes. The cluster starts 30 parallel tasks at the same time, each task loads unified verification rules, and independently processes all data of a medical insurance center. Finally, the verification results of each node are aggregated to generate a global verification report.

[0091] In some embodiments, the computing device can determine the target institution based on a principle of full-quantity verification pass. The computing device determines an institution as a target institution only when all business data records of the institution pass the verification after completing the verification on all business data of the institution. If there is any record that does not pass the verification, the institution will not be included in the list of target institutions.

[0092] In other embodiments, the computing device can determine the target institution by using a candidate institution promotion mechanism based on automatic repair. The candidate institution refers to an institution whose verification pass rate of business data records reaches a preset pass rate threshold. The automatic repair refers to a process of automatically repairing the data records that do not pass the verification in the candidate institution based on the target data verification rule. The promotion mechanism refers to a process flow of converting the candidate institution to the target institution after the data repair and re-verification.

[0093] For example, the preset pass rate threshold can be 90%, 95%, 99%, etc. The embodiments of the present application do not limit the specific value of the preset pass rate threshold.

[0094] For example, the computing device sets the candidate threshold to 98%. When verifying the data of the E City Social Insurance Bureau, it is found that the pass rate is 98.5%, and the computing device marks it as a candidate institution. The computing device automatically performs data repair on the 1500 records that do not pass the verification. Then, after verifying and confirming that the business data records pass the verification, the computing device determines the E City Social Insurance Bureau as a target institution.

[0095] It should be understood that, by using the candidate institution promotion mechanism based on automatic repair, the inclusiveness of institution screening is improved under the premise of ensuring data quality, the rapid conversion of institutions whose data verification does not pass is achieved through automatic processing, and the implementation efficiency of cross-institution data sharing is effectively improved.

[0096] S104, data fusion is performed on the business data of the target institution based on a preset data quality control strategy, and shared business data is obtained.

[0097] The shared business data is used to support data access between institutions or external systems.

[0098] Specifically, the data quality control strategy aims to deeply process and optimize the packaging of the business data of the target institution, so as to improve the discoverability, manageability and use efficiency of the business data.

[0099] In some embodiments, the preset data quality control strategy can include a data integration strategy, a data security strategy, a data storage strategy, a data sharing strategy and a data life cycle management strategy.

[0100] Specifically, the data integration strategy includes a method of cleaning, converting and fusing the business data of the target institution.

[0101] Data integration strategy includes data standardization, data correlation fusion, data asset cataloging, etc.

[0102] Data compression processing refers to encoding and compressing numerical fields in business data to reduce storage and transmission overhead. The computing device analyzes the data distribution characteristics of the numerical values, uses dictionary encoding, incremental encoding or frame compression technology, etc. to convert the original numerical values into a compact binary format, while ensuring that the compression process is lossless and can be accurately restored.

[0103] Data correlation fusion refers to associating and merging data from different target institutions that describe the same entity through entity recognition and relationship discovery technology. The computing device builds a unified entity view across institutions based on determined matching keys or fuzzy matching algorithms, and solves data conflicts to form more complete and accurate master data.

[0104] Data asset cataloging refers to creating discoverable and manageable metadata archives for shared business data after governance. The computing device automatically extracts the structural information, business semantics and management information of the data and registers them in a unified data directory. This directory provides an entry for external systems to query, understand and apply for access to data assets, and is a key technical means to achieve effective discovery and reuse of data.

[0105] Data security strategy refers to technical rules for classifying and grading shared data, access control, encryption protection and desensitization processing. The computing device automatically labels data security levels through a sensitive information identification model and implements dynamic desensitization for core privacy fields: generalization processing in test environment, and reservation of some key fields in production environment. At the same time, based on attribute encryption technology, it implements hierarchical access control and uses national encryption algorithm for full encryption during data transmission and storage.

[0106] Data storage strategy refers to technical rules for persistently saving shared business data after governance according to purpose.

[0107] For example, the computing device stores data used for analysis and decision-making in a data warehouse according to a subject domain model, stores core entity data supporting real-time business services in a main database, and stores information such as blood relationship and quality describing the characteristics of the data in a metadata database. The strategy clearly defines the synchronization mechanism and life cycle of the three types of stored data. The main database, as the data basis for cross-institution collaboration, needs to ensure the timeliness of batch synchronization between it and the data warehouse, and ensure real-time updates to the metadata database.

[0108] The data sharing strategy refers to a technical specification for providing data services externally based on a hierarchical storage system, and the core is to establish an access permission system linked with the storage hierarchy. The computing device opens a data catalog query permission to authorized users through a metadata database, grants real-time read and write permissions of core data to external systems or institutions through a main database, and grants batch data read-only permissions to an analysis platform through a data warehouse. The strategy requires defining a clear access control list (ACL) for each data interface and establishing a permission application and automatic approval process.

[0109] The data lifecycle management strategy includes updating the lifecycle state of shared business data based on preset rules, and the lifecycle state includes a creation state, a sharing state, an archiving state, and a destruction state.

[0110] In one possible implementation, the computing device can implement the above-mentioned data quality control strategy as a data governance framework. The model implements strategy execution through a modular architecture, and its core feature is intelligent decision-making capability: the integrated strategy realizes intelligent data fusion through a pipeline with pattern recognition capability; the security strategy realizes dynamic access control through an integrated sensitive information recognition sub-model; the storage and sharing strategy realizes intelligent storage allocation through a load-aware router; and the lifecycle strategy is driven by a data utility evaluation model. The parameters of each strategy exist in the form of configurable metadata, and the computing device can dynamically adjust the execution logic according to the data characteristics when running the model, realizing the intelligentization and adaptive optimization of the governance process.

[0111] In another possible implementation, the computing device can implement the above-mentioned data quality control strategy as a multi-agent subsystem. The subsystem is composed of multiple agents that perform their respective functions: the integrated agent actively discovers and correlates multi-source data; the security agent dynamically assesses risks and implements control; the storage agent collaborates to complete resource allocation and scheduling; and the lifecycle agent drives data flow based on global state. Through communication and collaboration, each agent achieves its own strategy goal while collectively implementing autonomous and collaborative governance of the entire data lifecycle.

[0112] It should be understood that this implementation ensures the quality, security, and availability of shared data throughout its entire lifecycle by defining a complete and multi-dimensional data quality control strategy. The data integration strategy is responsible for transforming multi-source data into unified, clean, and usable data assets; the data security and sharing strategy regulates data access and control; and the data lifecycle management strategy realizes fine-grained control of data cost and value. This combined strategy collectively ensures that the final output of shared business data is reliable, secure, and sustainable.

[0113] The data processing method provided by the embodiments of the present application first takes the global preset data rule as a reference, fuses the private data structure rules of different institutions, and generates a data verification rule that can meet the cross-institution data interaction standard and is compatible with the existing heterogeneous data characteristics. On this basis, the business data of different institutions is verified according to the data verification rule, and the data from different institutions and with different structures can be automatically integrated into a data set conforming to the unified data standard, thereby realizing standardization while preserving data value. Finally, the business data of the target institution is fused by executing the preset data quality control strategy, and shared business data that can reliably support cross-institution business collaboration and decision analysis is obtained. In summary, compared with the data processing method relying on manual consultation and customized development in the related art, the preset data exchange rule is taken as a reference in the present application, the private rules of each institution are converted into verification logic that can be automatically executed, and then the automatic verification and standardized conversion of multi-source data are realized, and finally the shared data that can directly support cross-system business applications is formed, effectively solving the data island problem and improving the sharing efficiency of data.

[0114] The generation process of the target data verification rule is specifically introduced below in combination with the accompanying drawings of the specification.

[0115] In some embodiments, the process in which the computing device generates the target data verification rule through rule resolution and fusion is as shown in Figure 3 S102 specifically includes the following steps S201-S204. S201, performing feature analysis on each data structure rule, extracting common features among the rules, and generating a first data verification rule based on the common features.

[0116] The feature analysis refers to that the computing device scans and statistics the data structure rules of the institutions, aims to find the frequently appearing fields, format constraints or logical relationships (i.e., common features) therein, so as to identify the common data patterns and standards among the institutions. The first data verification rule refers to the common verification rule generated based on the common data patterns and standards.

[0117] In some embodiments, the computing device can use a common feature mining method based on feature vectors and clustering analysis. The computing device converts the structure rule items of each institution into numerical feature vectors, and identifies the rule patterns that are gathered in the vector space and cross multiple institutions through a clustering algorithm, and defines the common specifications represented by these patterns as common features.

[0118] In one possible implementation, the computing device maps the rule texts into semantic vectors using a word embedding model, and groups the full set of rule vectors using a distributed clustering algorithm. The computing device counts the number of institutions covered by each cluster, and identifies the rule semantics represented by the cluster centers that cover more than a pre-set threshold of institutions as common characteristics.

[0119] For example, the computing device discovers through clustering that the fields describing the concept of “personal account” in multiple institutions are clustered into a high-density cluster in the vector space, with their corresponding data types (integer), length constraints (18 bits), and encoding rules (not allowed to be null) etc. This cluster covers 80% of the institutions, and thus the computing device identifies the set of specific data specifications {field name semantics: “personal account”, data type: “integer”, length: 18, null value constraint: “NOTNULL”} as a common characteristic.

[0120] In some other embodiments, the computing device can use a common characteristic discovery method based on text similarity calculation. The core of this method is that when the rule descriptions of the same business concept in different institutions are highly similar in text form, it is considered that they represent the same underlying data specification. The computing device calculates the text similarity of the rule descriptions between institutions, and defines the common specification represented by the rule cluster that has a similarity exceeding a threshold and covers a sufficient number of institutions as a common characteristic.

[0121] In one possible implementation, the computing device uses the edit distance as the similarity measure of the rule description texts. The computing device first compares all the rule descriptions of the institutions pairwise, and calculates the edit distance between each pair of rule descriptions. Then, the computing device constructs a similarity graph of the rule descriptions based on the distances, and identifies the tight communities in the graph using a community discovery algorithm (such as a graph clustering algorithm). Finally, the computing device identifies the rules corresponding to the communities that cover more than a pre-set threshold of institutions as common characteristics.

[0122] The edit distance refers to the minimum number of single-character editing operations required to convert one string into another, including inserting a character, deleting a character, or replacing a character with another character.

[0123] For example, when comparing the rule descriptions of “deposit base”, the computing device finds that A institution describes it as “monthly average salary of employees in the last year”, B institution describes it as “monthly average salary of staff”, and C institution describes it as “monthly average salary of employees in the last year”. Through calculating the edit distance and clustering, these three highly similar descriptions are grouped into the same community. Since this community covers 75% of the institutions participating in the comparison, the computing device extracts the core semantics “based on the monthly average salary in the last year” of the community as a common characteristic.

[0124] S202. Cross-compare the rules of each data structure to identify conflicting rule items and generate a set of conflicting rules.

[0125] Cross-comparison refers to the process by which computing devices perform pairwise or many-to-many comparisons of the private data structure rules of all organizations to identify rule items with consistent field semantics but contradictory attribute constraints (data type, value range, format, etc.). Conflicting rule items refer to incompatible attribute requirements existing in different organization rules for the same business semantic field, such as defining mutually exclusive data types, value ranges, or formats for the same field. A conflict rule set is a structured collection formed by classifying and organizing all conflicting rule items according to semantic fields. Its purpose is to clarify the contradictions between rules from multiple organizations and provide a basis for subsequent conflict resolution.

[0126] In some embodiments, the computing device can achieve systematic cross-matching by constructing a rule conflict detection matrix. The rows and columns of this matrix represent all participating institutions, and the computing device automatically schedules and executes the rule matching task for each pair of institutions in the matrix.

[0127] One possible implementation involves the computing device instantiating the conflict detection matrix using a task queue and a distributed computing framework. The computing device defines each cell in the matrix (i.e., each specific pair of institutions) as an independent alignment task and submits it to a task queue. Multiple worker nodes in the cluster then pull tasks from the queue, execute the rule-based alignment logic for the assigned institution pairs in parallel, and write the detected conflict results back to a shared storage system.

[0128] In some embodiments, during the cross-comparison of rules across data structures, the computing device may employ a constraint logic-based conflict detection method. The computing device converts the rules of different structures into formal constraints and identifies logical conflicts by checking the satisfiability of these constraints.

[0129] For example, Institution A defines the "Deposit Status" field as an integer, and its business rules explicitly require the value to be an enumeration set {1, 2, 3}, corresponding to "Normal," "Sealed," and "Closed," respectively. Institution B defines the same "Deposit Status" field as an integer, but its business rules define the enumeration set as {0, 1}, corresponding to "Normal" and "Suspension of Deposit," respectively. The computing device detects through logic checks that the value 1 is assigned different semantics in the business rules of the two institutions ("Normal" for Institution A, "Suspension of Deposit" for Institution B). This mutual exclusion of core business semantics means that no data value can simultaneously satisfy both constraints without violating one institution's business rules. Therefore, the computing device identifies a pair of conflicting rule items and records their conflict type as "Business Logic Conflict."

[0130] For example, the format rule defined by organization A for the "personal ID number" field is 18 digits (regular expression: ^[0-9]{17}[0-9Xx]$), while the format rule defined by organization B for the same field is 15 digits (regular expression: ^[0-9]{15}$). The computing device detects by logic that an 18-digit ID number string cannot pass the 15-digit format check of organization B, and vice versa. These two format constraints are mutually exclusive at the instance level, and there is no data value that can satisfy both format requirements. Therefore, the computing device identifies this as a pair of conflicting rule items, and records the conflict type as "format constraint conflict".

[0131] S203, conflict resolution of the conflict rule set based on the data exchange rule is performed to generate a second data validation rule.

[0132] The conflict resolution refers to a process in which the computing device processes each specific conflict instance (i.e., information of multiple mutually contradictory rule items and conflict logic of each record) recorded in the conflict rule set according to the global data exchange rule.

[0133] Specifically, the conflict resolution can include machine automatic resolution and manual arbitration resolution. The machine automatic resolution refers to automatic processing of the conflict by the computing device according to a preset resolution strategy. The manual arbitration resolution refers to a process in which human judgment is introduced by the computing device for assistance when the machine cannot automatically decide or the conflict involves complex business trade-offs. In addition, the machine automatic resolution and the manual arbitration resolution can be executed in parallel or in series.

[0134] In some embodiments, the process of manual arbitration resolution includes: the computing device screens out complex conflict instances that cannot be covered by the preset resolution strategy or have a confidence level lower than a threshold from the conflict rule set, presents them on an interactive interface, and attaches rule details of the conflicting parties, potential resolution scheme recommendations, and business impact analysis. The computing device receives a final arbitration instruction made by an authorized user through the interface, and generates or adjusts the corresponding second data validation rule according to the instruction.

[0135] In some embodiments, in the process of performing conflict resolution of the conflict rule set based on the data exchange rule, the computing device can adopt a standard priority coverage strategy. The computing device directly adopts the explicit definition of the data exchange rule for the conflict business concept as the benchmark for solving the conflict, and uses the benchmark to define the coverage of the original rule definitions of all conflicting parties.

[0136] One possible implementation involves the computing device traversing the set of conflict rules. For each conflict record in the set, the computing device first locates the corresponding global definition in the data exchange rules, and then generates a new, mandatory data validation rule. This new rule requires the data to satisfy the global definition and may include execution logic to convert the original non-compliant data into a standard format.

[0137] For example, the conflict rule set contains a record indicating a conflict between organizations A, B, and C regarding the format of the "Date" field ("DD / MM / YYYY", "MM / DD / YYYY", and "YYYY-MM-DD" respectively). The data exchange rule is uniformly defined as "YYYY-MM-DD". The computing device uses an authoritative ruling to generate a unique second data verification rule: "The date field is in YYYY-MM-DD format; otherwise, it needs to be converted according to the mapping relationship."

[0138] In some embodiments, when performing conflict resolution on a set of conflict rules based on data exchange rules, the computing device may adopt a strategy under a unified mapping. When the conflict mainly manifests as inconsistencies in value sets such as encoding and classification, the computing device uses a global standard value set as a benchmark to establish a mapping relationship between the value sets of all conflicting parties and the global value set, and generates unified verification and conversion rules based on this mapping.

[0139] One possible implementation is that the computing device creates a mapping table for each business concept with conflicting value sets. The generated second data validation rule encapsulates logic for querying this mapping table, used to uniformly map the input value of either conflicting party to a standard value during data processing.

[0140] For example, regarding the "Gender" field, the conflict set records show that organization A uses {1,2}, organization B uses {M,F}, and organization C uses {Male,Female}. The data exchange rule defines the standard value as {Male,Female}. The computing device creates a mapping table: 1->Male, 2->Female, M->Male, F->Female, Male->Male, Female->Female, and generates a second data validation rule based on this: "Receive input values ​​and output the standard value 'Male' or 'Female' by querying the mapping table."

[0141] In some embodiments, when resolving conflicts based on data exchange rules for a set of conflict rules, the computing device may also employ a recommendation strategy based on historical resolution records. The computing device maintains a historical resolution decision database. When processing a new conflict, the computing device retrieves similar resolved cases in the database that are similar in conflict pattern and business scenario, and presents their resolution solutions as high-priority recommendations to a human operator. After confirmation, these solutions are applied to the current conflict.

[0142] In some embodiments, in the conflict resolution of the conflict rule set based on the data exchange rule, the computing device can further perform a consistency check of the resolved rule. After generating the second data check rule, the computing device performs a logical consistency check with the existing first data check rule that is not in conflict, ensuring that the newly introduced rule does not cause new conflicts with the existing rule system, thereby maintaining the cohesion and conflict-free nature of the entire check rule set.

[0143] S204, based on the first data check rule and the second data check rule, generating a target data check rule.

[0144] In some embodiments, in the conflict resolution of the conflict rule set based on the data exchange rule, the computing device can further perform a consistency check of the resolved rule. After generating the second data check rule, the computing device performs a logical consistency check with the existing first data check rule that is not in conflict, ensuring that the newly introduced rule does not cause new conflicts with the existing rule system, thereby maintaining the cohesion and conflict-free nature of the entire check rule set.

[0145] In some embodiments, in the conflict resolution of the conflict rule set based on the data exchange rule, the computing device can further perform a consistency check of the resolved rule. After generating the second data check rule, the computing device performs a logical consistency check with the existing first data check rule that is not in conflict, ensuring that the newly introduced rule does not cause new conflicts with the existing rule system, thereby maintaining the cohesion and conflict-free nature of the entire check rule set.

[0146] For example, the computing device can establish an index structure with "table name. field name" as the key. When processing the target "accumulation fund account table. account status", the computing device can quickly locate all the check rules that act on this field by traversing the index, for example, a first data check rule requires that its "value cannot be empty", and another second data check rule requires that its "value needs to belong to the predefined enumeration set {normal, sealed, customer}".

[0147] In some embodiments, in the conflict resolution of the conflict rule set based on the data exchange rule, the computing device can further perform a consistency check of the resolved rule. After generating the second data check rule, the computing device performs a logical consistency check with the existing first data check rule that is not in conflict, ensuring that the newly introduced rule does not cause new conflicts with the existing rule system, thereby maintaining the cohesion and conflict-free nature of the entire check rule set.

[0148] In one possible implementation, the computing device builds a directed graph to formalize the dependency relationship by analyzing the logical premises among the rules. The nodes represent individual validation rules, and the edges represent the dependency relationship between the rules. The computing device processes the graph using a topological sorting algorithm to generate a linear rule execution sequence. This sequence ensures that when a data record is validated, the validation of all the preconditions is completed before the validation of the subsequent rules that depend on them.

[0149] It should be understood that this implementation can systematically accommodate existing data characteristics and resolve rule conflicts through the two-stage rule generation mechanism of "commonality extraction" and "conflict resolution". This not only improves the comprehensiveness and efficiency of the generation of validation rules, but also ensures that the generated rules have both broad inclusiveness and strict consistency, providing accurate and reliable basis for subsequent data verification and ensuring the feasibility of data integration from the source.

[0150] In some embodiments, the computing device can generate target data validation rules using a rule optimization method driven by historical business data, as shown in Figure 4 S204 specifically includes the following steps S301-S304: S301, based on the first data validation rule and the second data validation rule, a third data validation rule is generated.

[0151] In some embodiments, in the process of generating a third data validation rule based on the first data validation rule and the second data validation rule, the computing device can use a rule semantic fusion and logical expression synthesis method. Instead of simply physically stacking the two types of rules, the computing device performs a deep semantic analysis on them, and combines the constraint conditions that are logically combinable for the same data object into a more concise and efficient composite rule.

[0152] For example, when processing the "monthly deposit amount" field, the computing device finds that the first data validation rule contains constraint C1: "value is numeric", and the second data validation rule contains constraint C2: "value is greater than 0". The computing device generates a third data validation rule through syntax analysis and merging, and the logical expression of the third data validation rule is: "monthly deposit amount IS NUMERIC AND monthly deposit amount>0".

[0153] In some embodiments, in the process of generating a third data validation rule based on the first data validation rule and the second data validation rule, the computing device can use a final consistency check method for rule conflicts. Although the second data validation rule is derived from conflict resolution, there may still be subtle logical inconsistencies that are not covered by the previous process when merging with the first data validation rule. The computing device performs a final global consistency check at this step to ensure that there are no contradictions within the third data validation rule set.

[0154] In one possible implementation, the computing device employs a satisfiability modulo theories solver (SMT solver) to perform formal verification on the integrated rule set. The computing device converts all the third data checking rules into assertions in the SMT-LIB (SMT-LIB) standard, and requests the SMT solver to check the satisfiability of the assertions. If the solver returns "unsatisfiable", the computing device locates the specific rule pair that causes the conflict, and triggers a repair process.

[0155] For example, after integration, one rule requires the value of the "account status" field to be in the set {1, 2, 3}, while another associated rule requires a specific operation to be performed when the "account status" is 4. The SMT solver will identify the contradiction, and the computing device will then automatically disable or modify the latter according to the preset priority (for example, the second data checking rule is prioritized), thereby ensuring the internal consistency of the third data checking rule set.

[0156] S302, acquire historical business data of different institutions as test data.

[0157] Specifically, the historical business data can include business data records covering different business types, different time periods, and different data quality states. Different time periods should include recent data and historical data to reflect changes in business rules; different data quality states should cover fully compliant data, known problem data, and boundary case data.

[0158] In one possible implementation, the computing device uses a hierarchical sampling method to acquire test data. The computing device first identifies key classification dimensions in the business data of each institution, and then samples according to the distribution proportion of these dimensions to ensure that the test data can comprehensively cover different business scenarios and data characteristics.

[0159] For example, the computing device extracts historical contribution records for the past three years from the contribution management system as test data according to dimensions such as contribution status, unit type, and regional distribution, which include data records in different states such as normal contribution, freezing, and cancellation.

[0160] S303, verify the test data based on the third data checking rules, and acquire rule verification feedback information.

[0161] The rule verification feedback information includes rule misjudgment records that occur when the test data is verified.

[0162] In some embodiments, the rule misjudgment records can include false positive records and false negative records.

[0163] Wherein, the false positive refers to the case that the test data is actually correct but is wrongly judged as unqualified by the third data checking rule. The false negative refers to the case that the test data actually has quality problems but is wrongly judged as qualified by the third data checking rule.

[0164] In some embodiments, the computing device can obtain the rule verification feedback information by constructing a double verification system combining automated checking with manual review. The computing device first performs automated checking on all test data using the third data checking rule to generate preliminary checking results; then extracts part of the data from the checking results based on a preset strategy and submits them to manual review; finally, generates a structured difference report as the rule verification feedback information by systematically comparing the automated checking results and the manual review results.

[0165] In one possible implementation, the computing device constructs the double verification system by integrating a rule engine and a workflow engine, and realizes the automated generation of feedback information. The rule engine is responsible for performing automated checking and outputting a checking result set with metadata (such as confidence). The workflow engine drives the manual review process and internally has a result comparator. The comparator receives the automated checking results and the manual review results as input, performs record-level matching and state comparison, automatically identifies records with inconsistent states (i.e. false positives and false negatives), and packages the rule verification feedback information file containing all difference details, statistical data and associated context.

[0166] For example, the computing device sends 100,000 test data into the rule engine for automated checking. The workflow engine then starts, and automatically extracts 1,500 records to generate review tasks according to the confidence threshold. After the social security experts complete the review, the internally built result comparator automatically runs and compares the two result sets one by one. The comparator finds that among the records judged as “not passed” by the rule engine, 50 are marked as “actually qualified” by the manual review (false positives); among the records judged as “passed” by the rule engine, 15 are marked as “actually having problems” by the manual review (false negatives). The comparator then automatically generates a JSON-formatted rule verification feedback information report.

[0167] In other embodiments, the computing device can obtain the rule verification feedback information by using an automated verification method based on a test case library. The computing device maintains a test case library containing known correct data and known incorrect data, and directly generates the rule verification feedback information by automatically comparing the verification results of the test cases by the third data checking rule with the expected results.

[0168] In one possible implementation, the computing device establishes a standardized test case management system. The system includes a test case storage module, an automated execution module, and a result analysis module. The test case storage module stores test data with explicit quality labels (correct / incorrect) and error type annotations; the automated execution module invokes a rule engine to perform batch verification of test cases; and the result analysis module compares rule verification results with preset labels, automatically counts misjudgment cases, and generates performance indicator reports.

[0169] For example, the computing device selects 2000 pieces of data labeled as "correct" and 1000 pieces of data labeled as "specific type of error" from a test case library, and performs batch verification through a rule engine. The result analysis module finds that 30 of the 2000 correct data are misjudged as incorrect by the rule, and 25 of the 1000 incorrect data are not identified by the rule. The system automatically generates rule verification feedback information, including detailed distribution of various types of errors, rule recall rate, and accuracy rate, and other key indicators.

[0170] In S304, the verification threshold or logical condition in the third data verification rule is optimized based on the rule verification feedback information, to obtain a target data verification rule.

[0171] The verification threshold refers to a critical numerical value in the rule for determining whether the data is qualified. The logical condition refers to a composite expression formed by connecting multiple judgment conditions with logical operators in the rule. Optimization refers to adjusting the threshold or logical condition based on the statistical analysis results of false positives and false negatives, aiming to achieve the best balance between false positive rate and false negative rate.

[0172] In some embodiments, in the execution of the optimization of the third data verification rule based on the rule verification feedback information, the computing device can use a threshold optimization method based on false positive and false negative distribution. The computing device focuses on analyzing the feedback in the rule verification feedback information for numerical field verification rules, and finds the optimal threshold point that can minimize the overall error rate by analyzing the distribution characteristics of false positive samples and false negative samples on the numerical field.

[0173] In one possible implementation, the computing device constructs a loss function with the current threshold as a variable, which takes into account the cost of false positives and false negatives. The computing device then uses a numerical optimization algorithm, such as the gradient descent method, to find the threshold parameter that minimizes the value of the loss function through iterative calculation, and updates the original rule definition using this optimized parameter.

[0174] For example, for the rule of "monthly deposit amount > 0", the feedback information shows that some unusually high values are missed due to no upper limit, but setting a fixed upper limit generates false positives. The computing device sets the loss function as (false positive number * W1 + missed number * W2) by analyzing the distribution, and obtains the optimal upper limit threshold of 5800 yuan by optimization algorithm, so as to optimize the original rule to "monthly deposit amount > 0 AND monthly deposit amount <= 5800". In some embodiments, the computing device can use a decision tree-based conditional logic optimization method. When the rule verification feedback information shows that there is a systematic defect in the complex logic condition, the computing device trains a decision tree model using false positive and missed samples, and converts the learned decision path into a new logical condition expression to reconstruct the original rule and generate a new data validation rule.

[0175] In some embodiments, in the execution of the rule verification feedback information based on the optimization of the third data validation rule, the computing device can use a decision tree model based conditional logic optimization method. When the rule verification feedback information shows that the rule logic based on simple enumeration or single condition has high false positives or high misses in processing complex business scenarios, the computing device trains a classification model using the sample data in the feedback, and converts the learned decision boundary into a new and more accurate logic condition.

[0176] In one possible implementation, the computing device uses the decision tree algorithm to train all the problem samples (including false positives and misses) and their correct labels identified in the rule verification feedback information as a training set. By pruning the generated decision tree to prevent overfitting, the decision path is finally translated into a structured logic in the form of "IF-THEN-ELSE" to replace the ineffective logic module in the original third data validation rule.

[0177] For example, for the "unit type" validation rule, the original rule uses a fixed enumeration list (such as {"enterprise", "government", "enterprise"}), resulting in a large number of new units (such as "non-enterprise units") being missed. The computing device uses a decision tree to learn a combination judgment logic based on "organization code prefix", "registered capital range" and "industry classification code" from sample data, and then generates a new rule that is more robust and covers a wider range. In some embodiments, the computing device can also perform multi-rule collaborative optimization. The computing device establishes a rule influence propagation model, analyzes the impact of single rule adjustment on other associated rules, and finds the globally optimal rule parameter combination through a multi-objective optimization algorithm.

[0178] In some embodiments, after the third data verification rule is optimized based on the feedback information of the rule verification, the computing device can further perform regression verification of the optimized rule. The computing device uses the sample data subset in S302 or a new verification data set to re-verify the optimized data verification rule, to ensure that the optimization operation indeed improves the rule performance (such as reducing the false positive rate and the false negative rate) and does not introduce new logical conflicts, thereby ensuring the quality reliability of the target data verification rule.

[0179] It should be understood that this implementation introduces a rule verification and optimization closed loop based on historical data, which tests and calibrates the initially generated rule using real business data samples. By analyzing false positives and false negatives, the discriminant threshold and logical conditions of the rule can be optimized, thereby effectively improving the practicality of the data verification rule, reducing false positives and false negatives of compliant data in actual application, and further enhancing the automation level of the data sharing process.

[0180] In an exemplary embodiment, Figure 5 A flowchart of a method for generating a (target) data verification rule is provided for the embodiments of the present application, which specifically includes the following content: Requirement stage: The computing device obtains the data structure rules of each institution and the preset data exchange rules, and completes the preliminary rule definition. In some embodiments, the computing device can also obtain the rule opinions and requirements submitted by each institution, analyze and fuse with the above rules, and complete rule input.

[0181] Design stage: Based on the various rules obtained in the requirement stage, the computing device performs rule design and generates preliminary third data verification rules.

[0182] Development stage: The computing device develops a rule verification tool (such as the verifier in S103 above) based on the above third data verification rules. This stage realizes the specific function construction of the rule in the tool.

[0183] Test stage: The rule verification tool developed is used to verify the historical business data of each institution, and the verification feedback information is obtained. Based on the feedback information, the third data verification rule is iteratively optimized to obtain the final data verification rule, and the rule verification tool is modified synchronously. In some embodiments, the computing device can also use the rule verification tool to verify the preset data exchange rules with each other, to ensure that the constraints in the data verification rule meet the data exchange rules. In this case, the computing device needs to submit the rule verification tool to the publishing institution of the data exchange rules, i.e., the supervision platform.

[0184] Production stage: the rule verification tool that has completed the test is deployed to the production environment (i.e., the computing device) through the rule import part for rule use. In this stage, the computing device also needs to periodically evaluate the actual verification data generated by the tool running, regularly evaluate the effectiveness of the rules, and continuously optimize the data checking rules according to the evaluation results. The objects of the periodic evaluation include the business data adaptability of the institution and the compliance of the data exchange rules formulated by the regulatory platform.

[0185] In some embodiments, for the institutions whose verification results are not passed in S103, the computing device can also send data structure update information to the institutions according to the data checking rules, to instruct the institutions to update the data structure of the business data to comply with the data standard. The process is as shown in FIG. 4, before S104, the method can further include the following steps S401-S403: Figure 6 S401, for at least one to-be-updated institution whose data structure verification result of the business data is not passed from the different institutions, determine the data structure update information of each to-be-updated institution based on the data structure verification result of each to-be-updated institution and the target data checking rule.

[0186] The process of verifying the institution can refer to S103 above, which will not be described in detail here.

[0187] The to-be-updated institution can include an institution whose full-amount verification is not passed, or an institution whose verification pass rate does not reach a candidate threshold, or an institution that fails to be promoted in a candidate institution promotion mechanism.

[0188] The data structure update information refers to a structured adjustment scheme that needs to be performed to convert the data of the to-be-updated institution to comply with the sharing standard, including but not limited to field mapping relationship, format conversion rule, code value conversion table, and adjustment suggestion of data level structure.

[0189] In some embodiments, in the step of determining the data structure update information of each to-be-updated institution based on the verification result of each to-be-updated institution and the data checking rule, the computing device can use a difference comparison method. The computing device compares the actual data structure of the to-be-updated institution with the standard structure required by the data checking rule field by field, identifies the differences between the two in field definition, data type, constraint condition, etc., and generates an update scheme based on the differences.

[0190] ​In one possible implementation, the computing device constructs a structural difference analysis engine. The engine includes three core modules: a structure parser, a difference detector, and a solution generator. The structure parser is responsible for parsing the metadata information of the institution to be updated, the difference detector compares it with the standard data model, and the solution generator automatically generates the corresponding data structure update script and configuration document according to the difference type.

[0191] For example, for the fund center of F City to be updated determined in S401, the computing device finds that the main problems are concentrated in the inconsistent date format and the mismatched code value by analyzing the verification results. Through difference comparison, the computing device finds that the institution uses the "DD / MM / YYYY" date format, while the standard requires "YYYY-MM-DD"; at the same time, its deposit status uses letter coding instead of standard numerical coding. Based on this, the data structure update information generated by the computing device includes: date format conversion function, letter-numeric code mapping table, and the corresponding data conversion SQL script.

[0192] In some embodiments, the data structure update information can also include business impact assessment information generated by the computing device for the institution to update the business data structure. The business impact assessment information refers to the quantitative analysis and prediction of the impact on business continuity, system performance, and upstream and downstream dependencies that may be caused by executing the data structure update.

[0193] In one possible implementation, the computing device establishes a multi-dimensional impact assessment model. The input of the model is the list of data fields required to be updated by the institution to be updated, the update type (such as format conversion, code mapping) of each field, and the total number of records of the data to be updated. Then, the model calculates the data migration amount based on the total number of data records and the field update complexity, simulates the conversion execution time based on the migration amount and the rule complexity to assess the system downtime requirement, identifies the affected business function modules by analyzing the field business characteristics, and assesses the upstream impact by detecting the data interaction relationship. Finally, the output includes the data migration amount, the estimated downtime, the list of affected business functions, and the business impact assessment information including the analysis results of dependent systems.

[0194] S402, send each corresponding data structure update information to each institution to be updated.

[0195] In some embodiments, the process of sending data structure update information to the institution to be updated by the computing device can be divided into two ways: active push or passive provision, to adapt to the technical receiving capacity and processing habits of different institutions.

[0196] One possible implementation involves the computing device employing a proactive push strategy. Based on preset notification rules, the computing device automatically sends the updated data to the receiving address designated by each institution by either calling the message receiving interface registered by the institution or by encapsulating the updated data structure information into standardized data packets. During the transmission process, the computing device digitally signs and encrypts the updated information to ensure its integrity, authenticity, and confidentiality.

[0197] For example, the computing device proactively pushes its data structure update information package to the message middleware of the F City Housing Provident Fund Center through an established government data exchange channel. This information package is encapsulated in JSON format and contains core content such as field mapping rules, format conversion scripts, and code-value lookup tables. After the push is completed, the computing device confirms the receipt status through a callback interface and records the push log.

[0198] Another possible implementation involves a passive provision strategy for computing devices. These devices provide standardized web download interfaces or file service interfaces, allowing technical personnel from various organizations to access them on demand. After authentication, personnel can query, browse, and download data structure update information documents and related technical materials specifically generated for their organization.

[0199] In some embodiments, after sending data structure update information to the organization to be updated, the computing device can also receive feedback information from the organization. The computing device provides a feedback receiving channel for the organization to submit confirmation receipts, technical consultations, or adjustment suggestions regarding the update information, thereby forming a two-way communication loop and providing a basis for possible subsequent scheme optimization.

[0200] S403. In response to the received business data of the organization to be updated, perform structural verification on the business data of the organization to be updated. If the structural verification result is passed, determine the organization to be updated as the target organization.

[0201] Specifically, the verification process for organizations awaiting updates can be found in section S103 above, which will not be elaborated upon here.

[0202] It should be understood that this implementation constructs a dynamic and scalable closed-loop data access mechanism. It not only passively filters compliant data sources but also proactively provides precise guidance on updating data structures to organizations that have not yet passed verification, and re-verifies their subsequent data submissions. This mechanism effectively promotes the standardization of data standards among participating organizations, gradually expands the scale of qualified data sources, and thus achieves the self-improvement and sustainable development of the data sharing ecosystem.

[0203] The following section, in conjunction with the accompanying diagrams, provides a detailed explanation of the process of conducting business situation analysis based on shared business data.

[0204] In some embodiments, the computing device can perform cross-period multi-dimension business trend analysis, by obtaining state snapshot data of shared business data at different time periods, analyzing the change trend and correlation of each business dimension, and generating an analysis report for business decision-making, as shown in Figure 7 The method can further include the following steps S501-S503: S501, obtaining state snapshot data of shared business data at different time periods.

[0205] The state snapshot data includes feature parameters obtained by aggregating the shared business data according to the preset dimensions.

[0206] The feature parameters are quantitative values extracted from the business data that can reflect the status of a specific business dimension. The state snapshot data is a set of feature parameters obtained by multi-dimension quantitative statistics of shared business data at the end of the preset period.

[0207] Specifically, the preset dimensions serve to establish a structured analysis framework for subsequent trend comparison, correlation analysis, and root cause positioning. By multi-dimension aggregation, the raw data is converted into observation perspectives with clear business meaning, so that the data changes can not only be quantified, but also be positioned and attributed.

[0208] For example, the preset period refers to day, week, month, or quarter, etc. The specific value of the preset period is not limited in the embodiments of the present application.

[0209] In some embodiments, the preset dimensions can include: organization level dimensions, time period dimensions, business type dimensions, etc. Specifically, the organization level dimensions are used for data aggregation from the management range perspective, which aims to support subsequent horizontal comparison and regional difference analysis to identify performance differences and development imbalance problems between different organizations. The time period dimensions are used for data aggregation from the time series change perspective, which aims to support subsequent trend fitting and periodicity analysis to judge the business development trend and predict the future trend. The business type dimensions are used for data aggregation from the business composition perspective, which aims to support subsequent structure change and correlation analysis to understand the contribution and mutual influence of different business types.

[0210] Correspondingly, the feature parameters can include the following categories: Scale indicators: used to measure the business volume, including the number of service subjects, the number of business transactions, the total amount of fund circulation, etc.; quality indicators: used to evaluate the quality of business data, including data accuracy, information completeness, and specification compliance rate, etc.; efficiency indicators: used to reflect the business processing efficiency, including business handling time, resource utilization rate, and exception handling period, etc.; growth indicators: used to reflect the business development changes, including the year-on-year growth rate, the same period growth rate, and the market share change, etc.

[0211] For example, for the provident fund service scenario, the characteristic parameters can include: the number of depositing units (scale type index), the total monthly deposit amount (scale type index), the account information accuracy rate (quality type index), the efficiency of handling the transfer and connection of out-of-area (efficiency type index), the annual deposit amount year-on-year growth rate (growth type index), and the like.

[0212] For example, for the medical insurance settlement service scenario, the characteristic parameters can include: the number of designated medical institutions (scale type index), the total monthly settlement amount (scale type index), the diagnosis information completeness rate (quality type index), the hospitalization expense settlement period (efficiency type index), the number of insured persons year-on-year growth rate (growth type index), and the like.

[0213] In some embodiments, the computing device can obtain the state snapshot data of the shared business data at different time points through a timed batch processing aggregation method.

[0214] In one possible implementation, the computing device implements batch processing aggregation by establishing an association between a timed trigger mechanism and the distributed computing framework. The computing device takes a preset time point as a trigger condition and automatically starts a parallel computing job when the condition is met. The job first obtains the shared business data at the current time point, then divides the data into multiple independent computing tasks according to the preset dimension definition, completes the calculation of all characteristic parameters by executing these tasks, and finally automatically assembles the calculation results to form a state snapshot.

[0215] For example, in the provident fund cross-institution business collaboration scenario, in order to analyze the daily development trend of the deposit business in a specific region, the computing device needs to generate a daily-end business state snapshot. The computing device configures to trigger a parallel computing job at UTC time 00:00:00 every natural day. The job first obtains all provident fund deposit business data from the persistent storage up to 24:00:00 on the current day, and then automatically divides the data into multiple computing tasks (such as “A province-normal deposit-day”, “B province-supplementary deposit-day”, etc.) according to the combination of the institution level dimension (province, city), the business type dimension (normal deposit, supplementary deposit), and the time period dimension (day). These tasks are distributed to the computing cluster for parallel execution, and the total number of depositing units (scale type index), the cumulative deposit amount (scale type index), and the per capita deposit amount (efficiency type index) are calculated, respectively. After all the tasks are completed, the computing device merges the output results of each task according to the dimensions, assembles a complete state snapshot with “2024-12-31 24:00:00” as the timestamp, and stores it in the database. In this way, the computing device can automatically generate a daily business data state snapshot at the end of each natural day, providing accurate time series data basis for subsequent business trend analysis.

[0216] In other embodiments, the computing device can also obtain state snapshot data of shared business data at different times through a streaming incremental computing method.

[0217] Streaming incremental computation refers to a computing device that does not recalculate all metrics at every moment. Instead, it continuously monitors changes in business data and only calculates incremental data added since the last snapshot. This incremental result is then merged with the complete snapshot state of the previous moment to efficiently generate a new snapshot for the current moment. This method significantly reduces the amount of data processed per computation cycle.

[0218] One possible implementation involves deploying a state maintenance and incremental merging engine on the computing device. This engine comprises two core components: an incremental calculator that processes new data in the real-time data stream and calculates the change in metrics (i.e., the increment) since the last snapshot; and a state merger that internally maintains the complete snapshot state from the previous moment. When it receives the incremental calculation result, it merges it with the existing state to generate and output a new complete snapshot, while simultaneously updating the state it maintains.

[0219] For example, in a medical insurance settlement status analysis scenario, the computing device needs to generate a daily business status snapshot. Suppose that the snapshot generated by the computing device on December 30, 2024, records a cumulative settlement amount of 1 billion yuan. Throughout December 31, the incremental calculator continuously receives settlement transactions, calculating a new settlement amount of 50 million yuan for that day. At 24:00 on that day, the status merger merges the new 50 million yuan with the baseline status of 1 billion yuan from December 30, generating a new snapshot for December 31 (cumulative settlement amount of 1.05 billion yuan), and simultaneously updates the baseline status to 1.05 billion yuan. Through this incremental calculation method, the computing device does not need to reprocess all historical data at the end of each day; it only needs to calculate the incremental data for the current day to quickly generate a new snapshot.

[0220] S502. Perform trend fitting and comparative analysis on the state snapshot data to determine the evolution pattern and correlation of the corresponding feature parameters in different dimensions.

[0221] Trend fitting refers to using mathematical models to approximate time-series data of characteristic parameters in order to reveal the patterns of their changes over time.

[0222] Comparative analysis refers to cross-comparing characteristic parameter data from different dimensions or time intervals to identify differences and consistency.

[0223] Evolutionary pattern refers to the regular changes in characteristic parameters over time, used to describe the development pattern of a single indicator.

[0224] The correlation relationship refers to mutual influence or common change between different characteristic parameters in statistics, and is used to reveal the interaction mechanism between multiple indexes.

[0225] In some embodiments, the evolution mode can include a growth mode, a fluctuation mode, a periodic mode, etc.

[0226] The growth mode represents that the characteristic parameter value presents a continuous upward trend over time, such as stable growth of the total monthly contribution of the provident fund. The fluctuation mode represents that the characteristic parameter value fluctuates up and down within a certain range without a significant monotonic trend, such as daily medical insurance settlement amount fluctuating around the mean. The periodic mode represents that the characteristic parameter value presents regular repeated changes at fixed time intervals, such as periodic peak of the social security insured number at the end of the quarter. Embodiments of the present application do not limit the specific mode splitting of the evolution mode.

[0227] In some embodiments, the correlation relationship can include a positive correlation relationship, a negative correlation relationship, a nonlinear relationship, etc.

[0228] The positive correlation relationship represents that the growth of one characteristic parameter value is accompanied by the synchronous growth of another characteristic parameter value, such as the positive correlation between the contribution base and the contribution amount. The negative correlation relationship represents that the growth of one characteristic parameter value is accompanied by the decline of another characteristic parameter value, such as the negative correlation between the medical insurance reimbursement ratio and the personal self-payment amount. The nonlinear relationship represents that there is a non-monotonic complex correlation between the characteristic parameters, such as the first increase and then decrease relationship between the age and the medical insurance use frequency.

[0229] In some embodiments, the method of the computing device for trend fitting can include linear regression analysis, exponential smoothing processing, autoregressive integrated moving average model modeling, etc. on time series data.

[0230] In some embodiments, the method of the computing device for comparative analysis can include same period comparison (such as this year and last year same period comparison), cycle period comparison (such as this month and last month comparison), dimension comparison (such as comparison between different regions), etc.

[0231] In one possible implementation, the computing device adopts a multi-stage trend analysis method to perform trend fitting to determine the evolution mode of the characteristic parameters corresponding to different dimensions. The process specifically includes: first, pre-processing the time series data of the characteristic parameters, including missing value filling and outlier processing; then calculating the local trend characteristics using the sliding window method, including the mean, slope and volatility within the window; then using the exponential smoothing method to aggregate the trend characteristics, and distinguishing the long-term trend and short-term fluctuation; finally, dividing the sequence into a growth mode, a fluctuation mode or a periodic mode based on the slope threshold and the fluctuation range through a trend classifier, and outputting the strength index and the duration parameter of each mode.

[0232] Specifically, the sliding window method refers to dividing the time series data into multiple continuous subsequences of fixed length, and calculating the statistical features independently in each window. The window length can be configured according to the business cycle characteristics, such as setting the window size according to the 7-day, 30-day, etc. period. The trend classifier refers to a classification device based on a preset threshold rule or a machine learning model, which automatically classifies the time series into corresponding evolution modes by analyzing the numerical range of the trend features. The classifier can configure different decision thresholds to adapt to the sensitivity requirements of different business scenarios.

[0233] In another possible implementation, the computing device uses a multi-level correlation analysis method to perform comparative analysis and determine the correlation relationship of the feature parameters corresponding to different dimensions. The process specifically includes: first, standardizing the feature parameter data to eliminate the dimension effect; then calculating the covariance matrix and correlation coefficient matrix between the indicators to identify linear correlation characteristics; then using principal component analysis for dimension reduction to extract the main correlation dimensions; finally, through clustering analysis, the correlation relationship is divided into positive correlation, negative correlation and nonlinear relationship, etc., and the correlation strength index and significance level are output, and the correlation relationship network atlas is used to visualize the complex correlation structure between indicators.

[0234] Specifically, principal component analysis refers to a statistical method of converting a group of variables that may have correlations into a group of linearly uncorrelated variables through orthogonal transformation, which is used here to extract the most representative correlation dimensions from multi-dimensional feature parameters. The correlation relationship network atlas refers to a visualization model that displays the correlation relationship between feature parameters in the form of a graph structure, where nodes represent feature parameters, edges represent correlation relationships, and edge properties and directions represent correlation strength and type.

[0235] S503, generating a business situation analysis report based on the evolution mode and correlation relationship of the feature parameters corresponding to different dimensions.

[0236] The business situation analysis report is used to support business decision-making.

[0237] The business situation analysis report refers to a structured document formed after analyzing and interpreting the evolution mode and correlation relationship of the feature parameters, and its core function is to convert data analysis results into knowledge that can guide business actions, providing decision support for management personnel.

[0238] In some embodiments, the computing device can generate a business situation analysis report through a report template automatic filling method. The computing device preloads a report template containing chapter structure, chart type and text description framework, and automatically selects corresponding expression sentences, visual charts and data indicators to fill into the template according to the evolution mode and correlation relationship analysis results output in step S502, forming a complete business situation analysis report.

[0239] In one possible implementation, the computing device employs an intelligent template engine based on semantic analysis. The engine implements report generation through the following steps: first, semantic parsing of the analysis results output by S502, extracting key analysis elements including evolution mode type, correlation strength, characteristic parameter values, etc.; then selecting the most suitable report template from the preset template library through a template matching algorithm; then using natural language generation technology to convert the analysis elements into description paragraphs that conform to the business context; finally, automatically generating a comprehensive report containing trend charts, statistical tables, and textual analysis.

[0240] In other embodiments, the computing device can also enhance the utility of the business situation analysis report through key indicator early warning. Based on the evolution mode and correlation identified by S502, the computing device automatically identifies key indicators that need to be focused on and their change trends, and highlights abnormal situations and potential risks of these indicators in the report.

[0241] In one possible implementation, the computing device establishes a threshold rule library and a warning trigger mechanism. This implementation specifically includes: first, presetting the reasonable fluctuation range and threshold standard of various characteristic parameters; then comparing the difference between the current indicator value and the threshold range in real time; when detecting indicator abnormalities or trend deviations, automatically generating prominent early warning prompts and improvement suggestions in the report.

[0242] For example, when the computing device discovers through trend analysis that the medical insurance fund expenditure growth rate in a certain region has exceeded the income growth rate for three consecutive months, and correlation analysis shows that this phenomenon is strongly positively correlated with the population aging indicator, the computing device will automatically add a fund risk warning section to the business situation analysis report, clearly stating: "The current fund expenditure growth rate has exceeded the income growth rate by 15%, and it is expected that fund pressure will occur within 6 months, suggesting that medical insurance policies be adjusted in a timely manner or financial subsidies be increased", and providing specific control suggestions and measures for decision-makers.

[0243] It should be understood that this implementation deeply mines the potential value of shared business data, transforming it from a simple information carrier into a strategic asset supporting decision-making. Through trend analysis and correlation mining of historical state snapshots, the inherent laws and potential problems of business development can be revealed, generating forward-looking and instructive business situation analysis reports, ultimately realizing data-driven decision-making, and improving the intelligent management and service level of the entire public service system.

[0244] As Figure 8 A structural schematic diagram of a data processing apparatus provided by an embodiment of the present application is shown in FIG. 1. As Figure 8 shown, the data processing apparatus includes an acquisition module 801 and a processing module 802.

[0245] The acquisition module 801 is configured to acquire data structure rules of different institutions and business data. The data structure rules are used to define data organization forms of corresponding institutions. The business data is business data in the public service field.

[0246] The processing module 802 is configured to generate target data check rules based on the data structure rules and preset data exchange rules. The data exchange rules are used to define data structure standards for data interaction between institutions. The business data of the target institutions is filtered out from the different institutions based on the data structure verification results. The business data of the target institutions is fused based on the preset data quality control strategy to obtain shared business data. The shared business data is used to support data access between institutions or external systems.

[0247] In some other embodiments, the processing module 802 is specifically configured to analyze features of each data structure rule, extract common features between the rules, and generate first data check rules based on the common features. Each data structure rule is cross-compared to identify conflict rule items in each data structure rule, and a conflict rule set is generated. The conflict rule set is conflict-resolved based on the data exchange rules to generate second data check rules. The target data check rules are generated based on the first data check rules and the second data check rules.

[0248] In some other embodiments, the processing module 802 is specifically configured to generate third data check rules based on the first data check rules and the second data check rules. The acquisition module 801 is further configured to acquire historical business data of different institutions as test data. The processing module 802 is specifically configured to verify the test data based on the third data check rules, acquire rule verification feedback information, and the rule verification feedback information includes rule misjudgment records occurring when the test data is verified. The verification threshold or logic condition in the third data check rules is optimized based on the rule verification feedback information to obtain the target data check rules.

[0249] In some other embodiments, the processing module 802 is further configured to determine data structure update information of each to-be-updated institution based on the data structure verification results of each to-be-updated institution and the target data check rules, for at least one to-be-updated institution whose data structure verification result is not passed. The corresponding data structure update information of each to-be-updated institution is sent to each to-be-updated institution. In response to the received business data of the to-be-updated institution, the structure of the business data of the to-be-updated institution is verified, and the to-be-updated institution is determined as the target institution when the structure verification result is passed.

[0250] In some embodiments, the data quality control strategy includes a data integration strategy, a data security strategy, a data storage strategy, a data sharing strategy, and a data lifecycle management strategy. The data integration strategy includes a method of cleaning, transforming, and fusing business data of the target institution. The data lifecycle management strategy includes updating the lifecycle state of the shared business data based on preset rules, and the lifecycle state includes a creation state, a sharing state, an archiving state, and a destruction state.

[0251] In some embodiments, the acquisition module 801 is further configured to acquire state snapshot data of the shared business data at different time points, and the state snapshot data includes feature parameters of the shared business data aggregated according to preset dimensions. The processing module 802 is further configured to perform trend fitting and comparative analysis on the state snapshot data to determine the evolution mode and the correlation of the feature parameters corresponding to different dimensions. Based on the evolution mode and the correlation of the feature parameters corresponding to different dimensions, a business trend analysis report is generated, and the business trend analysis report is used to support business decision-making.

[0252] The data processing apparatus provided by the embodiments of the present application can perform the method shown in the method embodiments, and the implementation principles and beneficial effects can refer to the related description in the method embodiments, which will not be repeated here.

[0253] Figure 9 A structural schematic diagram of a data processing device provided by an embodiment of the present application is shown in FIG. 8. As shown in FIG. 8, the data processing device includes a memory 901, a transceiver 902, and at least one processor 903. Figure 9

[0254] The transceiver 902 is configured to interact with other devices to realize the transmission and reception of data.

[0255] For example, in the embodiments of the present application, the transceiver 902 can be specifically configured to acquire the data structure rules and the business data of each institution, or to send the respective data structure update information to each to-be-updated institution.

[0256] The memory 901 is configured to store computer program code, which includes computer instructions. The computer instructions run in the above-mentioned data processing device to implement the method shown in the above-mentioned method embodiments. For example, the memory can include a random access memory (RAM), and can also include a non-volatile memory (NVM), such as at least one disk memory, and can also be a U disk, a mobile hard disk, a read-only memory, a magnetic disk or an optical disk, etc.

[0257] ​The processor 903 can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The processor 903 can also be other general processors. The general processor can be a microprocessor or the processor can also be any conventional processor.

[0258] The memory 901, the transceiver 902 and the processor 903 are communicatively connected. For example, the memory 901, the transceiver 902 can be connected with the processor 903 through a system bus and complete mutual communication. The system bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, an industry standard architecture (ISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0259] Optionally, the memory 901 can be independent or integrated with the processor 903. When the memory 901 is independently arranged, the memory 901 and the processor 903 are connected through a system bus.

[0260] The embodiment of the present application further provides a chip for running instructions, which is used for executing the technical solution of the data processing method in the above embodiment.

[0261] The embodiment of the present application further provides a computer readable storage medium, which stores computer instructions. When the computer instructions are executed by a processor, the technical solution of the data processing method in the above embodiment is implemented. Specifically, when the computer instructions are executed by the processor, the data processing device can execute the technical solution of the data processing method in the above embodiment.

[0262] The embodiment of the present application further provides a computer program product, which comprises a computer program stored in a computer readable storage medium, at least one processor can read the computer program from the computer readable storage medium, and the at least one processor can implement the technical solution of the data processing method in the above embodiment when executing the computer program.

[0263] The computer readable storage medium can be realized by any type of volatile or nonvolatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. The computer readable storage medium can be any available medium that can be accessed by a general or special purpose computer.

[0264] An exemplary computer readable storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the computer readable storage medium. Of course, the computer readable storage medium can also be an integral part of the processor. The processor and the computer readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the computer readable storage medium can also exist as discrete components in an electronic control unit or a host device, and the embodiment of the present application does not limit this.

[0265] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative, for example, the division of the modules is only a logical function division, and actual implementation can have another division manner, for example, a plurality of modules can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed ones can be indirect coupling or communication connection through some interfaces, devices or modules, and can be electrical, mechanical or other forms.

[0266] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, i.e., may be located in one place, or may be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to implement the embodiments of the present application.

[0267] In addition, the functional modules in each embodiment of the present application can be integrated in one processing unit, or each module can be physically present alone, or two or more modules can be integrated in one unit. The unit formed by the above modules can be realized in the form of hardware or in the form of hardware plus software functional units.

[0268] The integrated modules realized in the form of software functional modules can be stored in a computer readable storage medium. The software functional modules stored in a storage medium include a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute part of the steps of the method of each embodiment of the present application.

[0269] It should be understood that the steps of the method disclosed in the embodiments of the present application can be directly embodied as hardware processor execution, or executed by a combination of hardware and software modules in the processor.

[0270] Those of ordinary skill in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction-related hardware. The aforementioned program can be stored in a computer readable storage medium. The program, when executed, executes steps including the above-mentioned method embodiments; and the aforementioned storage medium includes ROM, RAM, magnetic or optical disc, and various media that can store program codes.

[0271] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the above embodiments, or make equivalent replacements for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A data processing method, characterized in that, The method includes: Acquire the data structure rules and business data of different institutions; the data structure rules are used to define the data organization form of the corresponding institution; the business data is business data in the public service sector. Based on the data structure rules and the preset data exchange rules, target data verification rules are generated; the data exchange rules are used to define the data structure standards for data interaction between organizations. Based on the target data verification rules, the business data of the different institutions are verified for data structure, and based on the data structure verification results, the target institutions whose data structure verification results are passed are selected from the different institutions. Based on a preset data quality control strategy, the business data of the target organization is fused to obtain shared business data; the shared business data is used to support data access between organizations or external systems.

2. The method according to claim 1, characterized in that, The step of generating target data verification rules based on the data structure rules and preset data exchange rules includes: Feature analysis is performed on each of the data structure rules to extract common features among the rules, and a first data verification rule is generated based on the common features. By cross-comparing the various data structure rules, conflicting rule items in each data structure rule are identified, and a set of conflicting rules is generated. Based on the data exchange rules, the conflict rule set is resolved to generate a second data verification rule; Based on the first data verification rule and the second data verification rule, a target data verification rule is generated.

3. The method according to claim 2, characterized in that, The step of generating target data verification rules based on the first data verification rule and the second data verification rule includes: Based on the first data verification rule and the second data verification rule, a third data verification rule is generated; Obtain historical business data from different institutions as test data; The test data is verified based on the third data verification rule to obtain rule verification feedback information; the rule verification feedback information includes records of rule misjudgments that occur when verifying the test data; Based on the rule verification feedback information, the verification threshold or logical conditions in the third data verification rule are optimized to obtain the target data verification rule.

4. The method according to claim 1, characterized in that, Before performing data fusion on the target organization's business data based on a preset data quality control strategy to obtain shared business data, the method further includes: For at least one organization selected from the different organizations whose data structure verification result for business data fails, data structure update information for each organization to be updated is determined based on the data structure verification result of each organization to be updated and the target data verification rule. Send the corresponding data structure update information to each of the aforementioned organizations to be updated; In response to the received business data of the organization to be updated, a structural verification is performed on the business data of the organization to be updated. If the structural verification result is passed, the organization to be updated is identified as the target organization.

5. The method according to claim 1, characterized in that, The data quality control strategy includes: data integration strategy, data security strategy, data storage strategy, data sharing strategy, and data lifecycle management strategy; The data integration strategy includes methods for cleaning, transforming, and integrating the business data of the target organization; The data lifecycle management strategy includes: updating the lifecycle status of the shared business data based on preset rules; the lifecycle status includes: creation status, sharing status, archiving status, and destruction status.

6. The method according to claim 1, characterized in that, The method further includes: Acquire state snapshot data of shared business data at different times; the state snapshot data includes feature parameters aggregated according to preset dimensions of the shared business data. The state snapshot data is subjected to trend fitting and comparative analysis to determine the evolution pattern and correlation of the corresponding feature parameters in different dimensions; Based on the evolution patterns and correlations of the feature parameters corresponding to the different dimensions, a business situation analysis report is generated; the business situation analysis report is used to support business decision-making.

7. A data processing apparatus, characterized in that, include: The acquisition module is used to acquire the data structure rules and business data of different organizations; the data structure rules are used to define the data organization form of the corresponding organization. The business data refers to business data in the public service sector. The processing module is used to generate target data verification rules based on the data structure rules and preset data exchange rules; the data exchange rules are used to define the data structure standards for data interaction between institutions; the processing module performs data structure verification on the business data of the different institutions based on the target data verification rules, and selects the target institutions whose data structure verification results are passed from the different institutions based on the data structure verification results. Based on a preset data quality control strategy, the business data of the target organization is fused to obtain shared business data; the shared business data is used to support data access between organizations or external systems.

8. A data processing device, characterized in that, include: A memory and at least one processor; the memory is communicatively connected to the processor; the memory is used to store computer program code, the computer program code including computer instructions; when the processor executes the computer instructions, the data processing device causes the data processing device to perform the method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, are used to implement the method as described in any one of claims 1-6.

10. A computer program product, characterized in that, When the computer program product is run on a computer / executed by the computer's processor, it implements the method as described in any one of claims 1-6.