Cross-system Asset Data Reconciliation Method, System, Device and Storage Medium

Through semantic coding and large-scale model analysis, the problem of inconsistency between the data of the enterprise asset management system and the financial system is solved, the accuracy and security of data reconciliation are improved, and sensitive data leakage is avoided.

CN120181244BActive Publication Date: 2025-07-22QINGDAO PORT INT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510645148.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-07-22
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

When enterprises introduce new management systems, it is difficult for existing technology to effectively solve the synchronization problem of asset data in different systems, resulting in inconsistency of data and the sensitivity of financial data increases the risk of leakage.

Method used

The asset data is encoded using a semantic encoder, the difference data is input into the large model for analysis, and the inference text is verified through preset rules to ensure the accuracy and security of the data.

Benefits of technology

It improves the accuracy of cross-system asset data reconciliation, reduces the risk of leakage of sensitive data, and realizes the reliability of automatic data correction and verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181244B_ABST
    Figure CN120181244B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing, and specifically provides a cross-system asset data reconciliation method, system, device, and storage medium, including: respectively obtaining the asset data of the first system and the second system; respectively encoding the asset data of the first system and the asset data of the second system by using a semantic encoder to obtain the first sample data of the first system and the second sample data of the second system; screening out the difference data between the first sample data and the second sample data; inputting the difference data into a large model to obtain the inference text of the large model and the correct data corresponding to the difference data output by the large model; using a preset rule to verify the inference text, and if it is confirmed that the inference text passes the verification, it is determined that the correct data is credible. The present invention converts asset data into encoded vectors, avoiding the leakage of sensitive data; and performs comparison in the form of encoded vectors, improving the redundancy for naming deviations of different systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a cross-system asset data reconciliation method, system, device, and storage medium. Background Art

[0002] When an enterprise introduces a new management system, the asset ledger of the original system is imported into the new system during the introduction phase. For example, when introducing an enterprise asset management system, the ledger of the financial system is imported into the enterprise asset management system. Over time, the operating processes of the two systems are different, resulting in the asset information and status being out of sync, that is, the data is out of sync.

[0003] To ensure the consistency of asset data, enterprises need to invest a large amount of human resources in comparing asset information and finding the reasons. Existing data comparison methods such as the Euclidean distance calculation method and the cosine similarity calculation method can achieve the calculation and matching of data similarity. However, the expression of each record information of the same asset may be different in different systems, so the error rate of the existing similarity matching method is relatively high.

[0004] Moreover, financial data is relatively sensitive. If financial data is stored in a third-party reconciliation system for a long time for the third-party reconciliation system to call, there is a risk of data leakage. Summary of the Invention

[0005] Aiming at the above deficiencies of the prior art, the present invention provides a cross-system asset data reconciliation method, system, device, and storage medium to solve the above technical problems.

[0006] In a first aspect, the present invention provides a cross-system asset data reconciliation method, including:

[0007] Respectively obtain the asset data of the first system and the second system;

[0008] Use a semantic encoder to encode the asset data of the first system and the asset data of the second system respectively to obtain the first sample data of the first system and the second sample data of the second system;

[0009] Screen out the difference data between the first sample data and the second sample data;

[0010] Input the difference data into a large model to obtain the inference text of the large model and the correct data corresponding to the difference data output by the large model;

[0011] Use a preset rule to verify the inference text. If it is confirmed that the inference text passes the verification, then determine that the correct data is credible.

[0012] In an optional embodiment, respectively obtaining the asset data of the first system and the second system includes:

[0013] Obtain the first asset data set from the external interface of the first system;

[0014] Obtain the second asset data set from the external interface of the second system;

[0015] The first system is an asset management system, and the second system is a financial system;

[0016] The first asset data set includes the static data and dynamic data of the assets in the first system;

[0017] The second asset data set includes the static data and dynamic data of the assets in the second system;

[0018] The dynamic data in the first asset data set and the second asset data set both include the asset status and asset business generated by the assets within a period of time, and the static data both include the asset identity information.

[0019] In an alternative embodiment, encode the asset data of the first system and the asset data of the second system respectively using a semantic encoder, including:

[0020] Perform data cleaning and text splicing on the asset data to convert the identity fields, status values, and business fields in the asset data into text-format information;

[0021] Use the semantic encoder to perform staged encoding on the asset data after data cleaning and text splicing. In the first stage, generate corresponding asset identity codes according to each asset identity information of the asset data. In the second stage, generate corresponding asset status codes according to each status information of the asset data. In the third stage, generate corresponding asset business codes according to each business information of the asset data;

[0022] Divide the asset identity codes, asset status codes, and asset business codes of the same asset into the same data group;

[0023] Sort the asset status codes and asset business codes in the same data group according to the timestamps of the corresponding original data to obtain a status code sequence and a business code sequence;

[0024] Align the status code sequence and the business code sequence by filling with previous value vectors or zero vectors;

[0025] Store the asset identity codes and the corresponding status code sequences and business code sequences in a time series database.

[0026] In an alternative embodiment, screen out the difference data between the first sample data and the second sample data, including:

[0027] Obtain the first sample data and the second sample data from the time series database. Both the first sample data and the second sample data include multiple asset identity codes and corresponding multiple status code sequences and business code sequences;

[0028] Perform an intersection operation on the asset identity codes of the first sample data and the second sample data, and output the asset identity codes that do not belong to the intersection as difference data;

[0029] Deduplicate the status code sequences corresponding to the same asset identity code in the first sample data and the second sample data, and compare the consistency of the two status code sequences after deduplication. If the two are inconsistent, output these two status code sequences as difference data;

[0030] Perform a zero vector removal process on the business code sequences corresponding to the same asset identity code in the first sample data and the second sample data, and compare the consistency of the two business code sequences after processing. If the two are inconsistent, output these two business code sequences as difference data.

[0031] In an alternative embodiment, input the difference data into a large model to obtain the inference text of the large model and the correct data corresponding to the difference data output by the large model, including:

[0032] Import the difference data and the associated data of the difference data into a pre-constructed prompt template to obtain a prompt. The prompt template includes the requirements for exporting the inference text;

[0033] Input the prompt into the large model. The large model has been pre-trained, and the output layer of the large model includes a decoder;

[0034] Obtain the inference text generated by the large model and the correct data corresponding to the difference data.

[0035] In an alternative embodiment, extract key elements from the inference text using natural language processing technology. The key elements include data source, field type, and logical relationship;

[0036] Obtain the system functions and data quality indicators of the first system and the second system;

[0037] Use the rule execution algorithm to retrieve preset rules one by one from the rule library, and verify the key elements based on the system functions and data quality indicators.

[0038] In an alternative embodiment, the preset rules include:

[0039] The system core function priority principle, which stipulates that the credibility of data related to the system core function is higher than the credibility of data related to the non-core function of the system;

[0040] The principle of data update mechanism limits that the credibility of real-time updated data is higher than that of non-real-time updated data;

[0041] The principle of audit log traceability ability limits the positive proportional relationship between the integrity of audit logs and data credibility;

[0042] The principle of external evidence verification limits that the data with external evidence has higher credibility;

[0043] The principle of error occurrence frequency limits the inverse proportional relationship between the error occurrence frequency of the system and the credibility of the system's data;

[0044] The principle of data consistency range limits the verification of data credibility by using the logical consistency of data with other relevant data.

[0045] In a second aspect, the present invention provides a cross-system asset data reconciliation system, including:

[0046] An acquisition module for respectively acquiring the asset data of the first system and the second system;

[0047] An encoding module for respectively encoding the asset data of the first system and the asset data of the second system by using a semantic encoder to obtain the first sample data of the first system and the second sample data of the second system;

[0048] A first processing module for screening out the difference data between the first sample data and the second sample data;

[0049] A second processing module for inputting the difference data into a large model to obtain the inference text of the large model and the correct data corresponding to the difference data output by the large model;

[0050] A verification module for verifying the inference text by using a preset rule, and if it is confirmed that the inference text passes the verification, it is determined that the correct data is credible.

[0051] In a third aspect, there is provided a device, including:

[0052] A memory for storing a cross-system asset data reconciliation program;

[0053] A processor for implementing the steps of the cross-system asset data reconciliation method provided in the first aspect when executing the cross-system asset data reconciliation program.

[0054] In a fourth aspect, there is provided a computer-readable storage medium, on which a cross-system asset data reconciliation program is stored, and when the cross-system asset data reconciliation program is executed by a processor, the steps of the cross-system asset data reconciliation method provided in the first aspect are implemented.

[0055] The beneficial effects of the present invention are as follows. The cross-system asset data reconciliation method, system, device, and storage medium provided by the present invention encode asset data using a semantic encoder, converting the asset data into encoded vectors, thus avoiding the leakage of sensitive data. Moreover, when comparing data from different systems, the comparison is performed in the form of encoded vectors, improving the redundancy for naming deviations in different systems and thereby enhancing the reconciliation accuracy. In addition, by analyzing the differential data using a large model, automatic data error correction can be achieved, and at the same time, the reasoning of the large model is verified through preset rules to improve the accuracy of automatic data error correction.

[0056] In addition, the design principle of the present invention is reliable and the structure is simple, having a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0058] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.

[0059] Figure 2 It is a schematic flowchart of the method according to an embodiment of the present invention in an application scenario.

[0060] Figure 3 It is a schematic block diagram of the system according to an embodiment of the present invention.

[0061] Figure 4 It is a schematic structural diagram of a device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0064] The following explains the key terms that appear in the present invention.

[0065] BERT (Bidirectional Encoder Representations from Transformers), that is, bidirectional encoder representations based on transformers. BERT is based on the Transformer architecture, which is a deep learning model architecture for processing sequential data. It uses the self-attention mechanism to capture the semantic relationships between different positions in the text. The bidirectional feature of BERT is reflected in that it can consider the context information of a word simultaneously during the training process, that is, the context from left to right and from right to left, so as to learn a more comprehensive and richer semantic representation.

[0066] Chain of Thought is an important concept in the field of natural language processing in recent years, which helps to improve the reasoning ability and performance of language models. Chain of Thought refers to a series of intermediate reasoning steps that connect the input of a problem with the final answer. Simply put, when the model solves a problem, it does not directly give the answer, but gradually derives and shows the thinking process. For example, for the math problem "Xiaoming has 5 apples, Xiaohong gives him 3 more, and then he eats 2. How many apples does Xiaoming have now?" The Chain of Thought can be "First calculate the number of apples Xiaoming has after getting the apples given by Xiaohong, 5 + 3 = 8; then calculate the number of apples Xiaoming has after eating 2, 8 - 2 = 6", and the final answer is obtained through such step-by-step reasoning.

[0067] Large model, that is, Large Pre-trained Model, refers to a deep learning model pre-trained based on large-scale data, which has wide applications and important influences in multiple fields such as natural language processing and computer vision. Large models with Chain of Thought such as Deepseek, by imitating the human thinking process, using reinforcement learning and "Chain of Thought" technology, guide the model to autonomously solve problems. It will think deeply before answering questions, generate a long internal Chain of Thought, break the problem into smaller steps and solve them one by one, showing powerful capabilities beyond previous models in solving complex problems such as science, coding, and mathematics.

[0068] The cross-system asset data reconciliation method provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the cross-system asset data reconciliation system runs in the computer device.

[0069] Figure 1 It is a schematic flowchart of the method of an embodiment of the present invention. Among them, Figure 1The executing entity can be a cross-system asset data reconciliation system. According to different requirements, the order of steps in this flowchart can be changed, and some can be omitted.

[0070] As Figure 1 shown, the method includes:

[0071] S1. Obtain the asset data of the first system and the second system respectively.

[0072] For the first system and the second system, it is necessary to investigate the types of data interfaces they provide respectively. Common ones are RESTful API, SOAP API, etc. According to the interface documentation, develop corresponding data request modules. For example, if it is a RESTful API, use the GET or POST method of the HTTP protocol, carry necessary authentication information (such as API Key, OAuth token) and send a request to the specified interface endpoint to obtain the asset data.

[0073] The asset data formats returned by the two systems may be different, such as JSON, XML or CSV. Use data parsing libraries, such as Jackson (Java) for JSON, json library (Python); JAXB (Java) for XML, lxml library (Python); OpenCSV (Java), pandas library (Python) for CSV, to parse the data into an internal data structure that can be processed by the program, such as an object or a data frame.

[0074] Clean the obtained original asset data, remove duplicate records, handle missing values (strategies such as filling, deleting can be adopted), and correct incorrect data formats (such as standardizing date formats). For example, use the DataFrame API of the data processing framework Spark, and use the dropDuplicates method to remove duplicates and the na.fill method to fill missing values.

[0075] S2. Use semantic encoders to encode the asset data of the first system and the asset data of the second system respectively, to obtain the first sample data of the first system and the second sample data of the second system.

[0076] A pre-trained word vector model can be selected, such as a language model based on the Transformer architecture, such as the BERT, GPT series. Taking BERT as an example, use its pre-trained weights and fine-tune in the field of asset data.

[0077] Convert the text fields in the asset data (such as asset descriptions, asset tags) into vector representations. If using BERT, first tokenize the text, use the WordPiece tokenization algorithm to split the long text into sub - word units, and then convert these sub - words into corresponding word embedding vectors. For numerical asset data, after normalizing it to a specific range (such as [0, 1]), it can be concatenated with the text vector or processed separately.

[0078] Prepare a training dataset that contains a large number of asset data samples and their corresponding annotations (if any). Use the training data to fine - tune the selected semantic encoder and optimize the model parameters to adapt to the characteristics of the asset data. After training, input the asset data of the first system and the second system into the fine - tuned semantic encoder respectively to obtain the first sample data and the second sample data, which are represented in vector form and contain the semantic information of the asset data.

[0079] S3. Filter out the difference data between the first sample data and the second sample data.

[0080] Adopt methods such as cosine similarity and Euclidean distance to calculate the similarity between the first sample data vector and the second sample data vector. For example, use the cosine_similarity function in the Scikit - learn library to calculate the cosine similarity. Set a similarity threshold, such as 0.8, and consider the sample pairs with a similarity lower than this threshold as having differences.

[0081] For numerical features, directly compare the differences in the corresponding values, such as calculating the absolute value of the difference or the relative difference. For text features, in addition to vector similarity, text edit distance algorithms (such as Levenshtein distance) can also be used to measure the degree of text difference. The data samples with significant differences will be selected as the difference data by comprehensively considering the vector similarity and the results of feature difference analysis.

[0082] S4. Input the difference data into the large model to obtain the inference text of the large model and the correct data corresponding to the difference data output by the large model.

[0083] Select a large model with the ability of chain of thought, such as DeepSeek. Connect to the large model through the API interface and send a request containing the difference data. The request format needs to follow the regulations of the large model. For example, organize the difference data into a JSON format, including fields such as data description and data samples.

[0084] Design effective prompts to guide the large model for reasoning. The prompt should clearly describe the task, such as "Analyze the following differential asset data, give the possible reasons for the differences, and provide the correct data". At the same time, some examples can be provided to help the large model better understand the task and the requirements for the output format.

[0085] The response returned by the large model contains the reasoning text and the possible correct data. Parse the response to extract the reasoning text (such as the large model's analysis and explanation of the reasons for the differences) and the correct data part. If the response format is JSON, use the corresponding parsing library to extract the required field values.

[0086] S5. Use preset rules to verify the reasoning text. If it is confirmed that the reasoning text passes the verification, then determine that the correct data is trustworthy.

[0087] The preset rules can be formulated based on domain knowledge and business logic. For example, the reasoning text should contain reasonable analysis of the reasons for the differences, such as keywords like "data entry error", "system update causing data inconsistency", etc.; the structure of the reasoning text should be reasonable, including problem description, reason analysis, and suggestion parts. At the same time, rules are set for the format, value range, etc. of the correct data. For example, the asset amount should be greater than zero and meet the financial data precision requirements.

[0088] Use tools such as regular expressions and text matching algorithms to verify the reasoning text. For example, use regular expressions to match whether the preset keywords for the reasons for the differences exist in the reasoning text. For the correct data, verify it according to the set format and value range rules, such as using a data validation library (such as Hibernate Validator for Java) to verify the legality of the data. If both the reasoning text and the correct data pass the preset rule verification, then determine that the correct data is trustworthy.

[0089] In an embodiment of the present invention, based on step S1, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation scheme.

[0090] First, design an interface and obtain data through the interface.

[0091] Design a RESTful API interface for the asset management system (the first system) to obtain static asset data (such as asset numbers, categories, purchase dates, etc.) and dynamic data (such as maintenance records, status change logs).

[0092] The financial system (the second system) provides static asset data (such as original asset value, depreciation method) and dynamic data (such as depreciation calculation tables, expense allocation records) through a SOAP protocol interface.

[0093] Unify the asset identity identifiers of the two systems (such as asset numbers, serial numbers), establish a comparison table for static data fields, and ensure the consistency of static data.

[0094] Map dynamic data according to business types. For example, map "maintenance records" to "asset operations" and map "depreciation status" to "asset status".

[0095] Then, integrate the data.

[0096] Sort the dynamic data of the two systems by timestamp and remove duplicate records (such as redundant data caused by multiple synchronizations).

[0097] Unify the data format (for example, the date format is YYYY - MM - DD, and the currency unit is RMB yuan).

[0098] Match the asset status with business records based on a time range (such as monthly, quarterly). For example:

[0099] The "scrap application" in the asset management system needs to be associated with the corresponding "asset net value cleared" operation in the financial system.

[0100] The "depreciation completed" status in the financial system needs to be synchronized to the asset status in the asset management system as "depreciated".

[0101] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.

[0102] S201. Perform data cleaning and text splicing on the asset data to convert the identity fields, status values, and business fields in the asset data into text - formatted information.

[0103] (1) Remove duplicate records: Use a data processing framework, such as the DataFrame API of Apache Spark. Assume that the asset data is stored in a Spark DataFrame. By calling the df.dropDuplicates() method, this method will identify and delete exactly the same rows based on all columns of the DataFrame. If you only need to deduplicate based on specific columns (such as the identity field column identity_column), you can use df.dropDuplicates([identity_column]).

[0104] (2) Handle missing values:

[0105] If the identity field is a required field, a strategy of deleting records containing missing identity values can be adopted. In Spark, use df = df.filter(df[identity_column].isNotNull()). If a certain degree of incompleteness is allowed, consider filling in a special identifier such as "unknown", using df = df.na.fill("unknown", [identity_column]).

[0106] If the status value is missing, it can be filled according to the business logic. For example, if the status value is a count, it can be filled with 0; if it is a range value, it can be filled with the middle value of the range. In the pandas library of Python, for the DataFrame df storing asset data, df[status_column] = df[status_column].fillna(0) can be used (assuming filling with 0).

[0107] If the business field is missing and there is a default business value, it can be filled. For example, if the business field is the business type and the default value is "general", then df = df.na.fill("general", [business_column]).

[0108] (2)Text splicing:

[0109] ‌Identity text‌: Concatenate the identity fields into a natural language description. ‌Example‌: Asset type: Server | Department: IT Department | Geographic location: Beijing Data Center.

[0110] ‌Status text‌: Convert the status value into a descriptive text. ‌Example‌: Running status: Normal | CPU usage: 75% | Memory occupancy: 60% | Latest alert: None.

[0111] ‌Business text‌: Generate a coherent description based on the business fields. ‌Example‌: Business type: Data migration | Impact scope: Financial system | Operation record: Migration started on 2023-01-05.

[0112] S202. Use a semantic encoder to encode the asset data that has undergone data cleaning and text splicing in stages. In the first stage, generate corresponding asset identity codes according to each asset identity information of the asset data. In the second stage, generate corresponding asset status codes according to each status information of the asset data. In the third stage, generate corresponding asset business codes according to each business information of the asset data.

[0113] 1. Asset Identity Encoding

[0114] Objective: Generate a unique and static semantic encoding for each asset to identify asset attributes.

[0115] Processing Steps:

[0116] Encode the identity text of each asset using Sentence - BERT.

[0117] If the same asset appears multiple times (at different time points), it only needs to be encoded once (the identity information remains unchanged).

[0118] Code Example:

[0119] from sentence_transformers import SentenceTransformer

[0120] model = SentenceTransformer('paraphrase - mpnet - base - v2')

[0121] identity_text = "Asset type: Server | Department: IT | Location: Beijing Data Center"

[0122] identity_embedding = model.encode(identity_text)

[0123] 2. Asset Status Encoding

[0124] Objective: Generate a status semantic encoding that changes over time to reflect the real - time situation of the asset.

[0125] Processing Steps:

[0126] Sort the status text of each asset by timestamp.

[0127] Encode the status text for each time point separately.

[0128] Code Example:

[0129] status_texts =

[0130] "Running status: Normal | CPU usage: 75% | Memory occupancy: 60% | Recent alerts: None",

[0131] "Operating Status: Abnormal | CPU Usage: 95% | Memory Occupancy: 85% | Latest Alarm: CPU Overload"

[0132] status_embeddings = model.encode(status_texts).

[0133] 3. Asset Business Encoding

[0134] Objective: Generate business semantic encodings that change over time to reflect the dynamics of related businesses.

[0135] Processing Steps:

[0136] Sort the business texts of each asset by timestamp.

[0137] Encode the business texts for each time point separately.

[0138] Code Example:

[0139] business_texts =

[0140] "Business Type: Data Migration | Impact Scope: Financial System | Operation Record: Migration started on 2023-01-05",

[0141] "Business Type: System Maintenance | Impact Scope: All | Operation Record: Maintenance completed on 2023-01-10"

[0142] business_embeddings = model.encode(business_texts).

[0143] Among them, the semantic encoder uses Sentence-BERT, and its core architecture is:

[0144] Based on the pre-trained BERT model, the underlying parameters of the BERT model are shared by using the Siamese network. Among them, the Siamese network adopts the single-tower mode.

[0145] Add a pooling operation after the BERT output layer to generate a fixed-length sentence vector: CLS vector: directly use the output of the [CLS] token. MEAN vector: calculate the mean of all Token outputs. MAX vector: take the maximum value of each Token output.

[0146] The basic model parameters include:

[0147] BERT-base: 12-layer Transformer, with a hidden layer dimension of 768 and 12 attention heads.

[0148] The specific encoding principle is as follows: SBERT supports generating fixed-dimensional semantic vectors for a single sentence through an independent encoding structure (such as the single-tower mode in a siamese network), without relying on paired input or similarity calculation. Its core design includes a BERT pre-trained model + a pooling layer, directly outputting the embedding representation of the sentence.

[0149] After the input sentence passes through the BERT layer, MEAN pooling (the mean of all Token vectors) or CLS pooling (the first character vector) is used to generate a sentence vector. For example, the text field in a data table (such as "product description") is converted into a 768-dimensional vector for subsequent clustering or classification tasks; SBERT can process parameters in structured data, but the parameters need to be converted into natural language form (such as rewriting "temperature: 25°C" as "the current temperature is 25 degrees Celsius") and then input into the model for encoding.

[0150] Compared with the traditional BERT model, this semantic encoder has the following advantages:

[0151] It has independent encoding ability;

[0152] Single-sentence processing efficiency: SBERT uses a single-tower siamese network to support independent encoding of a single sentence (such as directly generating a 768-dimensional vector for the asset description "purchased a CNC machine in 2023"), while the traditional BERT relies on cross-encoding (requiring paired input of two sentences for interactive calculation), and the inference speed is increased by 3 - 5 times.

[0153] Offline vector pre-computation: Asset data can be pre-encoded and stored, and vectors can be directly called during real-time matching (such as reducing the matching time of a database of tens of thousands of assets from minutes to seconds).

[0154] The pooling strategy realizes multi-mode feature extraction. And the dynamic selection mechanism supports automatically switching the pooling method according to the data type (such as using CLS for asset identity information and MEAN for dynamic status logs).

[0155] Comparative advantages with general word vector models (such as Word2Vec, GloVe):

[0156] It has context awareness: accurately distinguishing scenarios with the same word but different meanings (such as generating different vectors for "closed circuit breaker" in EAM and "closed accounting period" in EAS); directly encoding composite fields (such as "double-declining balance depreciation method") without the need for word segmentation preprocessing, and the semantic integrity is increased by 35%.

[0157] With cross-system alignment capabilities:

[0158] Heterogeneous data mapping: Map the EAM device code "EQP-2024-001" and the EAS financial code "FA-2024-001A" to the same semantic space, with a cosine similarity > 0.934.

[0159] Noise resistance: Can still generate stable vectors for missing fields (such as the asset status record lacking a timestamp), reducing the error propagation rate to 2.1%.

[0160] S203. Divide the asset identity code, asset status code, and asset business code of the same asset into the same data group.

[0161] Assume that the asset identity code, asset status code, and asset business code are stored in the corresponding lists asset_identity_codes, asset_status_codes, and asset_business_codes, and the asset data is stored in a DataFrame df. The following code can be used to divide the codes of the same asset into the same data group:

[0162] data_groups = []

[0163] for index, row in df.iterrows():

[0164] identity_code = asset_identity_codes[index];

[0165] status_code = asset_status_codes[index];

[0166] business_code = asset_business_codes[index];

[0167] data_group = {"identity_code": identity_code, "status_code": status_code, "business_code": business_code};

[0168] data_groups.append(data_group).

[0169] S204. Sort the asset status codes and asset business codes in the same data group according to the timestamps of the corresponding original data to obtain a status code sequence and a business code sequence.

[0170] Input: identity_embedding (static) + status_embeddings (dynamic) + business_embeddings (dynamic) of the same asset.

[0171] Processing logic: Align the status codes and business codes according to the timestamps.

[0172] Assume there is a timestamp column timestamp_column in the DataFrame df of asset data. Sort the asset status codes and asset business codes in the same data group according to the timestamps:

[0173] import numpy as np;

[0174] for group in data_groups:

[0175] status_codes = group["status_code"];

[0176] business_codes = group["business_code"];

[0177] timestamps = df[timestamp_column].values;

[0178] status_sort_indices = np.argsort(timestamps)

[0179] business_sort_indices = np.argsort(timestamps);

[0180] group["status_code_sequence"] = [status_codes[i] for i in status_sort_indices];

[0181] group["business_code_sequence"] = [business_codes[i] for i in business_sort_indices].

[0182] S205. Align the status encoding sequence and the service encoding sequence by padding with the previous value vector or zero vector.

[0183] Among them, the status encoding sequence is padded with the previous value vector, and the service encoding sequence is padded with the zero vector.

[0184] Suppose the status encoding sequence is status_code_sequence, the service encoding sequence is business_code_sequence, and the lengths of the two sequences are different. Taking padding with the previous value vector as an example, if the status encoding sequence is shorter, the following code can be used for padding:

[0185] max_length = max(len(status_code_sequence), len(business_code_sequence));

[0186] padded_status_code_sequence = [];

[0187] for i in range(max_length):

[0188] if i<len(status_code_sequence):

[0189] padded_status_code_sequence.append(status_code_sequence[i]);

[0190] else:

[0191] padded_status_code_sequence.append(status_code_sequence[-1]).

[0192] Suppose the dimension of the encoding vector is vector_dim, and generate the zero vector zero_vector = np.zeros(vector_dim). For the shorter status encoding sequence:

[0193] max_length = max(len(status_code_sequence), len(business_code_sequence));

[0194] padded_status_code_sequence = [];

[0195] for i in range(max_length):

[0196] if i < len(status_code_sequence):

[0197] padded_status_code_sequence.append(status_code_sequence[i]);

[0198] else:

[0199] padded_status_code_sequence.append(zero_vector);

[0200] For example, if there is only one type of encoding at a certain time point (such as only status updates), the other item is filled with a zero vector or a previous value fill (depending on business requirements).

[0201] Example structure:

[0202] # Each time step of the time series is a triple;

[0203] time_series =

[0204] (timestamp1, identity_embedding, status_embedding1, business_embedding1),

[0205] (timestamp2, identity_embedding, status_embedding2, business_embedding2)];

[0206] S206. Store the asset identity encoding, the corresponding status code sequence, and the business code sequence in a time series database.

[0207] Use a time series database (such as InfluxDB) or a compressed format (Parquet) for storage.

[0208] Example of each line of data:

[0209] {"asset_id": "Asset-001",

[0210] "timestamp": "2023-01-01 10:00:00",

[0211] "identity_embedding": [0.12, -0.45,..., 0.78], / / 768-dimensional vector

[0212] "status_embedding": [0.34, 0.56, ..., -0.12],

[0213] "business_embedding": [-0.23, 0.67, ..., 0.89]}

[0214] ‌Index Optimization‌:

[0215] Create a composite index on asset_id and timestamp to accelerate queries by asset and time.

[0216] In one embodiment of the present invention, based on step S3, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.

[0217] S301. Obtain the first sample data and the second sample data from the time - series database. Both the first sample data and the second sample data include multiple asset identity codes and corresponding multiple status code sequences and business code sequences.

[0218] According to the query language Flux of InfluxDB, write a query statement to obtain the first sample data and the second sample data. Assume that the asset data is stored in a measurement named "asset_samples", and the first system and the second system are distinguished by the tag "system_type" (the values are "system1" and "system2" respectively). An example of the Flux statement to query the first sample data is as follows:

[0219] from(bucket: "your - bucket - name");

[0220] |>range(start: 0);

[0221] |>filter(fn: (r) =>r._measurement == "asset_samples" and r.system_type == "system1").

[0222] The query result of InfluxDB is a complex data structure. For the above query result, it is necessary to parse out the asset identity code, status code sequence, and business code sequence. Assume that the field names in the query result are "asset_identity_code", "status_code_sequence", and "business_code_sequence" respectively. In Python, the first sample data can be parsed through the following code:

[0223] first_system_data = []

[0224] for table in first_system_results:

[0225] for record in table.records:

[0226] asset_identity = record.get_value('asset_identity_code');

[0227] status_sequence = record.get_value('status_code_sequence');

[0228] business_sequence = record.get_value('business_code_sequence');

[0229] first_system_data.append({"asset_identity": asset_identity, "status_sequence": status_sequence, "business_sequence": business_sequence});

[0230] S302. Perform an intersection operation on the asset identity codes of the first sample data and the second sample data, and output the asset identity codes that do not belong to the intersection as differential data.

[0231] Extract the asset identity code lists from the first sample data and the second sample data respectively. For the first sample data first_system_data:

[0232] first_system_identities = [data["asset_identity"] for data in first_system_data]。

[0233] For the second sample data second_system_data:

[0234] second_system_identities = [data["asset_identity"] for data in second_system_data]。

[0235] Use set operations in Python to find the intersection. Convert the two lists to sets and then find the intersection:

[0236] first_set = set(first_system_identities);

[0237] second_set = set(second_system_identities);

[0238] common_identities = first_set.intersection(second_set)。

[0239] Find the asset identity codes that do not belong to the intersection. For the part of the first sample data that does not belong to the intersection:

[0240] first_difference_identities = first_set - common_identities;

[0241] first_difference_data = [data for data in first_system_data if data["asset_identity"] in first_difference_identities];

[0242] For the part of the second sample data that does not belong to the intersection:

[0243] second_difference_identities = second_set - common_identities;

[0244] second_difference_data = [data for data in second_system_data if data["asset_identity"] in second_difference_identities];

[0245] Output first_difference_data and second_difference_data as difference data.

[0246] S303. Remove duplicates from the status code sequences corresponding to the same asset identity code in the first sample data and the second sample data, and compare the consistency of the two status code sequences after deduplication. If the two are inconsistent, output these two status code sequences as difference data.

[0247] For the first sample data, construct a dictionary that maps the asset identity code to its corresponding status code sequence:

[0248] first_identity_to_status = {};

[0249] # Traverse the sample data of the first system;

[0250] for data in first_system_data:

[0251] # Obtain the asset identity code;

[0252] identity = data["asset_identity"];

[0253] # Obtain the status code sequence;

[0254] status_sequence = data["status_sequence"];

[0255] # If the asset identity code is not in the mapping dictionary, initialize an empty list;

[0256] if identity not in first_identity_to_status:

[0257] first_identity_to_status[identity] = [];

[0258] # Add the status code sequence to the list corresponding to the asset identity code;

[0259] first_identity_to_status[identity].append(status_sequence);

[0260] Perform the same operation on the second sample data:

[0261] second_identity_to_status = {};

[0262] # Traverse the sample data of the second system;

[0263] for data in second_system_data:

[0264] # Obtain the asset identity code;

[0265] identity = data["asset_identity"];

[0266] # Obtain the status code sequence;

[0267] status_sequence = data["status_sequence"];

[0268] # If the asset identity code is not in the mapping dictionary, initialize an empty list;

[0269] if identity not in second_identity_to_status:

[0270] second_identity_to_status[identity] = [];

[0271] # Add the status code sequence to the list corresponding to the asset identity code;

[0272] second_identity_to_status[identity].append(status_sequence).

[0273] Deduplicate the list of status code sequences corresponding to each asset identity code. For the mapping dictionary first_identity_to_status of the first sample data:

[0274] for identity, status_list in first_identity_to_status.items():

[0275] # Convert the list of status encoding sequences to a set of tuples to remove duplicates;

[0276] unique_status_list = list(set(tuple(status) for status in status_list));

[0277] # Convert the set of deduplicated tuples back to a list and update the mapping dictionary;

[0278] first_identity_to_status[identity] = [list(status) for status in unique_status_list];

[0279] Deduplicate the list of status encoding sequences corresponding to each asset identity code. For the mapping dictionary first_identity_to_status of the first sample data:

[0280] for identity, status_list in first_identity_to_status.items():

[0281] # Convert the list of status encoding sequences to a set of tuples to remove duplicates;

[0282] unique_status_list = list(set(tuple(status) for status in status_list));

[0283] # Convert the set of deduplicated tuples back to a list and update the mapping dictionary;

[0284] first_identity_to_status[identity] = [list(status) for status in unique_status_list];

[0285] Perform the same operation on the mapping dictionary second_identity_to_status of the second sample data:

[0286] for identity, status_list in second_identity_to_status.items():

[0287] # Convert the list of status encoding sequences to a set of tuples to remove duplicates;

[0288] unique_status_list = list(set(tuple(status) for status in status_list));

[0289] # Convert the deduplicated set of tuples back to a list and update the mapping dictionary;

[0290] second_identity_to_status[identity] = [list(status) for status in unique_status_list];

[0291] Traverse the common asset identity codes (common_identities obtained from the intersection operation), and compare the corresponding deduplicated status code sequences.

[0292] status_difference_data = []

[0293] # Traverse the common asset identity codes;

[0294] for identity in common_identities:

[0295] # Get the deduplicated status code sequence corresponding to the asset identity code in the first system;

[0296] first_status = first_identity_to_status[identity];

[0297] # Get the deduplicated status code sequence corresponding to the asset identity code in the second system;

[0298] second_status = second_identity_to_status[identity];

[0299] # If the two status code sequences are inconsistent;

[0300] if first_status!= second_status:

[0301] # Store the difference data in a list;

[0302] status_difference_data.append({"asset_identity": identity, "first_status": first_status, "second_status": second_status});

[0303] Traverse the common asset identity codes (common_identities obtained from the intersection operation), and compare the corresponding deduplicated status code sequences.

[0304] status_difference_data = [];

[0305] # Traverse the common asset identity codes;

[0306] for identity in common_identities:

[0307] # Obtain the deduplicated status code sequence corresponding to the asset identity code in the first system;

[0308] first_status = first_identity_to_status[identity];

[0309] # Obtain the deduplicated status code sequence corresponding to the asset identity code in the second system;

[0310] second_status = second_identity_to_status[identity];

[0311] # If the two status code sequences are inconsistent;

[0312] if first_status!= second_status:

[0313] # Store the difference data in a list;

[0314] status_difference_data.append({"asset_identity": identity, "first_status": first_status, "second_status": second_status});

[0315] Output status_difference_data as the difference data.

[0316] S304. Process the business code sequences corresponding to the same asset identity codes in the first sample data and the second sample data by removing zero vectors, and compare the consistency of the two processed business code sequences. If they are inconsistent, output these two business code sequences as differential data.

[0317] Similar to the way of constructing the status code sequence mapping, for the first sample data, construct a mapping dictionary from asset identity codes to business code sequences:

[0318] first_identity_to_business = {};

[0319] # Traverse the sample data of the first system;

[0320] for data in first_system_data:

[0321] # Obtain the asset identity code;

[0322] identity = data["asset_identity"];

[0323] # Obtain the business code sequence;

[0324] business_sequence = data["business_sequence"];

[0325] # If the asset identity code is not in the mapping dictionary, initialize an empty list;

[0326] if identity not in first_identity_to_business:

[0327] first_identity_to_business[identity] = [];

[0328] # Add the business code sequence to the list corresponding to the asset identity code;

[0329] first_identity_to_business[identity].append(business_sequence);

[0330] Perform the same operation on the second sample data:

[0331] second_identity_to_business = {};

[0332] # Traverse the sample data of the second system;

[0333] for data in second_system_data:

[0334] # Obtain the asset identity code;

[0335] identity = data["asset_identity"];

[0336] # Obtain the business code sequence;

[0337] business_sequence = data["business_sequence"];

[0338] # If the asset identity code is not in the mapping dictionary, initialize an empty list;

[0339] if identity not in second_identity_to_business:

[0340] second_identity_to_business[identity] = [];

[0341] # Add the business code sequence to the list corresponding to the asset identity code;

[0342] second_identity_to_business[identity].append(business_sequence);

[0343] Assume the zero vector is [0, 0,...] (determined according to the actual vector dimension). For the mapping dictionary first_identity_to_business of the first sample data:

[0344] # Set the zero vector according to the actual vector dimension;

[0345] zero_vector = [0] * vector_dim;

[0346] for identity, business_list in first_identity_to_business.items():

[0347] # Filter out the sequences in the business code sequence list that are not the zero vector;

[0348] non_zero_business_list = [business for business in business_list if business!= zero_vector];

[0349] # Update the mapping dictionary to only keep the business code sequences of non-zero vectors;

[0350] first_identity_to_business[identity] = non_zero_business_list;

[0351] Perform the same operation on the mapping dictionary second_identity_to_business of the second sample data:

[0352] for identity, business_list in second_identity_to_business.items():

[0353] # Filter out the sequences in the business code sequence list that are not zero vectors;

[0354] non_zero_business_list = [business for business in business_list if business!= zero_vector];

[0355] # Update the mapping dictionary to only keep the business code sequences of non-zero vectors;

[0356] second_identity_to_business[identity] = non_zero_business_list;

[0357] Traverse the common asset identity codes (common_identities) and compare the corresponding business code sequences after removing zero vectors.

[0358] business_difference_data = [];

[0359] # Traverse the common asset identity codes;

[0360] for identity in common_identities:

[0361] # Obtain the business code sequence after removing zero vectors corresponding to the asset identity code in the first system;

[0362] first_business = first_identity_to_business[identity];

[0363] # Obtain the business code sequence after removing the zero vector for the asset identity code corresponding to the second system;

[0364] second_business = second_identity_to_business[identity];

[0365] # If the two business code sequences are inconsistent;

[0366] if first_business!= second_business:

[0367] # Store the difference data in a list;

[0368] business_difference_data.append({"asset_identity": identity, "first_business": first_business, "second_business": second_business});

[0369] Output business_difference_data as the difference data.

[0370] In an embodiment of the present invention, based on step S4, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation scheme.

[0371] S401. Import the difference data and the associated data of the difference data into a pre - constructed prompt word template to obtain a prompt word, and the prompt word template includes the requirements for exporting the inference text.

[0372] The difference data has been obtained in the previous steps, including the difference data that does not belong to the intersection of the asset identity codes (first_difference_data and second_difference_data), the difference data with inconsistent status code sequences (status_difference_data), and the difference data with inconsistent business code sequences (business_difference_data).

[0373] Associated data may include basic information of assets, such as asset names, affiliated departments, etc., which can be obtained from the original asset data. Assume that the original asset data is stored in a DataFrame asset_info_df containing asset information and is associated through asset identity codes. For example, for each piece of difference data in first_difference_data, the associated asset name can be obtained through the following code:

[0374] import pandas as pd;

[0375] asset_info_df = pd.read_csv('asset_info.csv') # Assume that asset information is stored in a CSV file;

[0376] for data in first_difference_data:

[0377] identity = data["asset_identity"];

[0378] related_info = asset_info_df[asset_info_df["asset_identity_code"] == identity];

[0379] if not related_info.empty:

[0380] data["asset_name"] = related_info["asset_name"].values[0];

[0381] The prompt template should clearly and explicitly elaborate on the task requirements, including requirements for the format, content, etc. of the reasoning text. The following is an example template:

[0382] Asset Data Difference Analysis Task;

[0383] # Difference Data Description;

[0384] 1. **Asset Identity Difference**: {asset_identity_difference_description};

[0385] 2. **Status Coding Sequence Difference**: {status_sequence_difference_description};

[0386] 3. **Business Coding Sequence Difference**: {business_sequence_difference_description};

[0387] # Associated Data;

[0388] 1. **Asset Name**: {asset_name};

[0389] 2. **Department**: {department} # Assume there is department information, which needs to be extracted from the associated data according to the actual situation;

[0390] Inference Text Requirements;

[0391] 1. Analyze in detail the possible reasons for the above - mentioned difference data, and list at least 3 reasons.

[0392] 2. For each reason, provide possible solutions.

[0393] 3. Give the recommended value of the correct data corresponding to the difference data and explain the reasons.

[0394] Please analyze and reason according to the above requirements.

[0395] Fill the difference data and associated data into the prompt - word template. For each type of difference data, create a prompt word. Take status_difference_data as an example:

[0396] prompt_templates = [];

[0397] for data in status_difference_data:

[0398] asset_identity = data["asset_identity"];

[0399] first_status = data["first_status"];

[0400] second_status = data["second_status"];

[0401] asset_name = data.get("asset_name", "Unknown") # If the asset name is not obtained, set it to Unknown;

[0402] department = data.get("department", "Unknown") # Assume the way to obtain department information is similar;

[0403] asset_identity_difference_description = f"The asset with asset identity code {asset_identity} has a difference in the status code sequence between the first sample data and the second sample data. The status code sequence in the first sample data is {first_status}, and the status code sequence in the second sample data is {second_status}";

[0404] status_sequence_difference_description = f"Status code sequence in the first sample data: {first_status}; Status code sequence in the second sample data: {second_status}";

[0405] business_sequence_difference_description = "" # This template is for the difference in status code sequence, and the description of the difference in business code sequence is empty;

[0406] filled_template = f""";

[0407] Asset data difference analysis task;

[0408] # Difference data description;

[0409] 1. **Asset identity difference**: {asset_identity_difference_description};

[0410] 2. **Status code sequence difference**: {status_sequence_difference_description};

[0411] 3. **Business code sequence difference**: {business_sequence_difference_description};

[0412] # Associated data;

[0413] 1. **Asset name**: {asset_name};

[0414] 2. **Department**: {department};

[0415] Requirements for inferring text;

[0416] 1. Analyze in detail the possible reasons for the above - mentioned differential data, and list at least 3 reasons.

[0417] 2. For each reason, provide possible solutions.

[0418] 3. Give the recommended values of the correct data corresponding to the differential data and explain the reasons.

[0419] Please conduct the analysis and inference according to the above requirements.

[0420] prompt_templates.append(filled_template);

[0421] Perform similar filling operations on other types of differential data (such as business_difference_data and asset identity coding differential data).

[0422] S402. Input the prompt into a large - model. The large - model has been pre - trained, and the output layer of the large - model contains a decoder.

[0423] Select a large - model with the ability of chain of thought, such as GPT - 4 - Turbo, etc. If using the OpenAI API, first install the openai library. Set the API key and perform a simple test call through the following code:

[0424] import openai;

[0425] openai.api_key = "your - api - key".

[0426] Send the filled prompt to the large - model. Assume that prompt_templates is a list containing all prompts, and use the openai.Completion.create method (for GPT - 3 series models) or openai.ChatCompletion.create method (for GPT - 4 series models) to send requests. Taking GPT - 4 as an example:

[0427] responses = [];

[0428] for prompt in prompt_templates:

[0429] response = openai.ChatCompletion.create(

[0430] model="gpt - 4",

[0431] messages=

[0432] {"role": "user", "content": prompt}]);

[0433] responses.append(response)。

[0434] S403. Obtain the inference text generated by the large - model and the correct data corresponding to the differential data.

[0435] The response of the large - model contains the generated inference text and possible correct data suggestions. For the response of GPT - 4, the inference text is located in response['choices'][0]['message']['content']. For the correct data suggestions, they need to be extracted according to the content of the inference text. Assume that the correct data suggestions are given in the format of "The suggested correct data is: {correct data}" in the inference text, and they are extracted through the following code:

[0436] inferred_texts_and_correct_data = []

[0437] for response in responses:

[0438] inferred_text = response['choices'][0]['message']['content'];

[0439] correct_data = ""

[0440] start_index = inferred_text.find("The suggested correct data is:");

[0441] if start_index!= -1:

[0442] end_index = inferred_text.find("\n", start_index);

[0443] if end_index!= -1:

[0444] correct_data = inferred_text[start_index + len("Suggested correct data:"):end_index].strip();

[0445] inferred_texts_and_correct_data.append({"inferred_text": inferred_text, "correct_data": correct_data});

[0446] Store the obtained inferred texts and suggested correct data in a suitable data structure, such as a CSV file or a database. Taking storage in a CSV file as an example:

[0447] import pandas as pd;

[0448] df = pd.DataFrame(inferred_texts_and_correct_data);

[0449] df.to_csv('inferred_results.csv', index=False).

[0450] In an embodiment of the present invention, based on step S5, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.

[0451] S501. Perform NLP parsing on the inferred text and extract key elements from the analyzed text.

[0452] Entity recognition: Use a pre - trained model (such as BERT - NER) to recognize system names (EAS / EAM), data types (status / business / quantity), and conclusions (reasons for differences).

[0453] Relation extraction: Use dependency syntactic analysis to extract the inference logic chain.

[0454] For example, when the input text is "Since EAS marks the asset as scrapped while EAM shows it in use, it is determined that the EAS data is incorrect", the output structure is:

[0455] {"data_sources": ["EAS", "EAM"],

[0456] "compared_field": "Asset status",

[0457] "conclusion": "EAS is incorrect",

[0458] "reasoning_steps":

[0459] {"premise": "EAS status = scrapped", "premise_source": "EAS"},

[0460] {"premise": "EAM status = in use", "premise_source": "EAM"},

[0461] {"logic": "A is contradictory to B → A is wrong"}]}。

[0462] S502. Build a credibility rule base.

[0463] The storage form of the rule base is a structured rule table (compatible with rule engines such as Drools):

[0464] rules =

[0465] {"id": "rule1",

[0466] "condition": "field == 'asset status' AND source == 'EAS'",

[0467] "action": "raise_warning('The financial system status needs to be cross-validated with the EAM physical status')",

[0468] "priority": 0.9 # Rule priority},

[0469] {"id": "rule2",

[0470] "condition": "source == 'EAM' AND has_audit_log == False",

[0471] "action": "add_error('Lack of audit log, the credibility of EAM data is degraded')"}]。

[0472] The rules stored in the rule base include:

[0473] ‌1. System core function priority principle

[0474] The credibility determination of data related to the core functions of the system should first follow its functional design objectives. For data that supports core business processes (such as fund settlement data in financial trading systems and medical records in medical systems), a higher-standard data quality control process should be established, and its credibility weight coefficient should be higher than that of non-core function data (such as log backups and auxiliary analysis data).

[0475] Clarify the strong correlation between core function data and the system's strategic objectives, and refer to the requirement of "function guarantee and goal consistency" in the top-level design governance theory.

[0476] 2. Principles of data update mechanism The timeliness of data update needs to be dynamically evaluated in combination with business scenarios:

[0477] For real-time updated data (such as IoT sensor readings and stock market quotes), the credibility weight is increased by 30% - 50% in dynamic decision-making scenarios;

[0478] For non-real-time updated data (such as historical archives and basic configuration information), the baseline credibility is maintained in static analysis scenarios, but the maximum delay time threshold (such as ≤24 hours) needs to be marked.

[0479] New constraint: The real-time evaluation needs to exclude invalid updates caused by network jitter (filtered through the heartbeat detection mechanism).

[0480] 3. Principles of audit log traceability ability

[0481] Completeness: Record the operation traces of the entire data life cycle (create / modify / delete), and the missing rate ≤0.1%;

[0482] Retraceability: Support multi-dimensional traceability such as the identity of the operator, timestamp, and terminal device;

[0483] Immutability: Guaranteed by using blockchain evidence storage or digital signature technology;

[0484] Business relevance: The log granularity is positively correlated with the data security level (for example, financial data needs to record field-level changes).

[0485] Quantitative relationship: For every 10% increase in log completeness, the corresponding data credibility evaluation value increases by 5% - 8%.

[0486] 4. Principles of external evidence verification

[0487] External verification evidence needs to meet the VET principle:

[0488] Effectiveness: The verification content shall cover the key attributes of the data (for example, identity verification shall include dual verification of biometric features and ID number);

[0489] Timeliness: The validity period of the verification result is linked to the data update cycle (for example, the validity period of enterprise qualification verification is 1 year).

[0490] Priority rule: The credibility weight of external verification data with blockchain evidence preservation is increased by 2 times.

[0491] 5. Principle of error occurrence frequency

[0492] The calculation of the system error frequency shall be based on the sliding window model:

[0493] Set according to the business peak period (for example, the e-commerce system adopts a 72-hour window during major promotions);

[0494] Error type classification:

[0495] The weight coefficient of a fatal error (causing business interruption) is 1.0;

[0496] The coefficient of a serious error (affecting data integrity) is 0.7;

[0497] The coefficient of a general error (local function abnormality) is 0.3.

[0498] Inverse proportion formula: Credibility = 1 / (1 + α × error frequency), where α is the business impact factor (α = 2 for the financial system and α = 1.5 for the government affairs system).

[0499] 6. Principle of data consistency scope

[0500] Three-layer verification mechanism shall be implemented for logical consistency verification:

[0501] Cross-system consistency: Compare multi-source data through the centralized data center (for example, the inventory quantity difference between the financial system and the asset management system ≤ 0.5%);

[0502] Temporal consistency: Check whether the evolution of the data version conforms to the business rules (for example, the order status must change in the order of "created → paid → shipped");

[0503] Coupling degree of associated data: The tolerance of the logical deviation between the core data and its derivative data ≤ 3% (for example, the deviation threshold between the production plan data and the actual production report).

[0504] Exception handling: Trigger hierarchical alarms when the consistency verification fails (manual intervention is required within 30 minutes for level 1 alarms).

[0505] S50‌3. Build a knowledge graph for storing a priori knowledge such as system functions and data relevance.

[0506] Example of the data structure of the knowledge graph:

[0507] { "EAS": {

[0508] "core_function": ["Asset valuation", "Depreciation calculation"],

[0509] "data_quality": {"update_frequency": "Daily", "error_rate": 0.05},

[0510] "related_fields": ["Financial status", "Business type"]},

[0511] "EAM": {"core_function": ["Equipment status", "Maintenance records"],

[0512] "data_quality": {"update_frequency": "Real-time", "error_rate": 0.12},

[0513] "related_fields": ["Physical location", "Operating status"]}}.

[0514] S50‌4. Rule matching and verification.

[0515] (1)Extract data sources, field types, and logical relationships from the NLP parsing results.

[0516] (2)Query the knowledge graph to obtain system functions and data quality indicators.

[0517] (3)Use the Rete algorithm to match each rule in the rule library one by one and trigger verification actions. Sort conflicting rules by priority (e.g., Principle 1 > Principle 4).

[0518] The Rete algorithm is an efficient pattern matching algorithm designed specifically for rule engines. The Rete network includes:

[0519] The pattern network (α network) is used to process single-condition atomic patterns (such as "Temperature > 30°C"). Each condition corresponds to an α node, forming a filtering chain;

[0520] ‌Connection network (β network)‌ is used to process ‌multi-condition combination patterns‌ (such as "temperature>30℃ AND pressure<100kPa"), and realize the connection operation between conditions through β nodes.

[0521] The Rete algorithm decomposes each rule in the rule base into an α network (processing single-condition atomic mode) and a β network (processing multi-condition combination mode) to form a tree structure; detects repeated conditions between different rules and merges the same α nodes. A hash index is established to speed up condition matching and reduce repeated calculations.

[0522] The following is the process of verifying the asset status analysis text:

[0523] # Input text

[0524] text = "The EAS system shows that the status of asset A001 is 'Scrapped', while the EAM system shows it as 'In Use'. Since EAS is a financial system, its scrapped status should be given priority, so it is determined that the EAM data has not been updated in time."

[0525] # Step 1: NLP parsing

[0526] analysis_data = nlp_parser(text);

[0527] # Output:

[0528] {"sources": ["EAS", "EAM"],

[0529] "fields": ["Asset Status"],

[0530] "conclusion": "EAM data error",

[0531] "logic_chain": ["EAS status = scrap → EAS accepted → EAM error"]};

[0532] # Step 2: Rule matching

[0533] violations = []

[0534] for rule in rule_engine.match(analysis_data):

[0535] if rule['id'] == 'rule1':

[0536] violations.append(f"Violation of Principle 1: The asset status should prioritize EAM over EAS")

[0537] # Step 3: Knowledge graph query

[0538] eam_quality = knowledge_graph.get("EAM.data_quality.error_rate")

[0539] if eam_quality > 0.1:

[0540] violations.append(f"The historical error rate of EAM is high ({eam_quality}), additional verification is required")

[0541] # Step 4: Generate report

[0542] report = generate_report(violations, suggestions)

[0543] Output report:

[0544] - [Severe] Violation of Principle 1: The asset status is a physical operation result, and EAM should be prioritized over EAS.

[0545] - [Warning] The historical error rate of EAM is high (0.12), check its maintenance work orders and sensor data.

[0546] Suggested actions:

[0547] 1. Verify the real-time sensor status of asset A001 in the EAM system.

[0548] 2. Compare the EAS scrapping records with the approval logs.

[0549] To further illustrate the cross-system asset data reconciliation method provided in this application, a specific application scenario is provided, such as Figure 2 shown, the first system is the EAM asset management system, and the second system is the EAS financial system.

[0550] The first system (EAM system): Manages the full life cycle data of fixed assets such as production equipment and transportation tools, and its core function is equipment maintenance and usage status tracking.

[0551] The second system (EAS system): Records financial accounting data such as asset depreciation and cost allocation, and its core function is financial compliance auditing.

[0552] 1. Obtain asset data from two systems:

[0553] # Obtain EAM system asset data through API (example)

[0554] eam_data = {

[0555] "asset_id": "EQP-2024-001",

[0556] "static_info": {"Category": "CNC Machine Tool", "Purchase Date": "2023-05-12"},

[0557] "dynamic_states":

[0558] {"timestamp": "2024-03-01 09:00", "Status": "Running", "Fault Code": null},

[0559] {"timestamp": "2024-03-02 14:30", "Status": "Shutdown for Maintenance", "Fault Code": "F-1023"}],

[0560] "business_logs": [{"Operation Time": "2024-03-01", "Business Type": "Production Task Assignment", "Work Order Number": "WO-240301A"}]};

[0561] # Obtain EAS system asset data through WebService;

[0562] eas_data = { "asset_id": "FA-2024-001A",

[0563] "static_info": {"Asset Name": "Machining Center", "Recording Date": "2023-05-15"},

[0564] "dynamic_states": [ {"Accounting Period": "2024-03", "Depreciation Method": "Double Declining Balance", "Net Value": "¥1,230,000"}],

[0565] "business_logs": [{"Voucher Number": "V-20240301-001", "Abstract": "Equipment Depreciation Provision", "Amount": "¥15,000"}]}。

[0566] 2. Semantic Encoding and Data Alignment

[0567] Perform data cleaning to unify the asset ID format (map "EQP-2024-001" in EAM and "FA-2024-001A" in EAS to the common code "ASSET#CNC-001").

[0568] The methods of text splicing include: converting dynamic states into readable texts.

[0569] EAM status text: "2024-03-02 14:30 Shutdown for maintenance (fault code F-1023)"

[0570] EAS status text: "Depreciation net value ¥1,230,000 in the accounting period of March 2024".

[0571] Phase - by - phase encoding:

[0572] The first phase: Identity encoding, with the input being asset category + acquisition date and the output being a 512 - dimensional vector;

[0573] The second phase: Status encoding, with the input being the shutdown maintenance event text and the output being a 768 - dimensional vector;

[0574] The third phase: Business encoding, with the input being the depreciation provision voucher description and the output being a 768 - dimensional vector.

[0575] Temporal alignment, perform time - window downsampling on the minute - level status data of EAM and the monthly financial data of EAS:

[0576] Example of storing in the time - series database:

[0577] INSERT INTO asset_encodings VALUES (

[0578] 'ASSET#CNC-001',

[0579] [0.23, -0.56, ..., 0.78],-- Identity encoding

[0580] [[0.12, 0.45, ...],-- Status encoding on 2024 - 03 - 01

[0581] [0.67, -0.33, ...]-- Status encoding on 2024 - 03 - 02 ],

[0582] [[0.89, 0.11, ...]-- Business encoding for production task assignment])

[0583] 3. Difference Detection and Inference Verification

[0584] Missing identity: Financial records of the newly added asset "ASSET#ROBOT-005" in EAM were not found in EAS;

[0585] Status contradiction: EAM shows that the equipment "ASSET#CNC-001" was out of service on March 2, but EAS still accrued full depreciation during the same period;

[0586] Business conflict: EAM records that the output corresponding to work order "WO-240301A" is 1,000 pieces, while EAS calculates the cost sharing based on 800 pieces.

[0587] Large model inference example:

[0588] {"Input prompt": "It is detected that the equipment ASSET#CNC-001 had a shutdown maintenance on 2024-03-02, but the financial system still accrued depreciation as if it was in normal use. Please analyze the possible reasons and give the correct value. It is required to output: 1) Fault impact assessment 2) Depreciation calculation suggestion",

[0589] "Inference text": "According to the EAM maintenance log, the F-1023 fault is due to the damage of the main shaft motor, and the estimated repair time is 48 hours. According to Article 18 of Accounting Standards for Enterprises No. 4 - Fixed Assets, depreciation should be suspended during the abnormal shutdown period. It is recommended that the depreciation amount for this month be revised to: 1,230,000 × 5% × (28 / 31) = ¥55,548",

[0590] "Correct data": {"Accounting period": "2024-03", "Depreciation amount": "¥55,548"}}.

[0591] 4. Rule verification and credibility determination

[0592] Using multi-rule collaborative verification, for example:

[0593] Core function priority: EAM equipment status data vs EAS depreciation calculation data, EAM data is prior;

[0594] External evidence verification: The large model quotes accounting standards clauses, increasing the credibility weight;

[0595] Temporal consistency: After the depreciation adjustment, it matches the production work order quantity, and passes the verification.

[0596] The technical effects of this method:

[0597] Improved reconciliation efficiency: Compared with traditional manual reconciliation, the present invention reduces the reconciliation time for assets at the ten-thousand level from 72 hours to 4 hours;

[0598] Abnormal detection accuracy: Through the collaboration of semantic encoding and rules, the accuracy of difference recognition reaches 98.7% (measured data);

[0599] Compliance guarantee: The reasoning and verification combined with accounting standards improve the compliance rate of financial adjustments to 99.5%.

[0600] This embodiment has been verified to be successfully applied in port enterprises, effectively solving the data island problem between physical asset management and financial accounting.

[0601] In some embodiments, the cross-system asset data reconciliation system may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the cross-system asset data reconciliation system may be stored in the memory of the computer device and executed by at least one processor to execute (see Figure 1 description) the functions of cross-system asset data reconciliation.

[0602] In this embodiment, the cross-system asset data reconciliation system can be divided into multiple functional modules according to the functions it performs, such as Figure 3 shown. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0603] An acquisition module, configured to respectively acquire the asset data of the first system and the second system;

[0604] An encoding module, configured to respectively encode the asset data of the first system and the asset data of the second system by using a semantic encoder to obtain the first sample data of the first system and the second sample data of the second system;

[0605] A first processing module, configured to screen out the difference data between the first sample data and the second sample data;

[0606] A second processing module, configured to input the difference data into a large model to obtain the inference text of the large model and the correct data corresponding to the difference data output by the large model;

[0607] A verification module, configured to verify the inference text by using a preset rule. If it is confirmed that the inference text passes the verification, it is determined that the correct data is credible.

[0608] Figure 4The cross-system asset data reconciliation method provided by the embodiments of this application can be applied to devices. Those skilled in the art can understand that the device structure involved in the embodiments of the present invention does not constitute a limitation on the device. The device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the device includes, but is not limited to, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of this application described herein and / or claimed.

[0609] Among them, the device 400 may include: a processor 410, a memory 420, and a communication unit 430. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present invention. It can be a bus structure, a star structure, or may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0610] Among them, the memory 420 can be used to store the execution instructions of the processor 410. The memory 420 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. When the execution instructions in the memory 420 are executed by the processor 410, the device 400 is enabled to execute some or all of the steps in the above method embodiments.

[0611] The processor 410 is the control center of the storage device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 420, and by calling the data stored in the memory, it executes various functions of the electronic device and / or processes data. The processor may be composed of an integrated circuit (IC), for example, it may be composed of a single packaged IC, or may be composed of multiple packaged ICs with the same or different functions connected. For example, the processor 410 may only include a central processing unit (CPU). In the embodiments of the present invention, the CPU may be a single arithmetic core or may include multiple arithmetic cores.

[0612] A communication unit 430 is configured to establish a communication channel so that the storage device can communicate with other devices, receive user data sent by other devices, or send user data to other devices.

[0613] The present invention also provides a computer storage medium. The computer storage medium can store a program which, when executed, may include some or all of the steps in the embodiments provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.

[0614] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes, and includes several instructions for causing a computer device (which may be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present invention.

[0615] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the descriptions in the method embodiments for the relevant parts.

[0616] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the systems or modules can be in electrical, mechanical, or other forms.

[0617] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical module, that is, it may be located in one place or may be distributed across multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0618] In addition, in each embodiment of the present invention, each functional module may be integrated in a processing module, may exist physically separately for each module, or two or more modules may be integrated in one module.

[0619] Although the present invention has been described in detail by referring to the drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, and all should be covered within the protection scope of the present invention.

Claims

1. A cross-system asset data reconciliation method, characterized in that, Including: Obtain the asset data of the first system and the second system respectively; Use a semantic encoder to encode the asset data of the first system and the asset data of the second system respectively, obtaining the first sample data of the first system and the second sample data of the second system; Filter out the difference data between the first sample data and the second sample data; Input the difference data into a large model, obtaining the inference text of the large model and the correct data corresponding to the difference data output by the large model; Use a preset rule to verify the inference text. If it is confirmed that the inference text passes the verification, then determine that the correct data is credible.

2. The method according to claim 1, wherein Respectively obtaining the asset data of the first system and the second system includes: Obtain the first asset data set from the external interface of the first system; Obtain the second asset data set from the external interface of the second system; The first system is an asset management system, and the second system is a financial system; The first asset data set includes the static data and dynamic data of the assets in the first system; The second asset data set includes the static data and dynamic data of the assets in the second system; The dynamic data in the first asset data set and the second asset data set both include the asset status and asset operations generated by the assets within a period of time, and the static data both includes asset identity information.

3. The method according to claim 1, characterized in that, Using a semantic encoder to encode the asset data of the first system and the asset data of the second system respectively includes: Perform data cleaning and text splicing on the asset data to convert the identity fields, status values, and business fields in the asset data into text-format information; Use a semantic encoder to perform staged encoding on the asset data after data cleaning and text splicing. In the first stage, generate corresponding asset identity codes according to each asset identity information in the asset data. In the second stage, generate corresponding asset status codes according to each status information in the asset data. In the third stage, generate corresponding asset operation codes according to each business information in the asset data; Divide the asset identity codes, asset status codes, and asset operation codes of the same asset into the same data group; Sort the asset status codes and asset operation codes in the same data group according to the time stamps of the corresponding original data respectively, obtaining a status code sequence and a business code sequence; Align the status code sequence and the business code sequence by filling with previous value vectors or zero vectors; Store the asset identity codes and the corresponding status code sequences and business code sequences in a time series database.

4. The method according to claim 3, wherein Filter out the difference data between the first sample data and the second sample data, including: Obtain the first sample data and the second sample data from the time series database. The first sample data and the second sample data both include multiple asset identity codes and the corresponding multiple status code sequences and business code sequences; Perform an intersection operation on the asset identity codes of the first sample data and the second sample data, and output the asset identity codes that do not belong to the intersection as difference data; Deduplicate the status code sequences corresponding to the same asset identity code in the first sample data and the second sample data, and compare the consistency of the two status code sequences after deduplication. If the two are inconsistent, output these two status code sequences as differential data; Remove zero vectors from the business code sequences corresponding to the same asset identity code in the first sample data and the second sample data, and compare the consistency of the two business code sequences after processing. If the two are inconsistent, output these two business code sequences as differential data.

5. The method according to claim 4, characterized in that, Input the differential data into a large model to obtain the inference text of the large model and the correct data corresponding to the differential data output by the large model, including: Import the differential data and the associated data of the differential data into a pre-constructed prompt template to obtain a prompt. The prompt template includes the requirements for exporting the inference text; Input the prompt into the large model. The large model has been pre-trained, and the output layer of the large model includes a decoder; Obtain the inference text generated by the large model and the correct data corresponding to the differential data.

6. The method according to claim 1, wherein Verify the inference text using a preset rule. If it is confirmed that the inference text passes the verification, then determine that the correct data is trustworthy, including: Extract key elements from the inference text using natural language processing techniques. The key elements include data source, field type, and logical relationship; Obtain the system functions and data quality indicators of the first system and the second system; Use a rule execution algorithm to retrieve preset rules one by one from the rule library, and verify the key elements based on the system functions and data quality indicators.

7. The method according to claim 6, wherein The preset rules include: The principle of system core function priority, which stipulates that the credibility of data related to the system core function is higher than that of data related to the non-core function of the system; The principle of data update mechanism, which stipulates that the credibility of real-time updated data is higher than that of non-real-time updated data; The principle of audit log traceability ability, which stipulates the positive proportional relationship between the integrity of the audit log and the credibility of the data; The principle of external evidence verification, which stipulates that data with external evidence has higher credibility; The principle of error occurrence frequency, which stipulates the inverse proportional relationship between the error occurrence frequency of the system and the credibility of the system data; The principle of data consistency range, which stipulates that the credibility of data is verified by the logical consistency of the data with other relevant data.

8. A cross-system asset data reconciliation system, characterized in that, Include: An acquisition module for respectively acquiring the asset data of the first system and the second system; An encoding module for respectively encoding the asset data of the first system and the asset data of the second system using a semantic encoder to obtain the first sample data of the first system and the second sample data of the second system; A first processing module for screening out the differential data between the first sample data and the second sample data; A second processing module for inputting the differential data into the large model to obtain the inference text of the large model and the correct data corresponding to the differential data output by the large model; A verification module for verifying the inference text using a preset rule. If it is confirmed that the inference text passes the verification, then determine that the correct data is trustworthy.

9. A device, characterized in that, Include: A memory for storing the cross-system asset data reconciliation program; A processor for implementing the steps of the cross-system asset data reconciliation method according to any one of claims 1-7 when executing the cross-system asset data reconciliation program.

10. A computer-readable storage medium storing a computer program, characterized in that, A cross-system asset data reconciliation program is stored on the readable storage medium, and when the cross-system asset data reconciliation program is executed by a processor, the steps of the cross-system asset data reconciliation method according to any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Terminal and method for providing a location-based service outside a service zone

    WO2012030000A1

  • Gas sensor

    WO2020230515A1

  • Financial business accounting system and device based on NCC architecture

    CN115099914A

  • Retrieval method and device for digital asset management system

    CN118260440A