Data reproduction methods and apparatus, storage media, and computer equipment

By combining OCR and visual models, the problems of incomplete and inaccurate image data reproduction have been solved, achieving efficient and automated data reproduction and improving the accuracy and completeness of the data.

CN120726650BActive Publication Date: 2025-11-14RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511211855.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-14
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing technologies face difficulties in extracting raw data from image-based data, especially in accurately reconstructing image data with complex structures, resulting in incomplete and inaccurate data reproduction.

Method used

By combining optical character recognition (OCR) with a formatted data extraction model and a visual model, structured data extraction and reconstruction of image information can be achieved. The visual model is used to understand the table structure and logical relationships in the image, and the accuracy and integrity of the data are ensured by combining format verification and feedback retry mechanism.

Benefits of technology

It improves the efficiency and accuracy of data reproduction, automates the processing of image data, reduces labor and time costs, and ensures the integrity and consistency of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726650B_ABST
    Figure CN120726650B_ABST
Patent Text Reader

Abstract

This application discloses a data reproduction method, apparatus, storage medium, and computer device. The method includes: acquiring image information generated from original data in a specific data domain; performing optical character recognition (OCR) on the image information, and extracting structured data from the OCR results using a preset formatted data extraction model to obtain structured information corresponding to the image information; and reproducing the original data based on the structured information and the image information using a preset visual model to obtain reproduced data corresponding to the original data. This application breaks through the limitations of traditional data acquisition from images, improving acquisition efficiency and accuracy through OCR and a formatted data extraction model; and reproducing the original data completely and consistently by combining the visual model with structured and image information, thus achieving automated data reproduction and improving the efficiency and accuracy of data reproduction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a data reproduction method and apparatus, storage medium, and computer equipment. Background Technology

[0002] Currently, reproducing original data from image-based data faces numerous difficulties and challenges. On one hand, when original data exists in image form, the static visual information inherent in images lacks directly editable and parsable structured features, making accurate extraction of original data from them extremely complex. Traditional data acquisition methods often fail to directly extract valuable information from images; for example, simple text copy-paste operations cannot be applied to image content, severely hindering the first step in data reproduction—data acquisition.

[0003] On the other hand, some existing technologies have significant shortcomings in processing image data reproduction. While some methods based on simple character recognition can identify text in images, the results are often just simple text strings, lacking an understanding of the data's inherent structure and logical relationships. For example, for images containing complex structures such as tables and charts, simple character recognition can only obtain fragmented text information, failing to reconstruct the complete structure and relationships of the original data. This results in incomplete and inaccurate reproduced data, making it difficult to meet the needs of practical applications. Summary of the Invention

[0004] In view of this, the embodiments of this application provide a data reproduction method and apparatus, storage medium, and computer equipment, which effectively breaks down the information barriers between specific data domains and service systems while ensuring data security. This allows structured data that was originally unusable due to data isolation to participate in service analysis and service decision-making again, thereby improving the comprehensiveness and real-time response capability of enterprise data analysis.

[0005] According to one aspect of this application, a data reproduction method is provided, the method comprising:

[0006] Retrieve image information generated from raw data in a specific data domain;

[0007] Optical character recognition is performed on the image information, and structured data is extracted from the optical character recognition results using a preset formatted data extraction model to obtain the structured information corresponding to the image information.

[0008] By using a preset visual model, the original data is reproduced based on the structured information and the image information to obtain the reproduced data corresponding to the original data.

[0009] Optionally, the step of reproducing the original data based on the structured information and the image information using a preset visual model to obtain the reproduced data corresponding to the original data includes:

[0010] The structured information and the image information are input into the preset visual model. The preset visual model performs image parsing on the image information with the structured information as a constraint, and outputs the image parsing result within the constraint range of the structured information as the reproduction data corresponding to the original data.

[0011] Optionally, obtaining image information generated based on raw data in a specific data domain includes:

[0012] Query the original data in the specific data domain according to the query conditions;

[0013] The original data is visualized and the corresponding image information is determined based on the visualized data. The visualized data includes the image information and / or video information. The video information is converted into the image information through frame extraction.

[0014] Optionally, after reproducing the original data based on the structured information and the image information using a preset visual model to obtain the reproduced data corresponding to the original data, the method further includes:

[0015] The format of the reproduced data is validated according to the format conditions corresponding to the original data.

[0016] If the format verification fails, the reason for the failure is determined, and the original data is restored using the preset visual model based on the structured information, the image information, and the reason for the failure, until the reproduced data passes the format verification. The format conditions are obtained by acquiring the preset format conditions corresponding to the original data or by parsing the format of the image information.

[0017] Optionally, after reproducing the original data based on the structured information and the image information using a preset visual model to obtain the reproduced data corresponding to the original data, the method further includes:

[0018] If the image information includes multiple images, the reproduction data corresponding to each image information is merged and deduplicated to obtain the reproduction data corresponding to the original data.

[0019] Optionally, after obtaining the reproducible data corresponding to the original data, the method further includes:

[0020] The reproduced data is stored in a data warehouse, and a preset data repair model is invoked through a service call function pre-encapsulated in the data warehouse to repair the target reproduced data in the data warehouse;

[0021] Based on the repair results of the target reproduction data on the preset data repair model, the target reproduction data in the data warehouse is updated.

[0022] Optionally, the step of invoking a preset data repair model to repair the target reproducible data in the data warehouse by using a service call function pre-encapsulated in the data warehouse includes:

[0023] By using a service call function pre-encapsulated in the data warehouse, the target reproducible data that matches the data repair field in the reproducible data is queried. The target reproducible data is then concatenated with a prompt word template. The concatenated data reproducible prompt word is used to call the preset data repair model to repair the target reproducible data. The data reproducible prompt word is used to guide the preset data repair model to repair the target reproducible data.

[0024] Optionally, the step of concatenating the target reproduction data with the prompt word template, and using the concatenated data reproduction prompt word to call the preset data repair model to repair the target reproduction data, includes:

[0025] The target reproduction data, permission verification information, and prompt word template are concatenated. The concatenated data reproduction prompt word is used to call the preset data repair model to repair the target reproduction data. This allows the preset data repair model to perform access permission verification on the permission verification information in the data reproduction prompt word, and to repair the target reproduction data after the access permission verification is passed.

[0026] Optionally, before invoking a preset data repair model through a service call function pre-encapsulated in the data warehouse to repair the target reproducible data in the data warehouse, the method further includes:

[0027] Receive data reproduction instructions;

[0028] The data reproduction instruction is parsed for repair fields and permission verification information. If a repair field is parsed, it is used as the data repair field; otherwise, a preset data repair field is used as the data repair field.

[0029] Optionally, the data warehouse is stored in a service domain; after updating the target reproduction data in the data warehouse based on the repair results of the target reproduction data according to the preset data repair model, the method further includes:

[0030] The service domain responds to the execution request of the target service function and determines service requirement information based on the execution request of the target service function;

[0031] The system queries the data warehouse for target data that matches the service requirement information and executes the target service function based on the target data.

[0032] According to another aspect of this application, a data reproduction apparatus is provided, the apparatus comprising:

[0033] The image acquisition module is used to acquire image information generated based on raw data in a specific data domain.

[0034] The information recognition module is used to perform optical character recognition on the image information, and to extract structured data from the optical character recognition results through a preset formatted data extraction model to obtain the structured information corresponding to the image information.

[0035] The data reproduction module is used to reproduce the original data based on the structured information and the image information using a preset visual model, so as to obtain the reproduced data corresponding to the original data.

[0036] Optionally, the data reproduction module is further configured to:

[0037] The structured information and the image information are input into the preset visual model. The preset visual model performs image parsing on the image information with the structured information as a constraint, and outputs the image parsing result within the constraint range of the structured information as the reproduction data corresponding to the original data.

[0038] Optionally, the image acquisition module is further configured to:

[0039] Query the original data in the specific data domain according to the query conditions;

[0040] The original data is visualized and the corresponding image information is determined based on the visualized data. The visualized data includes the image information and / or video information. The video information is converted into the image information through frame extraction.

[0041] Optionally, the data reproduction module is further configured to:

[0042] The format of the reproduced data is validated according to the format conditions corresponding to the original data.

[0043] If the format verification fails, the reason for the failure is determined, and the original data is restored using the preset visual model based on the structured information, the image information, and the reason for the failure, until the reproduced data passes the format verification. The format conditions are obtained by acquiring the preset format conditions corresponding to the original data or by parsing the format of the image information.

[0044] Optionally, the data reproduction module is further configured to:

[0045] If the image information includes multiple images, the reproduction data corresponding to each image information is merged and deduplicated to obtain the reproduction data corresponding to the original data.

[0046] Optionally, the device further includes: a data repair module, used for:

[0047] The reproduced data is stored in a data warehouse, and a preset data repair model is invoked through a service call function pre-encapsulated in the data warehouse to repair the target reproduced data in the data warehouse;

[0048] Based on the repair results of the target reproduction data on the preset data repair model, the target reproduction data in the data warehouse is updated.

[0049] Optionally, the data repair module is further configured to:

[0050] By using a service call function pre-encapsulated in the data warehouse, the target reproducible data that matches the data repair field in the reproducible data is queried. The target reproducible data is then concatenated with a prompt word template. The concatenated data reproducible prompt word is used to call the preset data repair model to repair the target reproducible data. The data reproducible prompt word is used to guide the preset data repair model to repair the target reproducible data.

[0051] Optionally, the data repair module is further configured to:

[0052] The target reproduction data, permission verification information, and prompt word template are concatenated. The concatenated data reproduction prompt word is used to call the preset data repair model to repair the target reproduction data. This allows the preset data repair model to perform access permission verification on the permission verification information in the data reproduction prompt word, and to repair the target reproduction data after the access permission verification is passed.

[0053] Optionally, the data repair module is further configured to:

[0054] Receive data reproduction instructions;

[0055] The data reproduction instruction is parsed for repair fields and permission verification information. If a repair field is parsed, it is used as the data repair field; otherwise, a preset data repair field is used as the data repair field.

[0056] Optionally, the device further includes: a function execution module, used for:

[0057] The service domain responds to the execution request of the target service function and determines service requirement information based on the execution request of the target service function;

[0058] The system queries the data warehouse for target data that matches the service requirement information and executes the target service function based on the target data.

[0059] According to another aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the above-described data reproduction method.

[0060] According to another aspect of this application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described data reproduction method.

[0061] By employing the above technical solutions, the data reproduction method, apparatus, storage medium, and computer equipment provided in this application first convert the original data into image information; second, by leveraging OCR and a structured extraction model, key information in the image is extracted and expressed in a structured manner; finally, by introducing a visual understanding model, the ability to understand the image context is enhanced, improving the completeness and accuracy of data reproduction. This application breaks through the limitations of traditional data acquisition from images, improving acquisition efficiency and accuracy through OCR and a formatted data extraction model; by combining a visual model with structured and image information, the original data is reproduced completely and consistently, achieving automated data reproduction processing, improving the efficiency and accuracy of data reproduction, and reducing labor and time costs.

[0062] The technical solution of this invention can be applied to the transaction and delivery services of instant e-commerce platforms, such as flash sales, food delivery and retail.

[0063] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0064] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0065] Figure 1 A flowchart illustrating a data reproduction method provided in an embodiment of this application is shown;

[0066] Figure 2 A flowchart illustrating another data reproduction method provided in an embodiment of this application is shown;

[0067] Figure 3 A flowchart illustrating another data reproduction method provided in an embodiment of this application is shown;

[0068] Figure 4 A flowchart illustrating another data reproduction method provided in an embodiment of this application is shown;

[0069] Figure 5 A flowchart illustrating another data reproduction method provided in an embodiment of this application is shown;

[0070] Figure 6 A schematic diagram of the structure of a data reproduction device provided in an embodiment of this application is shown. Detailed Implementation

[0071] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0072] This embodiment provides a data reproduction method, such as Figure 1 As shown, the method includes:

[0073] Step 101: Obtain image information generated based on the original data in a specific data domain.

[0074] This application aims to address the challenge of efficiently and accurately restoring and utilizing original structured data when data isolation restrictions exist between specific data domains and service systems within an enterprise. By combining multi-dimensional technologies such as image processing, OCR recognition, structured information extraction, and visual model reconstruction, it achieves indirect recovery of data within a specific data domain without directly exposing sensitive data. The original data refers to structured data stored in a specific data domain within the enterprise, such as transaction records and user behavior logs. Due to various limitations, this data may not be directly exported in plaintext or transmitted to external service systems via API interfaces. Therefore, a feasible approach is to transform this original data into unstructured image information (such as PNG or JPG files) through visualization methods (e.g., data tables) and output it to an accessible service domain. Compared to traditional text anonymization methods, image formats are more difficult for programs to automatically parse, thus meeting data isolation requirements to a certain extent. Simultaneously, it provides basic materials for subsequent cross-domain data reproduction.

[0075] Step 102: Perform optical character recognition on the image information, and extract structured data from the optical character recognition results using a preset formatted data extraction model to obtain the structured information corresponding to the image information.

[0076] Next, OCR technology is used to recognize the text content in the image and convert it into a text string. Although traditional OCR has certain errors when dealing with complex fonts, blurry images, and messy layouts, it can still maintain a high recognition accuracy under controlled environments (such as standardized chart templates and fixed font styles). OCR outputs plain text, lacking the structured features of the original data (such as field names and table row and column relationships). To address this, a pre-trained formatted data extraction model is introduced (such as a structured information extraction model based on rule engines or deep learning, a large language model, a multimodal model, or a model obtained by knowledge transfer from a large model using distillation techniques, etc.) to extract data units with semantic structure (i.e., the structured information mentioned in this paper) from the text recognized by OCR, such as order number, amount, timestamp, and other field information.

[0077] Step 103: Using a preset visual model, reproduce the original data based on the structured information and the image information to obtain the reproduced data corresponding to the original data.

[0078] Finally, based on the obtained structured information, the original image information is further combined with a pre-defined visual model (such as a large visual model, a multimodal model, a convolutional neural network (CNN), a Transformer-based image understanding model, or a model obtained by knowledge transfer from a large model using distillation techniques, etc.) to model and analyze the image context. This model not only focuses on the text content itself but also understands visual semantic information such as chart types, table structures, and logical relationships in the image, thereby organizing the structured information extracted by the formatted data extraction model into a structured data format (such as JSON, XML, or database table structure). Ultimately, by fusing the structured text extracted by OCR with the contextual information understood by the visual model, "reproduced data" that closely resembles the original data form is reconstructed, achieving indirect restoration of the original data in a specific data domain.

[0079] By applying the technical solution of this embodiment, firstly, the raw data is transformed into image information; secondly, with the help of OCR and a structured extraction model, key information in the image is extracted and expressed in a structured manner; finally, by introducing a visual understanding model, the ability to understand the image context is enhanced, improving the completeness and accuracy of data reproduction. This embodiment breaks through the limitations of traditional data acquisition from images, improving acquisition efficiency and accuracy through OCR and a formatted data extraction model; by combining a visual model with structured and image information, the original data is reproduced completely and consistently, achieving automated data reproduction processing, improving the efficiency and accuracy of data reproduction, and reducing labor and time costs.

[0080] Furthermore, as a refinement and extension of the specific implementation methods of the above embodiments, and to fully illustrate the specific implementation process of this embodiment, another data reproduction method is provided, such as... Figure 2 As shown, the method includes:

[0081] Step 201: Query the original data in the specific data domain according to the query conditions; perform visualization data acquisition on the original data, and determine the image information corresponding to the original data based on the visualization data, wherein the visualization data includes the image information and / or video information, and convert the video information into the image information through frame extraction.

[0082] In this embodiment of the application, since the service system cannot directly access the original structured data (such as transaction records, user behavior logs, etc.) in a specific data domain, therefore, Figure 3As shown, the first step is to set query conditions within a specific data domain based on service requirements, and then retrieve the necessary raw data from the database of that specific data domain. This process is entirely completed within the specific data domain and does not involve data leakage. Furthermore, to achieve cross-domain data transmission, the retrieved raw structured data is converted into a visual form. Specifically, this may include: rendering the data into tables or other image formats; generating video content with dynamic display effects (such as carousel charts, animated demonstrations, etc.), and extracting frames from the video information to convert it into a set of static images for subsequent processing. For example, mobile devices can be used to record videos or take photos of table data within a specific data domain. If it is a video file, it can be extracted frame by frame and converted into image files. By presenting the raw data in the form of images or videos, the security risks of directly exporting data are avoided, and an operable unstructured carrier is provided for subsequent data reproduction.

[0083] Step 202: Perform optical character recognition on the image information, and extract structured data from the optical character recognition results using a preset formatted data extraction model to obtain the structured information corresponding to the image information.

[0084] In this embodiment, OCR technology is used to recognize text content in images and convert it into text information. OCR output text is typically an unstructured string, lacking the field structure and semantic information of the original data. Therefore, a pre-defined formatted data extraction model (e.g., a large language model trained on formatted data extraction) is used to perform semantic understanding and structured extraction of the OCR text. For example, fields such as "order number," "transaction amount," and "timestamp" are identified and organized into structured information; for instance, information related to the "order number" is grouped together, and information related to the "transaction amount" is grouped together. This process achieves the transformation from unstructured images to structured information, providing semantic constraints and a data foundation for subsequent visual model analysis.

[0085] Step 203: Input the structured information and the image information into the preset visual model, and use the preset visual model to perform image parsing on the image information with the structured information as a constraint. Output the image parsing result within the constraint range of the structured information as the reproduction data corresponding to the original data.

[0086] In this embodiment, a pre-defined visual model (e.g., a large visual model) receives raw image information and structured information as input. The structured information provides semantic guidance to the model, enabling it to more accurately understand visual semantic information such as table structures and field correspondences within the image. During the parsing process, the structured information serves not only as auxiliary information but also as a constraint on the visual model, thus avoiding the illusion problem inherent in large visual models. For example, the model can locate corresponding numerical regions in the image based on field names extracted by OCR and identify corresponding data points in charts. Finally, the visual model outputs the parsing result constrained by the structured information, i.e., the reproduced data. This reproduced data restores the true state of the original data as much as possible in terms of structure and content, possessing high accuracy and completeness.

[0087] Step 204: Perform format verification on the reproduced data according to the format conditions corresponding to the original data; if the format verification fails, determine the reason for the verification failure, and restore the original data again according to the structured information, the image information and the reason for the verification failure through the preset visual model until the reproduced data passes the format verification. The format conditions are obtained by obtaining the preset format conditions corresponding to the original data or by parsing the format of the image information.

[0088] In this embodiment, format conditions refer to the data format specifications that the original structured data should meet, such as field type (string, integer, date), field length, numerical range, and format expression (e.g., date format YYYY-MM-DD). Format conditions can be obtained in two ways: preset format conditions: directly obtaining the metadata or schema definition of the original data from a specific data domain; image format parsing: deducing the original data's structural form and field arrangement rules through layout analysis and table recognition of the image content. The reproduced data output by the visual model is compared one by one according to the preset format conditions to check whether it meets the original data's required format. For example, if the original data table has 5 rows and 10 columns, but the reproduced data does not, it is determined that the format verification fails. Another example is if a field value does not conform to the expected format (e.g., non-numeric characters appear in the amount field, or the time field has an incorrect format), it is determined that the format verification fails. The system records the specific reason for each verification failure, such as: "Field 'Order Number' contains illegal characters," "Field 'Transaction Time' has an incorrect format," "Numerical field 'Price' is outside the reasonable range," etc. After format verification fails, structured information, image information, and specific reasons for the failure can be input into the visual model to guide it to perform targeted re-analysis. The visual model, combining this feedback information, focuses on areas that may have been misjudged or missed, adjusts its parsing strategy, attempts to correct errors, and improves the accuracy of the reproduced data. This process can be iterated multiple times until the reproduced data fully meets the format requirements. This embodiment ensures the consistency of reproduced data quality through format verification and parsing retry. While traditional OCR + structured extraction methods can extract text information, they struggle to guarantee data integrity at the field level. The format verification step effectively identifies OCR misidentification problems caused by image blurring, font distortion, complex layout, etc., preventing erroneous data from flowing into downstream service systems. Furthermore, the feedback retry mechanism based on failure reasons enhances the system's fault tolerance and intelligence. The visual model no longer completes the parsing task in one go but can perform targeted optimization based on problems found in the previous round of results, gradually approaching the true data form. This closed-loop iterative mechanism significantly improves the success rate and accuracy of data reproduction.

[0089] Step 205: If the image information includes multiple images, the reproduction data corresponding to each image information is merged and deduplicated to obtain the reproduction data corresponding to the original data.

[0090] In this embodiment, in practical applications, the original data may not be fully presented by a single image due to its large volume or complex display format. For example, an image may only display one page of data, while the original data may contain multiple pages, video information may be extracted into multiple images after frame extraction, or data may be displayed in blocks or tables, such as multiple charts or multiple table screenshots. Therefore, the visualized data may contain multiple image information, each corresponding to a portion of the restored original data. Each image information is independently subjected to OCR recognition, structured extraction, visual model parsing, and format verification according to steps 202 to 204 to obtain its corresponding reproducible data fragment. Then, the reproducible data fragments corresponding to all image information are merged according to the logical structure of the original data. For example, paginated data may be merged according to page order; video frame data may be merged according to time or sequence order; and multiple table contents may be concatenated according to field or table structure. During the merging process, duplicate records or redundant fields may exist (such as repeated headers, titles, and static information displayed in multiple images). Therefore, a deduplication mechanism needs to be introduced: duplicate records are removed based on unique identifier fields (such as order numbers and user IDs); semantic comparison is performed on structured fields to identify and remove redundant content; contextual information is used to determine whether duplicate data is valid, avoiding accidental deletion. Ultimately, a logically consistent, structurally complete, and redundant set of reproducible data is output, restoring the overall form of the original data as much as possible. This ensures the accuracy and simplicity of the reproducible data. In multi-image scenarios, the repetitive nature of image display (such as repeated headers and static information display) can easily lead to redundant information in the recovered data. By introducing an intelligent deduplication strategy, the system can effectively identify and remove duplicate content, improving the usability of the reproducible data.

[0091] In the embodiments of this application, optionally, as shown... Figure 4 As shown, after obtaining the reproducible data corresponding to the original data, the method further includes:

[0092] Step 401: Store the reproduced data in a data warehouse, and use a pre-encapsulated service call function in the data warehouse to call a preset data repair model to repair the target reproduced data in the data warehouse.

[0093] Step 402: Update the target reproduction data in the data warehouse according to the repair results of the target reproduction data based on the preset data repair model.

[0094] In this embodiment, by introducing a data repair model invocation and a data reproduction update mechanism, centralized management, quality improvement, and continuous optimization of the reproduced data can be achieved, thus forming a closed-loop chain from data reproduction to data governance, enhancing the overall system's engineering availability and service support capabilities. Specifically, after completing the cross-domain recovery of the original data, the final reproduced data is uniformly written into a centralized data warehouse. Within the data warehouse or its integrated computing engine, service invocation functions are pre-encapsulated, which are responsible for triggering subsequent data repair logic. After the reproduced data is written to the data warehouse, the service invocation functions can be run automatically or as scheduled, thereby initiating a preset data repair model to analyze and repair specific target reproduced data. The data repair model can be a rule-based verification system, a statistical anomaly detection model, a machine learning prediction model, or a pre-trained large language model (or a small-scale model that has undergone knowledge distillation and knowledge transfer). For example, its main functions may include: verifying whether field values ​​conform to service logic (e.g., the amount cannot be negative); supplementing missing fields (e.g., inferring missing timestamps based on existing information); correcting numerical errors (e.g., identifying and correcting numerical misalignments caused by OCR recognition errors); standardizing units or formats (e.g., converting "1 million" to "1,000,000"); and converting non-standardized names to standardized names. After analyzing the target reproduced data, the data repair model outputs a repaired version of the data, such as more accurate field values, completed records, or corrected formats. These repair results will be fed back into the data warehouse and used to replace or supplement the original data, thereby forming more accurate, consistent, and usable data assets. For example, incremental updates or batch updates can be used to ensure that other service modules using the reproduced data are not affected; historical versions can be retained during the update process for easy audit tracking and problem backtracking; if the repair model outputs multiple candidate results, a manual review mechanism can be used for confirmation before updating. This application embodiment, through the combination of service call functions and the data repair model, enables the data warehouse to have the ability to automatically correct errors and govern data. Compared to traditional data reproduction solutions, the embodiments of this application can also proactively identify and correct logical errors, formatting issues, and semantic deviations in the data, thereby significantly improving the accuracy and applicability of the reproduced data.

[0095] Optionally, in step 401, the step of invoking a preset data repair model to repair the target reproducible data in the data warehouse using a service call function pre-encapsulated in the data warehouse includes: querying the target reproducible data that matches the data repair field in the reproducible data using the service call function pre-encapsulated in the data warehouse; concatenating the target reproducible data with a prompt word template; and using the concatenated data reproducible prompt word to invoke the preset data repair model to repair the target reproducible data. The data reproducible prompt word is used to guide the preset data repair model to repair the target reproducible data.

[0096] In this embodiment, such as Figure 5 As shown, the target reproduction data is concatenated with predefined prompt word templates to form data reproduction prompt words with clear semantic expression. This enables data repair models (such as large language models) to accurately understand the target and contextual logic of the current repair task based on natural language instructions, making the repair process closer to the real needs of the service scenario and improving the model's ability to understand and respond to complex error types. Through the prompt word mechanism, different fields, different data formats, and different industry repair needs can be flexibly adapted without modifying the model structure, enhancing the universality and scalability of the entire data repair process.

[0097] Optionally, in this embodiment, the step of concatenating the target reproduction data with the prompt word template and using the concatenated data reproduction prompt word to call the preset data repair model to repair the target reproduction data includes: concatenating the target reproduction data, permission verification information, and the prompt word template; using the concatenated data reproduction prompt word to call the preset data repair model to repair the target reproduction data; so that the preset data repair model performs access permission verification on the permission verification information in the data reproduction prompt word, and performs data repair on the target reproduction data after the access permission verification passes.

[0098] In this embodiment, permission verification information can be further introduced on the basis of the original prompt word mechanism. This allows the data repair model to not only focus on semantic content when processing reproduced data, but also to complete access control judgments on the call requests. This makes the entire data repair process more intelligent while also possessing higher security and controllability. Specifically, when generating data reproduction prompt words to guide the data repair model to execute tasks, in addition to including the target reproduced data itself, it is also concatenated with permission verification information into a preset prompt word template. The resulting complete prompt word not only includes the data content to be repaired, but also implicitly contains contextual information about whether the current request has the necessary access permissions. When the data repair model receives the prompt word, it first parses the permission verification information and verifies the caller's identity, role, or organization based on built-in logic or external authentication services. Only if it confirms that the caller has the corresponding access permissions will the model continue to perform repair operations on the target reproduced data; otherwise, it will refuse to respond or return an error message. The data repair model can pre-bind the access permission information of the subject with access permissions. When the model determines that the permission verification information belongs to the pre-bound access permission information, it continues to execute the subsequent data repair logic. Furthermore, a validity period for access permission information can be specified. Data repair logic will only proceed if the access permission information is matched within the valid timeframe. This embeds permission verification information into prompts, enabling fine-grained access control over the data repair process. Even when the data repair model is deployed in a shared environment, it ensures that different users or service modules can only access and repair data within their authorized scope, effectively preventing unauthorized access and accidental data manipulation.

[0099] Optionally, in this embodiment of the application, before calling a preset data repair model through a service call function pre-encapsulated in the data warehouse to repair the target reproducible data in the data warehouse, the method further includes: receiving a data reproducibility instruction; parsing the data reproducibility instruction for repair fields and permission verification information, wherein if a repair field is parsed out, the parsed repair field is used as the data repair field; otherwise, a preset data repair field is used as the data repair field.

[0100] In this embodiment, the execution of the service call function and the initiation of the data repair process can be triggered based on a data reproduction instruction. Specifically, it receives a data reproduction instruction from an external source (such as a server-side interface, a scheduled task, or manual operation). This instruction can be a structured request body (such as JSON format) or a task description in some form. The instruction is then parsed in two dimensions: the first dimension is the parsing of repair fields, attempting to extract the names of the fields specified by the user or task that need to be repaired, such as "amount" or "timestamp." If these fields are successfully parsed, they are used as data repair fields in the current repair task to limit the repair scope; if no repair fields are explicitly specified, the system automatically uses a preset set of default repair fields. This design supports flexible customization while ensuring usability in basic scenarios. The second dimension is the parsing of permission verification information, also extracting access control-related permission information from the instruction, such as user identity, organization affiliation, and role permission level. This information will be appended to the prompt when the data repair model is subsequently called, assisting the model in verifying the permissions of the current request and ensuring that the repair operation is executed within a reasonable scope. Through the above analysis process, the target fields and operation permissions of the task can be determined before data repair is performed, thereby achieving more refined task scheduling and security control.

[0101] It should be noted that, in this embodiment of the application, the service call function can also perform the following functions: perform length detection on the data reproduction prompt words to avoid the length of the data reproduction prompt words exceeding the input length limit of the data repair model, which would cause the call to fail; perform status detection on the return result of the data repair model, returning the result only when the call is successful and returning the corresponding failure status when the call fails; perform sensitive word detection on the return model of the data repair model to prevent sensitive words from being output to the repair data; and format the content returned by the data repair model to ensure that the data format is consistent.

[0102] Optionally, in this embodiment, the data warehouse is stored in a service domain; after updating the target reproduction data in the data warehouse based on the repair result of the target reproduction data according to the preset data repair model, the method further includes: the service domain responding to the execution request of the target service function, determining service requirement information based on the execution request of the target service function; querying the target data matching the service requirement information in the data warehouse, and executing the target service function based on the target data.

[0103] In this embodiment, the data reproduction and governance process is extended to the service execution stage, achieving a complete closed loop from data reproduction—data repair—data update—data application. By deploying a data warehouse in the service domain and dynamically querying and executing the required data based on service requests, the entire system not only possesses cross-domain data reproduction capabilities but can also directly support the operation of upper-layer service functions, enhancing the practical value and engineering implementation capability of this method in enterprise-level application scenarios. Specifically, the raw data in a specific data domain is restored and stored in the service domain's data warehouse in a high-quality, structured form through a preliminary process. Subsequently, when the service system receives an execution request for a target service function (such as generating reports, risk assessment, user profiling, transaction verification, etc.), it automatically enters the service response process. First, the context of the current service request is parsed to identify the data type, field range, service logic, etc., required by the service function, thereby generating service requirement information. This information may include key parameters such as user identifier, time range, service type, and field dependencies. Further, based on the service requirement information, data is retrieved in the service domain's data warehouse to find target data matching the current service request. These are high-quality structured data that have undergone cross-domain recovery, repair, and updates. Finally, the retrieved target data is used as input to drive the actual execution of the target service function. Overall, this not only breaks down data barriers between specific data domains and service domains, but also constructs a complete technical chain that starts from raw data in a data isolation environment, goes through intelligent recovery and quality governance, and ultimately supports the operation of the service system.

[0104] By applying the technical solution of this embodiment, the entire process uses unstructured carriers such as images or videos as the medium for cross-domain data transmission, effectively avoiding the security risks brought about by traditional direct data connection or API export methods. Simultaneously, by combining OCR recognition, structured information extraction, and visual modeling, the originally difficult-to-analyze image content is transformed into high-precision structured data, improving the completeness and accuracy of data reproduction. By introducing format verification and feedback retry mechanisms, data recognition errors caused by factors such as image quality, font style, or layout complexity can be automatically identified and corrected, thereby improving the robustness and adaptability of the data reproduction process. Furthermore, in scenarios involving multiple images or video frames, it has the ability to merge and deduplicate multiple restored segments. The introduction of a data warehouse not only provides a unified storage and management platform for reproduced data but also lays the foundation for subsequent data governance and service calls. Combined with preset service call functions and data repair models, intelligent repair operations can be performed on the data after it is entered into the warehouse, dynamically optimizing data quality and continuously updating the data warehouse content based on the repair results, ensuring that the server always accesses the latest and most accurate data version. Ultimately, by responding to service requests, parsing service requirements, querying and matching data, and driving the execution of service functions, a seamless connection from data reproduction to data services was achieved, breaking down data silos and allowing sensitive data that was originally limited to specific data domains to participate in the real-time operation and decision support of the online service system.

[0105] Furthermore, as Figure 1 To specifically implement the method, this application provides a data reproduction device, such as... Figure 6 As shown, the device includes:

[0106] The image acquisition module is used to acquire image information generated based on raw data in a specific data domain.

[0107] The information recognition module is used to perform optical character recognition on the image information, and to extract structured data from the optical character recognition results through a preset formatted data extraction model to obtain the structured information corresponding to the image information.

[0108] The data reproduction module is used to reproduce the original data based on the structured information and the image information using a preset visual model, so as to obtain the reproduced data corresponding to the original data.

[0109] Optionally, the data reproduction module is further configured to:

[0110] The structured information and the image information are input into the preset visual model. The preset visual model performs image parsing on the image information with the structured information as a constraint, and outputs the image parsing result within the constraint range of the structured information as the reproduction data corresponding to the original data.

[0111] Optionally, the image acquisition module is further configured to:

[0112] Query the original data in the specific data domain according to the query conditions;

[0113] The original data is visualized and the corresponding image information is determined based on the visualized data. The visualized data includes the image information and / or video information. The video information is converted into the image information through frame extraction.

[0114] Optionally, the data reproduction module is further configured to:

[0115] The format of the reproduced data is validated according to the format conditions corresponding to the original data.

[0116] If the format verification fails, the reason for the failure is determined, and the original data is restored using the preset visual model based on the structured information, the image information, and the reason for the failure, until the reproduced data passes the format verification. The format conditions are obtained by acquiring the preset format conditions corresponding to the original data or by parsing the format of the image information.

[0117] Optionally, the data reproduction module is further configured to:

[0118] If the image information includes multiple images, the reproduction data corresponding to each image information is merged and deduplicated to obtain the reproduction data corresponding to the original data.

[0119] Optionally, the device further includes: a data repair module, used for:

[0120] The reproduced data is stored in a data warehouse, and a preset data repair model is invoked through a service call function pre-encapsulated in the data warehouse to repair the target reproduced data in the data warehouse;

[0121] Based on the repair results of the target reproduction data on the preset data repair model, the target reproduction data in the data warehouse is updated.

[0122] Optionally, the data repair module is further configured to:

[0123] By using a service call function pre-encapsulated in the data warehouse, the target reproducible data that matches the data repair field in the reproducible data is queried. The target reproducible data is then concatenated with a prompt word template. The concatenated data reproducible prompt word is used to call the preset data repair model to repair the target reproducible data. The data reproducible prompt word is used to guide the preset data repair model to repair the target reproducible data.

[0124] Optionally, the data repair module is further configured to:

[0125] The target reproduction data, permission verification information, and prompt word template are concatenated. The concatenated data reproduction prompt word is used to call the preset data repair model to repair the target reproduction data. This allows the preset data repair model to perform access permission verification on the permission verification information in the data reproduction prompt word, and to repair the target reproduction data after the access permission verification is passed.

[0126] Optionally, the data repair module is further configured to:

[0127] Receive data reproduction instructions;

[0128] The data reproduction instruction is parsed for repair fields and permission verification information. If a repair field is parsed, it is used as the data repair field; otherwise, a preset data repair field is used as the data repair field.

[0129] Optionally, the device further includes: a function execution module, used for:

[0130] The service domain responds to the execution request of the target service function and determines service requirement information based on the execution request of the target service function;

[0131] The system queries the data warehouse for target data that matches the service requirement information and executes the target service function based on the target data.

[0132] It should be noted that other corresponding descriptions of the functional units involved in the data reproduction device provided in this application embodiment can be found by referring to... Figures 1 to 5 The corresponding descriptions in the method will not be repeated here.

[0133] This application also provides a computer device, specifically a personal computer, server, network device, etc. The computer device includes a bus, processor, memory, and communication interface, and may also include input / output interfaces and a display device. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores location information. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the various method embodiments.

[0134] Those skilled in the art will understand that the structure of the computer device described above is only a partial structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. A specific computer device may include more or fewer components, or combine certain components, or have different component arrangements.

[0135] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, having stored thereon a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0136] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0137] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0138] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, graphics processors, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0139] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0140] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A data reproduction method, characterized in that, The method includes: Retrieve image information generated from raw data in a specific data domain; Optical character recognition is performed on the image information, and structured data is extracted from the optical character recognition results using a preset formatted data extraction model to obtain the structured information corresponding to the image information. The structured information and the image information are input into a preset visual model. The preset visual model performs image parsing on the image information with the structured information as a constraint, and outputs the image parsing result within the constraint range of the structured information as the reproduction data corresponding to the original data.

2. The method according to claim 1, characterized in that, The process of obtaining image information generated from raw data in a specific data domain includes: Query the original data in the specific data domain according to the query conditions; The original data is visualized and the corresponding image information is determined based on the visualized data. The visualized data includes the image information and / or video information. The video information is converted into the image information through frame extraction.

3. The method according to claim 2, characterized in that, After obtaining the reproduction data corresponding to the original data, the method further includes: The format of the reproduced data is validated according to the format conditions corresponding to the original data. If the format verification fails, the reason for the failure is determined, and the original data is restored using the preset visual model based on the structured information, the image information, and the reason for the failure, until the reproduced data passes the format verification. The format conditions are obtained by acquiring the preset format conditions corresponding to the original data or by parsing the format of the image information.

4. The method according to claim 2, characterized in that, After obtaining the reproduction data corresponding to the original data, the method further includes: If the image information includes multiple images, the reproduction data corresponding to each image information is merged and deduplicated to obtain the reproduction data corresponding to the original data.

5. The method according to any one of claims 1 to 4, characterized in that, After obtaining the reproduction data corresponding to the original data, the method further includes: The reproduced data is stored in a data warehouse, and a preset data repair model is invoked through a service call function pre-encapsulated in the data warehouse to repair the target reproduced data in the data warehouse; Based on the repair results of the target reproduction data on the preset data repair model, the target reproduction data in the data warehouse is updated.

6. The method according to claim 5, characterized in that, The step of invoking a preset data repair model by calling a service call function pre-encapsulated in the data warehouse to repair the target reproduction data in the data warehouse includes: By using a service call function pre-encapsulated in the data warehouse, the target reproducible data that matches the data repair field in the reproducible data is queried. The target reproducible data is then concatenated with a prompt word template. The concatenated data reproducible prompt word is used to call the preset data repair model to repair the target reproducible data. The data reproducible prompt word is used to guide the preset data repair model to repair the target reproducible data.

7. The method according to claim 6, characterized in that, The step of concatenating the target reproduction data with the prompt word template, and using the concatenated data reproduction prompt word to call the preset data repair model to repair the target reproduction data includes: The target reproduction data, permission verification information, and prompt word template are concatenated. The concatenated data reproduction prompt word is used to call the preset data repair model to repair the target reproduction data. This allows the preset data repair model to perform access permission verification on the permission verification information in the data reproduction prompt word, and to repair the target reproduction data after the access permission verification is passed.

8. The method according to claim 7, characterized in that, Before invoking a preset data repair model through a service call function pre-encapsulated in the data warehouse to repair the target reproduction data in the data warehouse, the method further includes: Receive data reproduction instructions; The data reproduction instruction is parsed for repair fields and permission verification information. If a repair field is parsed, it is used as the data repair field; otherwise, a preset data repair field is used as the data repair field.

9. The method according to claim 5, characterized in that, The data warehouse is stored in a service domain; after updating the target reproduction data in the data warehouse based on the repair results of the target reproduction data according to the preset data repair model, the method further includes: The service domain responds to the execution request of the target service function and determines service requirement information based on the execution request of the target service function; The system queries the data warehouse for target data that matches the service requirement information and executes the target service function based on the target data.

10. A data reproduction device, characterized in that, The device includes: The image acquisition module is used to acquire image information generated based on raw data in a specific data domain. The information recognition module is used to perform optical character recognition on the image information, and to extract structured data from the optical character recognition results through a preset formatted data extraction model to obtain the structured information corresponding to the image information. The data reproduction module is used to input the structured information and the image information into a preset visual model, and to perform image parsing on the image information using the preset visual model with the structured information as constraints, and output the image parsing result within the constraints of the structured information as the reproduction data corresponding to the original data.

11. The apparatus according to claim 10, characterized in that, The image acquisition module is also used for: Query the original data in the specific data domain according to the query conditions; The original data is visualized and the corresponding image information is determined based on the visualized data. The visualized data includes the image information and / or video information. The video information is converted into the image information through frame extraction.

12. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 9.

13. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Receipt recognition method and device, equipment and storage medium

    CN113850175A