Data processing method and device, electronic equipment and computer readable storage medium

By combining optical character recognition with intelligent processing methods that utilize large language models and large visual models, the problems of low information extraction efficiency and insufficient accuracy are solved, achieving efficient and accurate character recognition.

CN122336788APending Publication Date: 2026-07-03PING AN INT FINANCIAL LEASING CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN INT FINANCIAL LEASING CO LTD
Filing Date
2026-05-09
Publication Date
2026-07-03

Smart Images

  • Figure CN122336788A_ABST
    Figure CN122336788A_ABST
Patent Text Reader

Abstract

This application relates to the field of data processing technology, applied in fintech and smart healthcare scenarios. It provides a data processing method, apparatus, electronic device, and computer-readable storage medium. The method includes: acquiring an image to be recognized; performing optical character recognition processing on the image to be recognized to obtain preliminary character recognition results; formatting the preliminary character recognition results to obtain formatted recognition information; performing field extraction processing on the formatted recognition information based on a preset large language model to obtain field extraction results; if the field extraction results indicate abnormal field extraction, performing recognition processing on the image to be recognized based on a preset large visual model to obtain character recognition results; and adjusting the character recognition results based on a preset strategy to obtain adjusted recognition results. Through the above technical solution, the efficiency and accuracy of information extraction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this application relate to, but are not limited to, the field of data processing technology, and are applied to financial technology and smart healthcare scenarios. In particular, they relate to a data processing method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] With continuous economic development, the financial and smart healthcare industries have also seen significant growth and expansion. In the financial sector, electronic bank acceptance bills are commonly used credit payment tools in commercial transactions, widely applied in trade settlements between enterprises. In scenarios such as financial leasing, factoring, and supply chain finance, staff need to extract key fields from electronic bank acceptance bills, including the bill number, issuance date, maturity date, drawer's full name, drawer's account number, bill amount, acceptor information, and endorsement details, for risk control review, contract verification, and business processing. In the smart healthcare industry, medical records can document and process various medical matters. Extracting information from these records can yield basic patient information, medical history and examination information, test results, and medication information. However, current information extraction primarily relies on manual verification, resulting in low efficiency. Summary of the Invention The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.

[0003] To address the problems mentioned in the background section, embodiments of this application provide a data processing method, apparatus, electronic device, and computer-readable storage medium that can improve the efficiency and accuracy of information extraction.

[0004] In a first aspect, embodiments of this application provide a data processing method, the data processing method comprising: Acquire the image to be recognized; The image to be identified is subjected to optical character recognition processing to obtain preliminary character recognition results; The preliminary character recognition results are formatted to obtain formatted recognition information; Based on a preset large language model, the formatted recognition information is processed by field extraction to obtain the field extraction result. In the event that the field extraction result indicates an abnormal field extraction, the image to be identified is processed based on a preset large visual model to obtain the character recognition result. The character recognition results are adjusted based on a preset strategy to obtain the adjusted recognition result.

[0005] Secondly, embodiments of this application also provide a data processing apparatus, the data processing apparatus comprising: The acquisition unit is used to acquire the image to be recognized; The first recognition unit is used to perform optical character recognition processing on the image to be recognized to obtain preliminary character recognition results; The conversion unit is used to format the preliminary character recognition result to obtain formatted recognition information. An extraction unit is used to perform field extraction processing on the formatted recognition information based on a preset large language model to obtain field extraction results. The second recognition unit is used to perform recognition processing on the image to be recognized based on a preset visual big model when the field extraction result indicates an abnormal field extraction, so as to obtain the character recognition result. The adjustment unit is used to adjust the character recognition result based on a preset strategy to obtain the recognition adjustment result.

[0006] Thirdly, embodiments of this application also provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the data processing method described in the first aspect above.

[0007] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions for performing the data processing method described in the first aspect above.

[0008] The data processing method according to the embodiments provided in this application has at least the following beneficial effects: First, an image to be recognized is acquired; then, optical character recognition processing is performed on the image to be recognized to obtain a preliminary character recognition result; next, the preliminary character recognition result is formatted to obtain formatted recognition information; then, field extraction processing is performed on the formatted recognition information based on a preset large language model to obtain a field extraction result; and if the field extraction result indicates an abnormal field extraction, the image to be recognized is processed based on a preset large visual model to obtain a character recognition result; finally, the character recognition result is adjusted based on a pre-set strategy to obtain a corresponding adjusted recognition result. Through the above technical solution, character recognition results can be intelligently identified from the image to be recognized, eliminating the need for manual review as in the past, thus improving the efficiency of information extraction; and during the recognition process, optical character recognition is used first, and if the recognition result obtained through optical character recognition is abnormal, a large visual model is used for recognition processing, thereby significantly improving the accuracy of information extraction. Attached Figure Description

[0009] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.

[0010] Figure 1 This is a schematic diagram of an application environment for a data processing method according to an embodiment of this application; Figure 2 This is a schematic flowchart of a data processing method provided in one embodiment of this application; Figure 3 yes Figure 2 A schematic diagram of a specific implementation of step S300; Figure 4 yes Figure 2 A schematic diagram of a specific implementation of step S400; Figure 5 Is it completed? Figure 2 A flowchart illustrating a specific implementation method following step S400; Figure 6 yes Figure 2 A flowchart illustrating the first specific implementation method of step S600; Figure 7 yes Figure 2 A flowchart illustrating the second specific implementation of step S600; Figure 8 yes Figure 2 A schematic diagram of the third specific implementation method of step S600; Figure 9 This is a schematic diagram of a data processing apparatus provided in one embodiment of this application; Figure 10 This is a schematic diagram of an electronic device provided in one embodiment of this application. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0012] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0013] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0014] AI is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. Artificial intelligence can simulate the information processes of human consciousness and thought. Furthermore, artificial intelligence utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results—the theories, methods, technologies, and application systems available for use.

[0015] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0016] Artificial intelligence, or AI, is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0017] The servers involved in artificial intelligence technology can be standalone servers or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0018] This application provides a data processing method, apparatus, electronic device, and computer-readable storage medium. In the data processing process, firstly, an image to be recognized is acquired; then, optical character recognition (OCR) is performed on the image to obtain a preliminary character recognition result; next, the preliminary character recognition result is formatted to obtain formatted recognition information; then, based on a preset large language model, field extraction processing is performed on the formatted recognition information to obtain field extraction results; and if the field extraction results indicate anomalies, a preset large visual model is used to process the image to be recognized to obtain character recognition results; finally, the character recognition results are adjusted based on a pre-set strategy to obtain corresponding adjusted recognition results. Through the above technical solution, character recognition results can be intelligently identified from the image to be recognized, eliminating the need for manual review as in the past, thus improving the efficiency of information extraction; and during the recognition process, optical character recognition is used first, and if the recognition result obtained through optical character recognition is abnormal, a large visual model is used for recognition processing, thereby significantly improving the accuracy of information extraction.

[0019] The data processing method provided in this application relates to the field of data processing technology. The data processing method provided in this application can be applied to a terminal or a server, and can also be software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0020] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0021] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0022] The data processing method provided in this application embodiment can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can acquire the image to be recognized from the client; then, it performs optical character recognition (OCR) on the image to obtain preliminary character recognition results; next, it formats the preliminary character recognition results to obtain formatted recognition information; then, based on a preset large language model, it performs field extraction processing on the formatted recognition information to obtain field extraction results; and if the field extraction results indicate anomalies, it performs recognition processing on the image to be recognized based on a preset large visual model to obtain character recognition results; finally, it adjusts the character recognition results based on a pre-set strategy to obtain corresponding adjusted recognition results. Through this technical solution, character recognition results can be intelligently identified from the image to be recognized, eliminating the need for manual review as in the past, thus improving the efficiency of information extraction. Furthermore, during the recognition process, it first uses OCR, and if the recognition results obtained through OCR are abnormal, it utilizes a large visual model for recognition processing, thereby significantly improving the accuracy of information extraction. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The embodiments of this application will be described in detail below through specific examples.

[0023] like Figure 2 As shown, Figure 2 This is a flowchart illustrating a data processing method provided in one embodiment of this application. The data processing method includes the following steps: Step S100: Obtain the image to be recognized.

[0024] The data processing method provided in this application first acquires an image to be identified during the data processing process. This image is the one from which information needs to be extracted. In the financial industry, the image to be identified can be an electronic bank acceptance bill, which includes fields such as bill number, issue date, maturity date, drawer's full name, drawer's account number, bill amount, acceptor information, and endorsement. Alternatively, the image to be identified can be a loan contract, which may include lender information, borrower information, contract number, loan type, loan amount, annual interest rate, repayment time, and default clauses. In the field of smart healthcare, the image to be identified can be a medical record, which may include basic patient information, basic medical records, physical examination information, diagnostic information, and prescription information.

[0025] It is worth noting that the image to be recognized can be obtained by the user taking a picture of the relevant invoice or contract, or by the user capturing a screenshot on their terminal device using screenshot software. The user can use a mobile phone, camera, or tablet to take the picture of the invoice or contract; there are no restrictions on this.

[0026] It is worth noting that during the acquisition of the image to be identified, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is always obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when this application embodiment needs to obtain sensitive personal information of the user, separate permission or consent from the user is obtained through pop-ups or redirection to a confirmation page. Only after explicitly obtaining the user's separate permission or consent is the necessary user-related data for the normal operation of this application embodiment acquired.

[0027] Step S200: Perform optical character recognition processing on the image to be recognized to obtain preliminary character recognition results.

[0028] The data processing method provided in this application first acquires the image to be recognized during the data processing process, and then performs optical character recognition processing on the image to be recognized to obtain preliminary character recognition results, in order to prepare for subsequent character recognition verification.

[0029] It is worth noting that in the process of optical character recognition processing of the image to be recognized, the image to be recognized is first preprocessed to obtain a preprocessed image. The preprocessing may include image enhancement processing, geometric correction processing, and layout analysis processing. After obtaining the preprocessed image, text detection processing can be performed on the preprocessed image to determine the coordinate frame of the text region. Then, text recognition processing is performed on the coordinate frame of the text region to obtain the corresponding preliminary character recognition results.

[0030] For example, in the financial industry, after obtaining an electronic bank acceptance bill, image enhancement, geometric correction, and layout analysis can be performed on the bill to obtain a preprocessed image. Then, text detection is performed on the preprocessed image to determine the coordinate frames of the text regions. Finally, character recognition processing is performed on the coordinate frames to obtain preliminary character recognition results corresponding to the electronic bank acceptance bill. Alternatively, in the field of smart healthcare, after obtaining an electronic medical document, image enhancement, geometric correction, and layout analysis can be performed on the document to obtain a preprocessed image. Then, text detection is performed on the preprocessed image to determine the coordinate frames of the text regions. Finally, character recognition processing is performed on the coordinate frames to obtain preliminary character recognition results corresponding to the electronic medical document.

[0031] Step S300: Format the preliminary character recognition results to obtain formatted recognition information.

[0032] The data processing method provided in this application embodiment, during the data processing process, after obtaining the preliminary character recognition result by optical character recognition processing of the image to be recognized, can perform formatting processing on the preliminary character recognition result to obtain formatted recognition information, in order to prepare for subsequent field extraction.

[0033] It is worth noting that in the process of formatting the preliminary character recognition results to obtain formatted recognition information, multiple text contents and corresponding coordinate information are first determined from the preliminary character recognition results; then, the multiple text contents are sorted according to the coordinate information; finally, the sorted text contents and corresponding coordinate information are structured and merged to obtain formatted recognition information, which can then be prepared for subsequent character recognition verification processing.

[0034] It is worth noting that formatting the preliminary character recognition results can yield formatted recognition information. The preliminary character recognition results contain text content and coordinate information. Therefore, the text content and coordinate information can be determined from the preliminary character recognition results first. Then, multiple text contents can be sorted according to the coordinate information. Finally, the sorted text contents and corresponding coordinate information are structured and merged to obtain the corresponding formatted recognition information.

[0035] like Figure 3 As shown, formatting the initial character recognition results to obtain formatted recognition information can include the following steps: Step S310: Determine multiple text contents and corresponding coordinate information from the preliminary character recognition results; Step S320: Sort multiple text contents according to coordinate information; Step S330: Perform structured merging processing on the sorted text content and corresponding coordinate information to obtain formatted recognition information.

[0036] For steps S310 to S330, in the process of formatting the preliminary character recognition results to obtain formatted recognition information, firstly, multiple text contents and corresponding coordinate information are determined from the preliminary character recognition results; then, the multiple text contents are sorted according to the coordinate information; finally, the sorted text contents and corresponding coordinate information are structured and merged to obtain formatted recognition information, which can then prepare for subsequent character recognition verification processing.

[0037] It is worth noting that the preliminary character recognition result obtained after optical character recognition (OCR) processing can contain multiple text contents and coordinate information. Subsequently, the multiple text contents can be sorted based on the coordinate information, for example, sorted by the Y-coordinate. Finally, the sorted text contents and corresponding coordinate information are structurally merged to obtain the corresponding formatted recognition information, preparing for subsequent field extraction. For example, in the financial industry, after obtaining the preliminary character recognition result from OCR on electronic bank acceptance bills, multiple text contents can be determined from the preliminary character recognition result, such as the bill number, drawer's account number, issue date, and bill amount, as well as the corresponding coordinate information. Then, the multiple text contents are sorted based on the Y-coordinate information. Finally, the sorted text contents and corresponding coordinate information are structurally merged to obtain the corresponding formatted recognition information. Alternatively, in the smart healthcare industry, after obtaining preliminary character recognition results from optical character recognition (OCR) on hospital records and receipts, multiple text contents can be determined from these preliminary results, such as receipt number, inpatient's name, discharge date, and treatment costs, as well as corresponding coordinate information. Then, the multiple text contents are sorted according to the Y-coordinate information. Finally, the sorted text contents and corresponding coordinate information are structured and merged to obtain the corresponding formatted recognition information.

[0038] Step S400: Based on the preset large language model, perform field extraction processing on the formatted recognition information to obtain the field extraction results.

[0039] The data processing method provided in this application, after formatting the preliminary character recognition results to obtain formatted recognition information, can perform field extraction processing on the formatted recognition information based on a pre-set large language model to obtain the corresponding field extraction results; subsequently, it can also determine whether the field extraction process is normal based on the field extraction results, so as to determine whether it is necessary to trigger the visual large model to recognize the image to be recognized, thereby greatly improving the reliability of field extraction.

[0040] It is worth noting that in the process of extracting fields from formatted recognition information based on a pre-defined large language model, the text content and coordinate information are first determined from the formatted recognition information. Then, the large language model uses the spatial relationship between the text content and coordinate information to perform field matching, thereby obtaining the corresponding field names and values. Finally, the field names and values ​​are determined as the field extraction results. Subsequent evaluation of the field extraction results can then confirm whether the previous field extraction was successful.

[0041] like Figure 4As shown, the field extraction process for formatted recognition information based on a preset large language model to obtain the field extraction results may include the following steps: Step S410: Determine the text content and coordinate information from the formatted recognition information; Step S420: Using the large language model, the spatial relationship between text content and coordinate information is used to perform field matching processing to obtain field names and corresponding field values. Step S430: Determine the field name and field value as the field extraction result.

[0042] For steps S410 to S430, in the process of extracting fields from the formatted recognition information based on a pre-defined large language model, the text content and coordinate information are first determined from the formatted recognition information. Then, the large language model is used to perform field matching based on the spatial relationship between the text content and coordinate information, thereby obtaining the corresponding field names and values. Finally, the field names and values ​​are determined as the field extraction results. Subsequently, the accuracy of the previous field extraction can be determined based on the field extraction results, thus ensuring the accuracy of the field extraction.

[0043] It is worth noting that the process of field matching using a large language model first requires extracting the text content and its corresponding coordinate information (such as the coordinates of the top left and bottom right corners of the bounding box) from the scanned document or image using OCR technology. The text content may represent field names (such as "name" or "date") or field values ​​(such as "Zhang San" or "2024-01-01"). The spatial relationship between them (such as left-right adjacency or top-bottom alignment) is key to determining the matching relationship. Next, the text content and coordinate information corresponding to each text content are jointly encoded into an input vector. The text content is converted into semantic embedding through a word embedding layer, while the coordinate information is converted into spatial location embedding through a dedicated coordinate encoder (such as a linear layer or positional encoding). Then, the two are concatenated or added together to form a vector. The system generates a representation that integrates semantic and spatial features. Then, the resulting vectors are ordered spatially according to their text content within the document to form a sequence, which is input into the Transformer module of the large language model. Inside the Transformer, a self-attention mechanism calculates the attention weights between any two text contents, thereby capturing their semantic relevance and spatial proximity. For example, the field name "Name" might have a high attention score with the field value "Zhang San" on the right. Through multiple stacked attention layers and feedforward networks, the large language model can learn complex correspondence patterns between field names and field values. Finally, the output layer of the large language model can directly generate structured field-value pairs (e.g., output in JSON format) through the decoder.

[0044] It is worth noting that when the field name and field value match each other, or when the matching threshold between the two is greater than the preset matching threshold, it can be determined that there was no abnormality in the previous field extraction process, and there is no need to trigger the visual big model to perform character recognition processing.

[0045] Step S500: In the case of anomalies in the field extraction result, the image to be recognized is processed based on the preset visual large model to obtain the character recognition result.

[0046] The data processing method provided in this application, after performing field extraction processing on formatted recognition information based on a preset large language model to obtain field extraction results, can perform recognition processing on the image to be recognized based on a preset large visual model when the field extraction results indicate field extraction anomalies, thereby obtaining the corresponding character recognition results. Through the above technical solution, character recognition in images becomes more reliable, improving the success rate of character recognition.

[0047] For example, in the financial industry, optical character recognition (OCR) is used to obtain preliminary character recognition results on electronic bank acceptance bills. However, if the preliminary character recognition results indicate anomalies in the previously extracted fields, a pre-set visual model will be triggered to reprocess the electronic bank acceptance bills for character recognition, thus avoiding the problem of low character recognition success rates caused by relying on a single character recognition method. Similarly, in the smart healthcare industry, OCR is used to obtain preliminary character recognition results on outpatient registration records. If the preliminary character recognition results indicate anomalies in the previously extracted fields, a pre-set visual model will be triggered to reprocess the outpatient registration records for character recognition, similarly avoiding the problem of low character recognition success rates caused by relying on a single character recognition method.

[0048] It is worth noting that the character recognition process based on a large visual model first involves adjusting the image to a uniform size and normalizing pixel values ​​through a preprocessing module. Next, the image is segmented into fixed-size image patches. Each patch is transformed into an initial visual embedding through a linear projection layer and added to a learnable positional encoding to preserve spatial order information. The embedding sequence is then input into a core vision module consisting of multiple Transformer encoder layers. Each layer captures long-distance dependencies between image patches using a multi-head self-attention mechanism and performs nonlinear transformations using a feedforward network to extract deep visual features rich in contextual semantics. The feature map or feature sequence output by the encoder is fed into a dedicated text decoding module, which generates character probability distributions position-by-position in an autoregressive manner while dynamically focusing on relevant regions in the image using an attention mechanism. Finally, the decoded features are mapped to a predefined character category space through classification to obtain the final character recognition result.

[0049] Step S600: Adjust the character recognition results based on the preset strategy to obtain the adjusted recognition results.

[0050] The data processing method provided in this application, when the field extraction result indicates an anomaly in the field extraction, after obtaining the character recognition result by recognizing the image based on a preset large visual model, can adjust the character recognition result based on a pre-set strategy. Through the above technical solution, the character recognition result can be adjusted and simplified to better meet the user's viewing and review needs.

[0051] like Figure 5 As shown, after performing field extraction processing on the formatted recognition information based on a preset large language model and obtaining the field extraction results, the following steps may also be included: Step S440: If the field extraction result indicates that the field extraction is normal, adjust the field extraction result based on a preset strategy to obtain the identification adjustment result.

[0052] The data processing method provided in this application embodiment performs field extraction processing on formatted recognition information based on a preset large language model. After obtaining the field extraction result, when the field extraction result indicates that the field extraction is normal, the field extraction result can be directly adjusted based on a preset strategy without going through the subsequent visual large model for character recognition processing. This greatly speeds up the efficiency of character recognition and makes the character recognition more reasonable.

[0053] For example, in the financial industry, if the field extraction results from a loan contract indicate normal field extraction, the results can be directly adjusted based on a preset strategy, eliminating the need for subsequent character recognition processing using a large visual model, thus improving character recognition efficiency. Similarly, in the smart healthcare industry, if the field extraction results from inpatient visit details indicate normal field extraction, the results can also be directly adjusted based on a preset strategy, eliminating the need for further character recognition processing using a large visual model, thereby improving character recognition efficiency as well.

[0054] like Figure 6 As shown, adjusting the character recognition results based on a preset strategy to obtain the adjusted recognition result can include the following steps: Step S610: Determine the primary key information from the character recognition results, wherein each primary key information corresponds to associated information on each page; Step S620: Based on the primary key information, perform information merging processing on the related information in each page to obtain the first adjustment result; Step S630: Use the first adjustment result as the identification adjustment result.

[0055] For steps S610 to S630, in the process of adjusting the character recognition results based on a pre-set strategy to obtain the adjusted recognition result, firstly, primary key information is determined from the character recognition results, where each primary key information corresponds to associated information on each page; then, the associated information on each page is merged according to the primary key information to obtain a first adjustment result; finally, the first adjustment result is used as the adjusted recognition result. This technical solution effectively simplifies the information in the character recognition results, facilitating subsequent viewing and review by the user.

[0056] For example, in the financial industry, for electronic bank acceptance bill business, the method for PDF to image conversion and multi-page concurrent processing uses the PyMuPDF library to convert each page of the PDF file into a 2x scaled high-definition PNG image. Then, a thread pool is used to perform concurrent recognition processing on all pages, with the concurrency level controlled by the MAX_WORKERS_PAGE parameter (default is 3). Each thread independently executes the dual-scheme adaptive recognition process. Parameters such as channel ID and scene ID are passed through method parameters rather than instance variables to ensure multi-threaded concurrency safety. After processing, the recognition results of each page are collected in page number order. In addition, cross-page information is merged by bill number. For cases where the information of the same bill is distributed across multiple pages, this method uses the bill number as the primary key to merge multiple records with the same bill number into a complete record. During merging, a "non-null value priority" strategy is used to supplement missing fields, and endorsement information is deduplicated and merged by endorser name to avoid duplication. For cross-page endorsement date completion, this method addresses situations where endorsement records are split across pages (e.g., the previous page has the endorser's name but no date, and the next page only begins with a date). It identifies isolated records with only a date and no endorser, and missing records with an endorser but no date. Based on positional distance, it performs optimal matching and completion, then removes the used isolated date records to achieve complete restoration of cross-page endorsement information. For endorsement information merging based on key fields, in addition to merging by invoice number, this method also performs secondary merging using a combination of three key fields: "invoice number + sub-interval start + sub-interval end". When one page contains endorsement information while another page lacks endorsement information for the same key field, the endorsement information is merged into the missing page, and other potentially missing fields are added, achieving a union of the two pages' information.

[0057] like Figure 7 As shown, adjusting the character recognition results based on a preset strategy to obtain the adjusted recognition result can include the following steps: Step S640: Filter out empty records from the character recognition results; Step S650: Delete empty records in the character recognition results to obtain the second adjustment result; Step S660: Use the second adjustment result as the identification adjustment result.

[0058] For steps S640 to S660, in the process of adjusting the character recognition results based on a preset strategy to obtain the adjusted recognition result, firstly, empty records are filtered out from the character recognition results; then, the empty records in the character recognition results are deleted to obtain the second adjusted result; finally, the second adjusted result is used as the adjusted recognition result. Through the above technical solution, empty records in the character recognition results can be effectively removed, and the character recognition results can be effectively optimized.

[0059] like Figure 8 As shown, adjusting the character recognition results based on a preset strategy to obtain the adjusted recognition result can include the following steps: Step S670: Select multiple subsets of information from the character recognition results; Step S680: Deduplication is performed on the identical subset information in the character recognition results to obtain the third adjustment result; Step S690: Use the third adjustment result as the identification adjustment result.

[0060] For steps S670 to S690, in the process of adjusting the character recognition results based on a preset strategy to obtain the adjusted recognition result, firstly, multiple subset information is selected from the character recognition results; then, duplicate information in the same subset information in the character recognition results is deduplicated to obtain a third adjusted result; finally, the third adjusted result can be used as the adjusted recognition result. Through the above technical solution, the character recognition results can be effectively optimized, the subset information can be merged and simplified, and preparations can be made for subsequent review and approval of the character recognition results.

[0061] For example, in the financial industry, a multi-level data cleaning and subset bill deduplication method was designed for electronic bank acceptance bill business to ensure the accuracy and uniqueness of the final output. For empty record cleaning and merging, this method classifies and judges each record in the identification results. Completely empty records (all fields, including endorsement information, are empty) are directly deleted; records with only endorsement information (main fields are empty but there is a valid endorser) are attempted to be merged into the previous record; records with only endorsement date (endorser is empty but date is not empty) are attempted to be added to the previous record for endorsement records lacking a date. For subset bill deduplication, this method judges the subset relationship between any two bill records; the subset relationship of values ​​is defined as follows: an empty value is a subset of any value, and a non-empty value A is a subset of B if and only if A equals B. The subset relationship of endorsement information is defined as follows: the endorser's name for every endorsement record of A can be found in B; the subset relationship of a bill is defined as follows: all major fields and endorsement information of A are subsets of B; when bill A is a proper subset of bill B (A is a subset of B but B is not a subset of A), delete A with less information and retain B with more complete information; when A and B are completely equivalent, retain the one with more fields or the one with a smaller index to avoid duplication.

[0062] In addition, for the complex structure of the bank acceptance bill number, an intelligent bill number parsing and sub-interval extraction method is designed. For the automatic extraction of sub-intervals in the bill number, for bills without sub-interval start, sub-interval end field names, directly identifying the values of the sub-interval start and sub-interval end fields has poor effects; since the position of the value of the sub-interval is below the value of the bill number, the value of the bill sub-interval is identified together as part of the bill number value. For example, for "531330551683920251029002198194 1006000001-1122567529", "531330551683920251029002198194" is the bill number, and "1006000001-1122567529" is the bill sub-interval, which often contains a sub-interval digital pattern connected by "-" or ",". The first part is the sub-interval start field, and the second part is the sub-interval end field; in the post-processing of this method, the sub-interval pattern in the bill number is detected by regular expression matching, and it is automatically separated into two fields: the pure bill number and the sub-interval, and the sub-interval start and sub-interval end fields are further split, and at the same time, the Chinese comma is unified as the English comma. For the priority processing of the sub-interval, if the large model directly recognizes the value of the bill sub-interval field, this value is used preferentially; otherwise, the sub-interval value automatically extracted from the bill number is used. The sub-interval string supports the parsing of two delimiters, "-" and ",", and the special processing of the "0" value (indicating that both the sub-interval start and end are "0"). For the cleaning of the bill number, blank characters such as spaces, tab characters, and line break characters introduced in the OCR recognition process are removed to obtain a standard bill number consisting of pure numbers.

[0063] In addition, the embodiments of the present disclosure also provide an algorithm for recognizing the amount of the bill described in Chinese capital letters and converting it into an Arabic numeral amount, and a robust intelligent conversion algorithm for Chinese capital amounts is adopted. The embodiments of the present disclosure adopt the form of recognizing the amount of the bill described in Chinese capital letters and converting it into an Arabic numeral amount, and at the same time implement a Chinese capital amount conversion algorithm based on integer operations, avoiding the problem of floating-point precision loss and significantly improving the accuracy of bill amount recognition; this algorithm supports two sets of character systems, capital numbers (zero, one, two, three, four, five, six, seven, eight, nine) and lowercase numbers (〇, one, two, three, four, five, six, seven, eight, nine), and supports all measurement units and their aliases such as "yuan / circle", "shi / ten", "bai / hundred", "qian / thousand", "wan / ten thousand", "yi / hundred million", "jiao / hair", "fen". Before conversion, irrelevant characters such as "Renminbi", "capital", parentheses, "¥" are automatically cleared, and suffixes such as "whole" and "positive" are removed. The core algorithm performs integer operations with "fen" as the smallest measurement unit, accumulates level by level according to the hierarchical structure of "hundred million → ten thousand → yuan", and finally divides by 100 to obtain the accurate yuan value, which is formatted as a string with two decimal places and output, completely solving the problem of floating-point precision loss.

[0064] The above technical solutions significantly improve recognition accuracy and robustness. Through a dual-scheme adaptive switching architecture, this method automatically and seamlessly switches to Scheme Two when Scheme One fails. Combined with multiple retry mechanisms within each scheme, this greatly enhances the success rate and robustness of recognition. Scheme One utilizes the precise text recognition capabilities of OCR and the semantic understanding and spatial reasoning capabilities of a large language model to achieve high-precision field location and extraction. Scheme Two leverages the end-to-end recognition capabilities of a large visual model to provide a reliable backup when Scheme One fails. The two schemes complement each other, effectively addressing real-world challenges such as diverse document formats and inconsistent image quality. Furthermore, by using a thread pool to concurrently process multi-page PDF documents, the original sequential page-by-page processing flow is transformed into parallel processing, significantly reducing processing time. Simultaneously, during concurrent processing, context information is passed via parameters rather than instance variables, ensuring data security and result accuracy in a multi-threaded environment. Furthermore, the multi-level cross-page merging mechanism (merging by bill number → merging by key field → cross-page endorsement date completion → empty record cleanup → subset deduplication) ensures that the information of the same bill of exchange distributed across multiple pages can be completely and accurately merged into a single record. In particular, the intelligent completion of cross-page endorsement records solves the problem of pagination of endorsers and endorsement dates, which traditional solutions cannot handle, significantly improving the completeness of endorsement information recognition. Moreover, the subset bill deduplication method effectively eliminates redundant records caused by repeated recognition across multiple pages; and multi-level data cleaning (null value standardization, invalid value filtering, English parentheses to Chinese parentheses, etc.) ensures a unified and standardized output data format. The conversion of uppercase amounts based on integer arithmetic avoids the problem of floating-point precision loss, ensuring absolute accuracy in amount conversion and significantly improving the accuracy of amount field recognition.

[0065] In addition, such as Figure 9 As shown, one embodiment of this application also provides a data processing apparatus 10, which includes: Acquisition unit 100 is used to acquire the image to be recognized; The first recognition unit 200 is used to perform optical character recognition processing on the image to be recognized to obtain preliminary character recognition results; The conversion unit 300 is used to format the preliminary character recognition results to obtain formatted recognition information. Extraction unit 400 is used to perform field extraction processing on formatted recognition information based on a preset large language model to obtain field extraction results; The second recognition unit 500 is used to perform recognition processing on the image to be recognized based on a preset visual large model when the field extraction result indicates an abnormal field extraction, so as to obtain the character recognition result. The adjustment unit 600 is used to adjust the character recognition result based on a preset strategy to obtain the recognition adjustment result.

[0066] The specific implementation of the data processing device 10 is basically the same as the specific embodiment of the data processing method described above, and will not be repeated here.

[0067] In addition, such as Figure 10 As shown, one embodiment of this application also provides an electronic device 700, which includes: a memory 720, a processor 710, and a computer program stored on the memory 720 and executable on the processor 710.

[0068] The processor 710 and memory 720 can be connected via a bus or other means.

[0069] The non-transient software program and instructions required to implement the data processing method of the above embodiments are stored in the memory 720. When executed by the processor 710, the data processing method of each of the above embodiments is executed.

[0070] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0071] Furthermore, one embodiment of this application provides a computer-readable storage medium storing computer-executable instructions that are executed by a processor 710 or a controller, for example, by a processor 710 in the above-described device embodiment, causing the processor 710 to perform the data processing method in the above-described embodiment.

[0072] The above embodiments can be used in combination, and modules with the same name in different embodiments may be the same or different.

[0073] The foregoing has described specific embodiments of this application; other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than those shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily have to follow the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0074] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and computer-readable storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0075] The apparatus, device, computer-readable storage medium and method provided in the embodiments of this application are corresponding to each other. Therefore, the apparatus, device and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, device and computer storage medium will not be described again here.

[0076] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Moreover, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used when writing program development code. The original code before compilation must also be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using the aforementioned hardware description languages ​​and programming it into an integrated circuit, the hardware circuit that implements the logic method flow can be easily obtained.

[0077] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0078] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0079] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing the embodiments of this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0080] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0081] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0082] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0083] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0084] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0085] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (FlashRAM). Memory is an example of computer-readable media.

[0086] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0087] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0088] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, A and B simultaneously, or B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0089] The embodiments of this application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. The embodiments of this application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.

[0090] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0091] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

[0092] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.

Claims

1. A data processing method, characterized in that, The data processing method includes: Acquire the image to be recognized; The image to be identified is subjected to optical character recognition processing to obtain preliminary character recognition results; The preliminary character recognition results are formatted to obtain formatted recognition information; Based on a preset large language model, the formatted recognition information is processed by field extraction to obtain the field extraction result. In the event that the field extraction result indicates an abnormal field extraction, the image to be identified is processed based on a preset large visual model to obtain the character recognition result. The character recognition results are adjusted based on a preset strategy to obtain the adjusted recognition result.

2. The data processing method according to claim 1, characterized in that, The initial character recognition result is formatted to obtain formatted recognition information, including: Multiple text contents and corresponding coordinate information are determined from the preliminary character recognition results; Based on the coordinate information, the multiple text contents are sorted. The sorted text content and the corresponding coordinate information are structurally merged to obtain the formatted recognition information.

3. The data processing method according to claim 1, characterized in that, The formatted recognition information is processed by extracting fields based on a preset large language model to obtain field extraction results, including: The text content and the coordinate information are determined from the formatted recognition information; Using the large language model, field matching is performed based on the spatial relationship between the text content and the coordinate information to obtain field names and corresponding field values. The field name and the field value are determined as the field extraction result.

4. The data processing method according to claim 1, characterized in that, After performing field extraction processing on the formatted recognition information based on a preset large language model to obtain the field extraction results, the method further includes: If the field extraction result indicates that the field extraction is normal, the field extraction result is adjusted based on a preset strategy to obtain the identification adjustment result.

5. The data processing method according to claim 1, characterized in that, The adjustment process based on a preset strategy to obtain the character recognition result includes: Primary key information is determined from the character recognition results, wherein each primary key information corresponds to associated information on each page; Based on the primary key information, the associated information in each page is merged to obtain the first adjustment result; The first adjustment result is used as the identification adjustment result.

6. The data processing method according to claim 1, characterized in that, The adjustment process based on a preset strategy to obtain the character recognition result includes: Empty records are filtered out from the character recognition results; The empty records in the character recognition results are deleted to obtain the second adjustment result; The second adjustment result is taken as the identification adjustment result.

7. The data processing method according to claim 1, characterized in that, The adjustment process based on a preset strategy to obtain the character recognition result includes: Multiple subsets of information are selected from the character recognition results; The identical subset information in the character recognition results is deduplicated to obtain a third adjustment result; The third adjustment result is used as the identification adjustment result.

8. A data processing apparatus, characterized in that, The data processing device includes: The acquisition unit is used to acquire the image to be recognized; The first recognition unit is used to perform optical character recognition processing on the image to be recognized to obtain preliminary character recognition results; A conversion unit is used to format the preliminary character recognition result to obtain formatted recognition information; An extraction unit is used to perform field extraction processing on the formatted recognition information based on a preset large language model to obtain field extraction results. The second recognition unit is used to perform recognition processing on the image to be recognized based on a preset visual big model when the field extraction result indicates an abnormal field extraction, so as to obtain the character recognition result. The adjustment unit is used to adjust the character recognition result based on a preset strategy to obtain the recognition adjustment result.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the data processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the data processing method according to any one of claims 1 to 7.