Method and System for Consistency Calibration and Standardized Output of AI Model Recognition Results

By standardizing the data format of yellow page number and using AI models for intelligent correction, the problem of complex format processing and time-consuming data sorting in yellow page number recognition is solved, and efficient, consistent and accurate output of number recognition results is achieved.

CN119917647BActive Publication Date: 2025-08-05BEIJING YULORE INNOVATION TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510405635.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-08-05
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The existing yellow page number identification technology has a lot of code to process format problems, and data sorting is time-consuming and lacks a data garbled pre-processing mechanism, which makes it difficult to guarantee the consistency and accuracy of the model output results.

Method used

By receiving yellow page number data from the preset format collection, using the format recognition algorithm and the input format unified algorithm to normalize the key fields, using the prompt word template construction algorithm to generate structured AI model input instructions, and using a fine-tuned general AI model to identify and verify, and initially standardize it according to the fixed field output format rules, intelligently correct it in combination with common error mapping tables and error case knowledge bases, and finally generate consistency calibration and standardized number recognition results.

Benefits of technology

It significantly improves the accuracy and consistency of yellow page number identification, reduces the workload of data operation and maintenance and developers, and realizes the system's continuous self-optimization and long-term identification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917647B_ABST
    Figure CN119917647B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for consistency calibration and standardized output of AI model recognition results. The method includes: obtaining yellow page number data in a standardized format; using an algorithm for constructing prompts stored to generate input instructions for a structured AI model, performing yellow page number data recognition and verification processing through a fine-tuned general AI large model, and converting the processing results into preliminary standardized AI recognition results that conform to preset specifications according to fixed field output format rules; according to the preliminary standardized AI recognition results, performing intelligent correction on the recognition results by matching with a common error mapping table and verifying with an error case knowledge base, calculating the correct rate index of the current batch, and outputting calibrated recognition results with confidence scores, as well as generating final consistent calibration and standardized number recognition results by using standardized output format conversion. The present invention effectively improves the accuracy and consistency of yellow page number recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method and system for consistency calibration and standardized output of AI model recognition results. Background Art

[0002] The yellow page number recognition technology is an important link in digitizing enterprise and merchant information. This technical field mainly involves multiple technical links such as data collection, recognition, verification, and standardized output. With the development of Internet technology and the application of artificial intelligence, yellow page number recognition has gradually developed from traditional manual entry to automated recognition. However, data format diversity and standardized output remain important technical challenges in this field.

[0003] Currently, the common yellow page number recognition technologies on the market mainly adopt the combination of OCR image recognition and rule matching, and extract information such as telephone numbers and addresses in various format documents through preset templates. For example, some systems use a matching algorithm based on regular expressions to recognize telephone numbers in different formats, while others use an entity extraction model based on deep learning to recognize and classify each field in merchant information.

[0004] The most relevant existing technology is a method for number recognition and verification based on an AI large model. This method uses artificial intelligence technology to batch process and recognize yellow page data. Its main technical principle is to take various format documents collected as input, perform natural language understanding and information extraction through an AI model, and output structured merchant information data. This method can handle certain format changes and data differences by training the model to learn various number formats and context relationships.

[0005] However, the existing technology has significant technical defects: First, dealing with yellow page number format problems requires a large amount of code and it is impossible to predict and handle all possible incorrect formats; second, data operation and maintenance personnel need to spend a lot of time organizing the data input into the large model, and developers also need to write a large amount of code to process the model output data; in addition, there is a lack of a preprocessing mechanism for data garbled situations during the model usage process, resulting in difficulties in ensuring the consistency and accuracy of the model output results. Summary of the Invention

[0006] In view of this, this application provides a method and system for consistency calibration and standardized output of AI model recognition results, which solves the problems in the existing technology that dealing with yellow page number formats requires a large amount of code, time-consuming data organization, and lack of a data garbled preprocessing mechanism.

[0007] The embodiment of this application provides a method for consistent calibration and standardized output of AI model recognition results, including: receiving yellow page number data in a preset format set, using a format recognition algorithm to identify the source type and format features of the yellow page number data, and using an input format unification algorithm to standardize key fields to obtain yellow page number data in a standardized format; using the input instructions of a structured AI model generated by a prompt template construction algorithm stored in a memory, performing yellow page number data recognition and verification processing through a fine-tuned general AI large model, and converting the processing results into preliminary standardized AI recognition results that conform to preset specifications according to fixed field output format rules; according to the preliminary standardized AI recognition results, using common error mapping table matching and error case knowledge base verification to intelligently correct the recognition results, calculating the correct rate index of the current batch, and outputting calibrated recognition results with confidence scores, and using standardized output format conversion to generate final consistent calibration and standardized number recognition results; according to the final consistent calibration and standardized number recognition results, using a performance index calculation model to calculate the overall recognition correct rate and scores of each dimension, identifying high-frequency error patterns to generate an error pattern classification report, and updating error correction rules and general AI large model optimization strategies.

[0008] Optionally, the step of using an input format unification algorithm to standardize key fields to obtain yellow page number data in a standardized format includes: based on the source type and format features of the yellow page number data, using a number type standard library to match and verify the numbers in the yellow page number data to obtain structured data with type identifiers; according to the structured data with type identifiers, using an automated format input template generation tool and a preset format template framework to convert and generate yellow page number data in a standardized format.

[0009] Optionally, performing yellow page number data recognition and verification processing through a fine-tuned general AI large model, and converting the processing results into preliminary standardized AI recognition results that conform to preset specifications according to fixed field output format rules includes: using the general AI large model to process the yellow page number data in the standardized format based on the structured AI model input instructions to obtain an original AI recognition result; according to the original AI recognition result, using preset fixed field output format rules to perform format standardization processing on the recognition results one by one, and outputting the preliminary standardized AI recognition results.

[0010] Optionally, intelligent correction of the recognition results is performed by matching with a common error mapping table and verifying with an error case knowledge base, the accuracy rate index of the current batch is calculated, and a calibrated recognition result with a confidence score is output, including: using a preset common error mapping table to match the preliminary standardized AI recognition results, identifying and correcting character confusion and format errors to obtain a first preliminarily calibrated recognition result; according to the first preliminarily calibrated recognition result, using the error case knowledge base for in-depth verification and correction, updating the error case knowledge base according to the verification and correction results, and outputting a second preliminarily calibrated recognition result; according to the second preliminarily calibrated recognition result, using a dynamic feedback algorithm to calculate the accuracy rate index of the current batch, and obtaining the calibrated recognition result with a confidence score.

[0011] Optionally, the step of using the dynamic feedback algorithm to calculate the accuracy rate index of the current batch and obtaining the calibrated recognition result with a confidence score includes: according to the second preliminarily calibrated recognition result, using a logarithmic function to process the quantity ratio of the input yellow page number data to the output recognition result to obtain a preliminary performance evaluation score, where the preliminary performance evaluation score reflects the data processing effect of the second preliminarily calibrated recognition result; according to the preliminary performance evaluation score, using a dynamic weight adjustment algorithm to calculate the evaluation indexes of each dimension and generate a multi-dimensional performance score, where the multi-dimensional performance score includes the correct rate of the number format, the correct rate of the address format, and the correct rate of the merchant information; according to the multi-dimensional performance score, using a preset confidence calculation model to comprehensively evaluate the reliability of the recognition result, and when the score of any dimension in the multi-dimensional performance is lower than the first preset threshold, recalibrating the recognition result of the corresponding dimension, and outputting the calibrated recognition result with a confidence score.

[0012] Optionally, using the performance index calculation model to calculate the overall recognition accuracy rate and the scores of each dimension, identifying high-frequency error patterns to generate an error pattern classification report, and updating the error correction rules and model optimization strategies includes: according to the finally consistent calibrated and standardized number recognition results, using the performance index calculation model to calculate the overall recognition accuracy rate and the scores of each dimension to obtain a model performance evaluation report; using an error pattern analysis algorithm to identify the high-frequency error patterns and abnormal situations in the model performance evaluation report, and outputting the error pattern classification report; according to the error pattern classification report, using a rule base update mechanism to update the error correction rules and model optimization strategies.

[0013] Optionally, before performing the identification and verification processing of the yellow page number data by the fine-tuned general AI large model, it further includes: constructing a linear-nonlinear hybrid bandit learning framework, including: constructing a hybrid bandit learning framework containing a linear model and a nonlinear model based on the yellow page number data in the standardized format to generate a hybrid processing framework; wherein the linear model is used to process data items that conform to the first preset structural rule, and the nonlinear model is used to process data items that require context understanding; according to the hybrid processing framework, introducing a time attention mechanism, using a time feature extractor to extract time tags and temporal patterns, and using an attention calculation unit to assign weights to data in different time periods to obtain the yellow page number data features with time weights; according to the yellow page number data features with time weights, using the UpperConfidenceBound algorithm to dynamically adjust the balance between exploring new strategies and exploiting known strategies, and outputting an optimized processing strategy for yellow page number data identification.

[0014] Optionally, before normalizing the keyword fields using the input format unification algorithm, it further includes: constructing an equivariant network preprocessing framework for Markov data, including: receiving the yellow page number data in the preset format set, constructing it into a Markov process, and using an equivariant network construction algorithm to generate an equivariant network structure containing a Markov transition modeler, an equivariant feature extractor, and a structural consistency verifier; according to the equivariant network structure, analyzing the generalization boundary of the equivariant network structure when processing the yellow page number data, and using the theoretical upper bound of generalization error calculation and empirical test results to generate a targeted data augmentation strategy; according to the data augmentation strategy, analyzing the preset complexity index and dependence structure of the yellow page number data, and dynamically adjusting the order of the Markov process and the network structure parameters according to the analysis results, and outputting an optimized equivariant network preprocessing model.

[0015] Optionally, before performing intelligent correction on the recognition result using the common error mapping table matching and error case knowledge base verification, it further includes: establishing an unbounded Gaussian differential privacy sampling mechanism, including: receiving the preliminary standardized AI recognition result, using the unbounded Gaussian mechanism and the differential privacy protection mechanism to add Gaussian-distributed noise to the yellow page number data to obtain data with differential privacy protection; according to the data with differential privacy protection, using the gradient descent method or reinforcement learning technology to dynamically adjust the parameters of the differential privacy sampling, optimize the sampling strategy, and output the optimized unbounded Gaussian differential privacy sampling parameters; according to the optimized unbounded Gaussian differential privacy sampling parameters, tracking the privacy consumption in each processing step through a privacy budget management mechanism, and automatically adjusting the processing strategy of the unbounded Gaussian differential privacy sampling mechanism when detecting that the privacy budget is lower than the preset threshold, and generating a final calibration processing strategy, which is used to guide the subsequent intelligent correction process.

[0016] The embodiment of the present application further provides a system for consistent calibration and standardized output of AI model recognition results, including a data preprocessing module, an AI model enhancement processing module, a post-processing calibration module, and a dynamic optimization module connected in sequence, where: The data preprocessing module is used to receive yellow page number data in a preset format set, and use a format recognition algorithm to identify the source type and format characteristics of the yellow page number data, and use an input format unification algorithm to standardize keyword fields, and output the yellow page number data in a standardized format to the AI model enhancement processing module; The AI model enhancement processing module is used to generate structured input instructions for the AI model using the stored algorithm based on the prompt template, perform yellow page number data recognition and verification processing through a fine-tuned general AI large model, and convert the processing result into a preliminary standardized AI recognition result that conforms to the preset specification according to the fixed field output format rule, and transmit the preliminary standardized AI recognition result to the post-processing calibration module; The post-processing calibration module is used to perform intelligent correction on the recognition result according to the preliminary standardized AI recognition result by using common error mapping table matching and error case knowledge base verification, calculate the correct rate index of the current batch, and obtain a calibrated recognition result with a confidence score, and use a standardized output format conversion to generate a final consistent calibration and standardized number recognition result, and send the final consistent calibration and standardized number recognition result to the dynamic optimization module; The dynamic optimization module is used to calculate the overall recognition correct rate and scores of each dimension according to the final consistent calibration and standardized number recognition result by using a performance index calculation model, identify high-frequency error patterns to generate an error pattern classification report, and update the error correction rule and the optimization strategy of the general AI large model.

[0017] The embodiment of the present application further provides a computer system, the computer system includes: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for consistent calibration and standardized output of the above-mentioned AI model recognition results.

[0018] The embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the method for consistent calibration and standardized output of the above-mentioned AI model recognition results.

[0019] The embodiment of the present application further provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the method for consistent calibration and standardized output of the above-mentioned AI model recognition results are implemented.

[0020] The present application has the following technical effects: By constructing a data preprocessing module and using an input format unification algorithm and an automated JSON input template generation tool, multi-source heterogeneous yellow page number data is converted into standardized AI model inputs, significantly reducing the impact of data input diversity on model output; By designing an AI model enhancement layer and using a constrained output template design and a fixed field output format rule, the fine-tuned large AI model can output structured and consistent recognition results, improving the standardization degree of model output; By implementing a post-processing calibration algorithm, establishing a common error mapping table and an error case knowledge base, and combining a dynamic feedback algorithm, the recognition results of the AI model are accurately corrected and confidence evaluated, significantly improving the accuracy of the recognition results; By establishing a dynamic optimization mechanism and using a performance metric calculation model and an error pattern analysis algorithm, a closed-loop feedback system from recognition results to rule updates is constructed, achieving continuous self-optimization of the system and ensuring the recognition accuracy during long-term operation; By innovatively designing a full-process standardized processing framework for yellow page number recognition and organically combining preprocessing, model processing, and post-processing, the technical bottleneck of traditional methods in processing diverse data is solved, significantly reducing the workload of data operation and maintenance and development personnel. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for use in the embodiments will be briefly introduced below. The accompanying drawings herein are incorporated into the specification and form a part of this specification. These accompanying drawings show embodiments consistent with the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following accompanying drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related accompanying drawings can be obtained based on these accompanying drawings without creative efforts.

[0022] Figure 1 is a flowchart of a method for consistency calibration and standardized output of AI model recognition results provided by an embodiment of the present application;

[0023] Figure 2 is an implementation flowchart of data preprocessing provided by an embodiment of the present application;

[0024] Figure 3 is an implementation flowchart of data recognition and verification processing provided by an embodiment of the present application;

[0025] Figure 4 is an implementation flowchart of generating number recognition results provided by an embodiment of the present application;

[0026] Figure 5 is a system structure diagram of consistency calibration and standardized output of AI model recognition results provided by an embodiment of the present application. Detailed implementation manners

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only some of the embodiments of the present disclosure, rather than all of the embodiments. The components of the embodiments of the present disclosure described and illustrated herein generally may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0028] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0029] The term "and / or" in this document merely describes an association relationship and indicates that three relationships may exist. For example, A and / or B may represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" in this document means any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C may represent any one or more elements selected from the set composed of A, B, and C.

[0030] The following will describe the specific implementation manners of the present application in detail with reference to the accompanying drawings. It should be understood that the following description is only for illustrative purposes and is not used to limit the protection scope of the present application.

[0031] As Figure 1 shown, the embodiments of the present application provide a method for consistency calibration and standardized output of AI model recognition results, including:

[0032] S1: Receive yellow page number data in a preset format set, and use a format recognition algorithm to identify the source type and format features of the yellow page number data, and use an input format unification algorithm to standardize keyword fields to obtain yellow page number data in a standardized format.

[0033] First, in the data preprocessing stage of the embodiments of the present application, diverse yellow page number data from different sources is received. This data may exist in various forms, including but not limited to various formats such as txt text files, Excel spreadsheets, CSV files, etc. These data have significant differences in format and structure, posing challenges to subsequent AI model processing.

[0034] Specifically, a dedicated format recognition algorithm is designed. This algorithm can automatically analyze the characteristics of the input data, including information in multiple dimensions such as file extensions, content format features, data separator types, etc. Through this intelligent analysis, the source type and format features of the yellow page number data can be accurately identified, providing a basis for subsequent standardization processing. For example, for the spreadsheet format, its header structure, data column types, and inter-column relationships will be recognized; for the text format, its separators and data organization methods will be analyzed.

[0035] Subsequently, an input format unification algorithm is applied to normalize the recognized data. The core function of this algorithm is to accurately extract key fields from the original data of various formats, such as phone numbers, area codes, addresses, merchant names, etc., and perform preliminary normalization processing on these fields. For example, for phone numbers, non-standard separators and spaces will be removed to unify the format; for address information, the address components will be decomposed and structured. This normalization processing ensures that data from different sources can be converted into a unified standard format, creating conditions for subsequent processing by the AI model.

[0036] As Figure 2 shown, S1 includes:

[0037] S1.1: Based on the source type and format features of the yellow page number data, use the number type standard library to match and verify the numbers in the yellow page number data, and obtain structured data with type identifiers.

[0038] After completing the preliminary identification of the data source and extraction of format features, it is necessary to further accurately identify the type and verify the validity of the key element - the phone number in the yellow page number data. The implementation of this step depends on a pre-established number type standard library, which contains standard format patterns for various types of numbers.

[0039] Specifically, during implementation, first, possible phone number fields are extracted from the preprocessed data. Then, using the pattern rules defined in the number type standard library, pattern matching and type identification are performed on these numbers. For example, according to characteristics such as the number of digits and prefixes of the number, it will be identified as different types such as mobile phone numbers, landline numbers, or special service numbers. During the identification process, specific verification rules will also be applied, such as the matching check of the area code and the number of digits, and the validity verification of the number prefix, to ensure the accuracy of the identification results.

[0040] For numbers with format recognition doubts, such as the fixed - line format differences between "023xxxxxxx" and "23xxxxxxx", they will be distinguished through a specific area - code verification algorithm. For example, it will check whether "023" is a valid area code and combine it with the subsequent number length to determine the correct format. Similarly, for the distinction between the area code "0955" and the service number "95", it will be identified by referring to the preset area - code database and service - number rules.

[0041] Through this process, finally, structured data with clear type identifiers is output. Each telephone number is assigned a clear type label (such as mobile phone number, landline number, service number, etc.) and a validity identifier, providing an important basis for subsequent processing.

[0042] S1.2: According to the structured data with type identifiers, use an automated format input template generation tool and a preset format template architecture to convert and generate yellow - page number data in a standardized format.

[0043] After obtaining the structured data with type identifiers, it is necessary to convert this data into a standardized format so that the AI model can process it efficiently. This step is achieved through an automated format input template generation tool and a preset format template architecture.

[0044] First, use an automated format input template generation tool to process the structured data. This tool pre - sets regular - expression templates for more than 30 common data formats, covering various telephone - number formats, area - code formats, address formats, etc. Through these templates, various fields in the data can be identified and extracted, even if the forms of these fields in the original data are diverse.

[0045] For example, when processing table data containing the number "010XXX456", merchant name "Test Merchant", address "Test Address", and longitude and latitude "Test Longitude, Test Latitude", it will automatically generate a JSON data structure in the following format:

[0046] {'phone': ['XXX456', 'XXX457'],

[0047] 'areaCode': '010',

[0048] 'address': 'Test Address',

[0049] 'lng': 123.123, 'lat': 96.233}

[0050] This standardized JSON format contains clearly defined fields and values and can be directly used to construct the prompt templates for inputting into the AI model. In this way, the data diversity of the input model is effectively reduced, and the accuracy of the model output format is improved.

[0051] During the conversion process, the preset format template architecture is also applied to ensure the consistency and integrity of the conversion results. These architectures define the fields that must be included in the standard output, the relationships between fields, and the format requirements for values.

[0052] For example, for telephone number data, the architecture requires that the area code and local number be stored separately and maintain a specific format; for address data, the architecture defines the components of the address and their organization method.

[0053] Through the above processing, the yellow page number data in the standardized format is finally output. These data are not only unified in structure but also highly consistent in terms of field definition and value representation, providing an ideal input for subsequent AI model processing.

[0054] Before normalizing the keyword fields using the input format unification algorithm, it also includes: constructing an equivariant network preprocessing framework for Markov data, including:

[0055] A1: Receive the yellow page number data in the preset format set, construct it into a Markov process, and use the equivariant network construction algorithm to generate an equivariant network structure including a Markov transition modeler, an equivariant feature extractor, and a structural consistency validator.

[0056] Constructing an equivariant network preprocessing framework for Markov data is an important enhancement to the data preprocessing module.

[0057] In this step, first, various types of yellow page number data in the preset format set are received. These data may have diverse sources and different formats. Innovatively, these data are regarded as a Markov process, that is, it is assumed that each element in the data (such as the digits in the number, the components in the address) has a conditional dependence relationship with its preceding and following elements. For example, in a telephone number, the area code part and the local number part usually follow different distribution rules, and the choice of the area code will affect the legality and distribution characteristics of the subsequent number.

[0058] Based on this Markov property, an equivariant network construction algorithm is used to generate a complex network structure, which consists of three core components: a Markov transition modeler, an equivariant feature extractor, and a structural consistency validator. The Markov transition modeler is mainly responsible for capturing the conditional dependencies between data elements. It can learn and represent statistical regularities such as those between area codes and subsequent numbers. The equivariant feature extractor is specifically used to extract format-independent data features, making it robust to minor changes in the data, such as changes in separators, order adjustments, etc. The structural consistency validator ensures that the processed data maintains its structural integrity and avoids the destruction of the data structure during processing.

[0059] A key feature of the equivariant network is its invariance or predictable transformation relationships with respect to specific transformations of the input data, such as translation, rotation, permutation, etc. This enables effective handling of various format changes when processing yellow page number data, such as different separator usages, digital order adjustments, etc., thus improving the robustness and accuracy of preprocessing.

[0060] A2: According to the equivariant network structure, analyze the generalization boundary of the equivariant network structure when processing yellow page number data, and use the theoretical upper bound of generalization error calculation and empirical test results to generate targeted data augmentation strategies.

[0061] After constructing the equivariant network structure, it is necessary to evaluate its generalization ability boundary when processing yellow page number data, that is, to determine the range of data changes that the network can reliably handle. This step is achieved through a combination of theoretical analysis and empirical testing.

[0062] At the level of theoretical analysis, use the method of calculating the upper bound of generalization error to evaluate the theoretical performance upper limit of the equivariant network when processing different types of data changes. This calculation takes into account factors such as the structural complexity of the equivariant network, the scale of the training data, the Markov properties of the data, and the equivariance constraints, and mathematically determines the theoretical boundaries of the network's generalization ability.

[0063] At the same time, a large number of empirical tests are also carried out. By inputting various variants of yellow page number data into the equivariant network, such as adding different levels of noise, changing the data format, adjusting the data structure, etc., the actual processing performance of the network under these changes is evaluated. Combining the test results with the theoretical analysis enables accurate identification of the strong and weak regions of the network's generalization ability.

[0064] Based on the above analysis results, targeted data augmentation strategies are generated. For data types with weak generalization ability of the equivariant network, such as non-standard number representations in specific formats or complex address structure changes, more variant samples of such data will be generated for enhanced training; while for data types for which the network already has good generalization ability, unnecessary enhancements will be reduced to optimize the usage efficiency of computing resources.

[0065] This data augmentation strategy not only improves the ability to process various non-standard data formats, but also enables the adaptability to new data forms that may appear in the future, enhancing the long-term usability and stability.

[0066] A3: According to the data augmentation strategy, analyze the preset complexity metrics and dependency structure of the yellow page number data, and dynamically adjust the order of the Markov process and the network structure parameters according to the analysis results, and output an optimized equivariant network preprocessing model.

[0067] After obtaining the targeted data augmentation strategy, further analyze the preset complexity metrics and dependency structure of the yellow page number data to achieve dynamic structure optimization of the Markov model. This enables the preprocessing module to adaptively process yellow page data with different complexities.

[0068] First, the complexity of the yellow page number data will be evaluated, considering factors such as the structural complexity of the data, the strength of the dependency relationships between elements, and the format diversity. For example, for telephone number data in a standard format, its complexity score is relatively low; while for address data containing non-standard formats, multiple nested structures, or special characters, its complexity score is relatively high.

[0069] Based on the complexity evaluation results, dynamically adjust the order of the Markov process. For data with simple structures and clear dependency relationships (such as telephone numbers in standard formats), a low-order Markov model, such as a first-order or second-order model, may be used, which can reduce the computational complexity while ensuring the processing quality; while for data with complex structures and rich dependency relationships (such as address expressions composed of multiple parts), a high-order Markov model, or even a variant model such as a hidden Markov model, may be used to better capture the long-distance dependency relationships in the data.

[0070] At the same time, the structure parameters of the equivariant network, such as the number of network layers, the number of nodes, and the connection weights, will also be adjusted according to the data characteristics. For data types that require stronger expressive power, the network complexity will be increased; while for simple data types, the network structure may be simplified to improve the processing efficiency.

[0071] Through this dynamic optimization, an equivariant network preprocessing model adapted to specific data characteristics is finally output. This model can optimize the usage efficiency of computing resources while ensuring the processing accuracy. This method not only improves the adaptability of the model, but also reduces unnecessary computational overhead, providing significant performance advantages for its practical applications.

[0072] S2: Generate the input instructions for the structured AI model using the stored algorithm for constructing based on the prompt template, perform the identification and verification processing of the yellow page number data through the fine-tuned general AI large model, and convert the processing results into the preliminary standardized AI recognition results that meet the preset specifications according to the fixed field output format rules.

[0073] In the AI model enhancement processing stage of implementing the present invention, the fine-tuning of the general AI large model is a key step to ensure the recognition accuracy. The general AI large model adopted in the present invention is based on the Transformer architecture, includes 12 layers of attention mechanisms and a 768-dimensional hidden layer, and uses the ReLU activation function. More than 500,000 manually annotated yellow page number data are used as the training set during the fine-tuning process, and these data cover various formats of telephone numbers, addresses, and merchant information.

[0074] The following hyperparameter settings are adopted during the fine-tuning process: the initial learning rate is set to 3e -5 , and the cosine annealing scheduling strategy is adopted; the batch size is 32; the number of training epochs is 5; the cross-entropy loss function is used; the AdamW optimizer is used, and the weight decay is 0.01. To prevent overfitting, 10% of random dropout is also introduced.

[0075] The model evaluation uses accuracy, precision, recall, and F1 score as the main indicators, and the comprehensive accuracy rate on the validation set reaches 95.2%. In particular, the recognition accuracy rate for standard format numbers is as high as 98.7%, and the recognition accuracy rate for abnormal format numbers also reaches 93.1%, which is significantly better than the traditional methods.

[0076] After completing the data preprocessing, enter the AI model enhancement layer processing stage. In this stage, first generate the input instructions for the structured AI model using the stored algorithm for constructing based on the prompt template. This algorithm receives the standardized JSON format yellow page number data output from the data preprocessing module as the basic input. These data have been processed with unified formats, contain key information such as numbers, area codes, addresses, longitude and latitude, etc., and have a consistent data structure.

[0077] Apply the preset algorithm for constructing the prompt template to analyze the standardized JSON data. This algorithm will select the corresponding prompt template according to the task type (such as number recognition, number verification, number appeal, or number correction). Each template contains the key instructions and format requirements for guiding the AI model to correctly understand and process the data.

[0078] Specifically, the input instructions of the constructed model usually include the following key parts: task description, input data format description, expected output format definition, processing rule description, and example demonstration. The task description clearly tells the model the specific task to be completed; the input data format description helps the model understand the meaning of each field in the JSON structure; the expected output format definition clearly stipulates the data structure and field requirements that the model should output; the processing rule description includes the processing principles for special cases; the example demonstration provides input-output comparison to help the model better understand the task requirements.

[0079] In the construction process, the prompt template algorithm particularly emphasizes the constraints on the output format. For example, it clearly specifies that the landline number must include the area code and seven digits, the mobile phone number must be eleven digits, and each number type needs to output key fields such as the original data, cleaned data, confidence level, and anomaly identification.

[0080] After generating the structured input instructions, input them into a general AI large model that has been specifically fine-tuned for processing. The fine-tuning process here is crucial. A large amount of standard yellow page number data has been used to specifically train the general AI large model, enabling it to accurately understand various formats and characteristics of the yellow page number data. The fine-tuning process mainly focuses on enabling the model to learn the standard data patterns, master the number format rules, and be able to accurately extract and verify key information from complex texts.

[0081] As Figure 3 shown, S2 includes:

[0082] S2.1: Use the general AI large model to process the yellow page number data in the standardized format based on the structured AI model input instructions to obtain the original AI recognition result.

[0083] When the structured AI model input instructions and the yellow page number data in the standardized format are ready, input these contents into the general AI large model that has been fine-tuned. This processing process is a key link in the core function and directly affects the quality of the final recognition result.

[0084] During the processing, the AI large model first analyzes the input instructions and data. The instruction part tells the model the type of task to be executed and the processing method, while the data part provides the yellow page number information to be analyzed. Due to the previous data preprocessing and structured instruction generation, the model can clearly understand the task requirements and data structure, greatly improving the processing efficiency and accuracy.

[0085] Subsequently, the model initiates a series of processing steps: First, it conducts an in-depth analysis of the input data to understand various fields and information contained therein; then, it performs specific processing on the data according to the task requirements, such as verifying the legality of the number format, identifying the number type, correcting obvious format errors, etc.; finally, it generates a processing result according to the preset output requirements.

[0086] Throughout the processing, the model makes full use of the domain knowledge obtained through fine-tuning to accurately analyze the yellow page number data. For example, the model can recognize that "010 - 123XXXXX" and "010123XXXXX" are actually the same number, just with different format representations; it can determine that "400 - 8XX - 88XX" is a service hotline rather than an ordinary landline number; it can accurately extract valid phone number information from mixed text, etc.

[0087] Meanwhile, the model also generates confidence scores during the processing, reflecting its confidence in the processing results. These scores provide important references for subsequent calibration and optimization. For example, for numbers with clear and standard formats, the model may give a higher confidence; while for non-standard formats or numbers containing ambiguous information, it may give a lower confidence.

[0088] Through this process, the original AI recognition results are output. These results contain the model's initial processing and understanding of the input data, but may not fully meet all the requirements of the standardized output format and need to be further processed for format standardization in the next step.

[0089] S2.2: According to the original AI recognition results, adopt the preset fixed-field output format rules to perform format standardization processing on each recognition result one by one, and output the preliminary standardized AI recognition results.

[0090] After obtaining the original AI recognition results, it is necessary to perform format standardization processing on these results to ensure the consistency and normativeness of the output results. This step is achieved by applying the preset fixed-field output format rules.

[0091] Although the original AI recognition results contain the model's processing results of the yellow page number data, there may be inconsistencies or incompleteness in the format and structure. For example, some fields may be missing, the formats are not unified, the representation methods of field values are inconsistent, etc. These inconsistencies will bring difficulties to subsequent processing and use, so standardization processing is required.

[0092] These raw results are processed one by one using preset fixed-field output format rules. These rules include format requirements and conversion rules for different fields, especially strict specifications for phone numbers. For example, landline numbers are required to be output in the format of "area code + 7-digit number", such as "010 - 123XXXXX"; mobile phone numbers must strictly follow the 11-digit number format standard, such as "139123XXXXX", and no additional separators are allowed. For fields such as addresses and merchant names, there are also corresponding format specifications.

[0093] In the specific implementation process, the corresponding format rules are applied to each field output by the AI model one by one. For the number field, it is checked whether it meets the preset format requirements, and if not, the corresponding format conversion is carried out; for the address field, it is structured and standardized to ensure the order and format of the address components are unified; for other fields such as longitude and latitude, merchant names, etc., there are also corresponding format standards.

[0094] In addition, it is ensured that the output results include all necessary standard fields, such as raw data, cleaned data, confidence level, and anomaly identification, etc., to ensure the integrity and consistency of the output. For necessary fields that the AI model fails to generate, default values are supplemented or marked as missing according to the rules.

[0095] Through this process, preliminary standardized AI recognition results that meet the preset specifications are output. These results are highly unified in format and structure, providing a good foundation for subsequent calibration processing. This standardization not only improves the readability and processability of the results but also creates conditions for error detection and correction in the next step.

[0096] Before the yellow page number data is identified and verified by the fine-tuned general AI large model, it also includes: constructing a linear-nonlinear hybrid bandit learning framework, including:

[0097] B1: According to the yellow page number data in the standardized format, construct a hybrid bandit learning framework containing a linear model and a nonlinear model to generate a hybrid processing framework; where the linear model is used to process data items that conform to the first preset structural rule, and the nonlinear model is used to process data items that require context understanding.

[0098] Before entering the AI model processing stage, by constructing a linear-nonlinear hybrid bandit learning framework, the model's processing ability and adaptability to yellow page number data can be significantly improved. This framework creates a hybrid structure that combines a linear model and a nonlinear model based on the standardized format of yellow page number data, enabling it to more flexibly handle various data processing scenarios.

[0099] The core feature of the hybrid processing framework is to intelligently allocate processing tasks according to data characteristics. Specifically, the linear model in it is mainly responsible for processing data items with high structuralization and clear patterns, such as standard format phone numbers, standardized area codes, etc. Such data usually conforms to the first preset structural rule, has clear format definitions and verification standards, and is suitable for processing using a linear model with high computational efficiency and concise logic. For example, to determine whether an 11-digit number is a valid mobile phone number, it can be quickly verified whether the first three digits are legal operator codes through a linear model.

[0100] At the same time, the non-linear model in it (mainly the large language model LLM) is specifically responsible for processing data items with high complexity and requiring context understanding, such as irregular format address descriptions, merchant descriptions with mixed information, etc. Such data usually requires understanding semantics, context, and implicit relationships, and simple rule matching is difficult to handle effectively, so it is necessary to utilize the powerful semantic understanding and context analysis capabilities of the non-linear model. For example, to extract structured address information from the text "On the first floor of Lenovo Building at the southwest corner of the intersection of North Third Ring West Road and Zhongguancun South Street", the deep semantic understanding ability of the non-linear model is required.

[0101] It will dynamically determine which model to allocate specific data to according to the characteristics of the input data and the historical processing effects. This intelligent allocation mechanism not only improves the processing efficiency but also gives full play to the expertise of various models and optimizes the overall processing effect. At the same time, for some complex data, it may also call both the linear and non-linear models simultaneously, and then synthesize the processing results of both to further improve the accuracy.

[0102] This hybrid architecture makes full use of the high efficiency of the linear model and the complex expression ability of the non-linear model, enabling it to more flexibly handle various yellow page data recognition scenarios, thereby improving the overall processing performance.

[0103] B2: According to the hybrid processing framework, introduce a time attention mechanism, use a time feature extractor to extract time markers and time series patterns, and adopt an attention calculation unit to assign weights to data in different time periods to obtain yellow page number data features with time weights.

[0104] Based on the hybrid slot machine learning framework, an innovative time attention mechanism is introduced, enabling the model to focus on the time dimension features of the data, thereby more effectively capturing the characteristics of yellow page number data evolving over time. This mechanism is particularly suitable for processing long-running yellow page data because the data format and features often change over time.

[0105] The temporal attention mechanism first extracts temporal markers and temporal patterns from the data through a temporal feature extractor. The changing trends of the features of the yellow page data processed in different time periods are recorded and analyzed, including newly added merchant types, changes in regional formats, or newly emerging number formats, etc. For example, it may be observed that a new telephone number prefix appears in a certain time period, or the address expression method in a specific region changes. These time-related patterns are recorded and form a temporal feature library.

[0106] Subsequently, the attention calculation unit of will assign weights to the data in different time periods. This process assigns higher weights to the recent data and data in historically similar scenarios based on the temporal markers and historical performance of the data. For example, if it is found that a new number format pattern has appeared in the recent month, then when processing the current data, higher attention will be given to these new patterns; at the same time, if special data patterns have occurred in a specific season (such as during holidays) in history, then the attention to these patterns will also be increased in similar time periods.

[0107] Through this weight assignment, the features of the yellow page number data with temporal weights are obtained, enabling the model to prioritize the most relevant temporal patterns. This feature representation with a temporal dimension enables the model to not only understand the current structure of the data but also perceive the temporal evolution trend of the data, thereby improving the ability to process emerging data patterns.

[0108] The temporal attention mechanism is implemented through three key components: a temporal feature extractor (which extracts temporal markers and temporal patterns from the data), an attention calculation unit (which assigns weights to the data in different time periods), and an adaptive fusion layer (which integrates the outputs of different models according to the attention weights). This mechanism enables to automatically adapt to the format changes and emerging patterns that evolve over time in the yellow page data, such as new number formats, changes in address representation methods, etc., greatly improving the long-term usability and adaptability of.

[0109] B3: According to the features of the yellow page number data with temporal weights, use the UpperConfidenceBound algorithm to dynamically adjust the balance between exploring new strategies and exploiting known strategies, and output an optimized processing strategy for yellow page number data recognition.

[0110] After obtaining the features of the yellow page number data with temporal weights, the core mechanism of hybrid slot machine learning is implemented: dynamically balancing exploration and exploitation. This step realizes an intelligent balance between "exploring new strategies" and "exploiting known effective strategies" through the UpperConfidenceBound (UCB) algorithm, ensuring both stable and efficient performance and continuous improvement and adaptation to new situations.

[0111] Specifically, the ratio between exploration and exploitation is dynamically adjusted based on the feedback of the processing results. When encountering new data patterns or a decline in the effectiveness of existing strategies, the proportion of exploration is increased to try new processing methods; while when a certain strategy is proven to be efficient and reliable, the utilization of that strategy is increased to improve processing efficiency.

[0112] When implementing this balance, the UCB algorithm takes into account both the historical performance and uncertainty of each processing strategy. The algorithm calculates a score for each possible processing strategy, which consists of two parts: the average historical performance of the strategy (i.e., the exploitation term) and the upper confidence bound representing uncertainty (i.e., the exploration term). The strategy with the highest score is selected for trial, thus naturally balancing exploration and exploitation.

[0113] For example, when processing the area code recognition task, multiple recognition strategies may be maintained simultaneously: rule-based matching, context-based inference, region information-based verification, etc. Through the UCB algorithm, which strategy to use is dynamically determined based on the historical performance and current uncertainty of these strategies. If a certain strategy (such as rule-based matching) has always performed well when processing a specific type of area code, it will be used more frequently; but if a new area code format appears, other strategies will also be given sufficient opportunities to try to discover more suitable processing methods.

[0114] Through this dynamic balance mechanism, both a high recognition accuracy rate can be maintained, and better processing methods can be continuously explored to achieve continuous self-optimization. This is particularly important for the long-term operation of yellow page number recognition, because data characteristics and patterns may change over time, and it is necessary to be able to adapt to these changes and continuously optimize its performance.

[0115] Finally, optimized yellow page number data recognition processing strategies are output, which combine historical experience and innovative attempts, can effectively process various yellow page data, and continuously optimize as the data characteristics change.

[0116] S3: According to the preliminary standardized AI recognition results, intelligent correction of the recognition results is carried out by using common error mapping table matching and error case knowledge base verification, the correct rate index of the current batch is calculated, and a calibrated recognition result with a confidence score is output, and a final consistent calibration and standardized number recognition result is generated by using standardized output format conversion.

[0117] As Figure 4 shown, S3 includes:

[0118] S3.1: Using a preset common error mapping table, match the preliminary standardized AI recognition results, identify and correct character confusion and format errors, and obtain the first preliminarily calibrated recognition result;

[0119] In the first step of the post - processing calibration phase, a preset common error mapping table is used to match and correct the preliminary standardized AI recognition results. This step targets the common character confusion and format errors that may occur during the AI model processing, and performs precise correction through pattern matching and replacement.

[0120] The common error mapping table is an important component. It contains more than 2,000 high - frequency error cases, covering various common character confusion and format error types. For example, the confusion between the number "0" and the letter "O" (e.g., "0XX0" should be "OXXO"), the confusion between the vertical bar "|" and the numbers "1" and the letter "I" (e.g., "10XX5|10XX6" should be "10XX5, 10XX6"), the incorrect handling of special characters (e.g., " 95XXX" should be "95XXX"), etc. These error patterns and their correct corresponding forms are pre - entered and continuously updated.

[0121] The calibration process uses the method of pattern matching and replacement. It will traverse each field in the recognition result, compare it with the patterns in the error mapping table. When a match is found, the error pattern will be replaced with the correct form according to the mapping relationship. This replacement is not limited to simple character replacement, but also includes more complex pattern conversions, such as the replacement of character combinations in specific context environments.

[0122] For example, when it is detected that the recognition result contains a character sequence like "13o", it will be recognized that the "o" is actually an incorrect recognition of the number "0", and it will be corrected to "130"; similarly, for a sequence like "I0XX6", it will be recognized that the "I" should be the number "1" and corrected accordingly. This correction is based on the analysis and summary of a large number of actual error cases and can effectively handle most common character confusion situations.

[0123] In addition, the details of each correction will be recorded, including information such as the original text, the recognized error type, the corrected text, etc. These records will be used for subsequent error analysis and mapping table update. In this way, the mapping table can be continuously improved and expanded to enhance the coverage ability for newly emerging error types.

[0124] Through this process, the first - stage preliminary calibrated recognition results are output. These results have corrected most common character confusion and format errors, laying a foundation for subsequent in - depth calibration. Through this precise error correction, the accuracy of the recognition results is significantly improved, reducing the manual correction workload of data operation and maintenance personnel.

[0125] S3.2: According to the recognition result of the first preliminary calibration, use the error case knowledge base for in-depth verification and correction, update the error case knowledge base according to the verification and correction results, and output the recognition result of the second preliminary calibration.

[0126] After the preliminary correction based on the common error mapping table, it enters a more complex in-depth calibration stage. This step uses the error case knowledge base to conduct a deeper verification and correction of the recognition result of the first preliminary calibration, dealing with complex error patterns that are difficult to solve through simple character mapping.

[0127] The error case knowledge base is another key component. Different from the common error mapping table, it contains more complex and context-related error patterns, which are usually difficult to solve through simple character mapping and require considering more semantic and context information. The error case knowledge base is automatically updated weekly, capable of continuously absorbing newly emerging error types and patterns, maintaining its effectiveness and applicability.

[0128] During the implementation process, first analyze each field and the overall structure of the recognition result to identify possible complex error patterns. For example, it may be detected that the address information contains inconsistent regional information, such as the province not matching the city; or it may be recognized that the merchant name contains possible misuse of industry abbreviations or professional terms. Then, match these potential error patterns with the cases in the knowledge base to find similar error types and solutions.

[0129] When a match is found, apply the corresponding correction strategy to correct the recognition result. These correction strategies usually not only consider local characters or words but also broader context and semantic relationships. For example, for the correction of address information, it will refer to the geographical database and address format specifications; for the correction of merchant information, it will consider the industry background and common expression methods.

[0130] In addition, the error case knowledge base will also be updated according to the correction results. When a certain error is successfully corrected, the error pattern and its correction strategy will be recorded to enhance the coverage of the knowledge base; if a new type of error that cannot be corrected is encountered, it will also be recorded as a reference for the update of the knowledge base. This dynamic update mechanism ensures that the knowledge base can continuously expand and improve to adapt to newly emerging error types.

[0131] An important feature of the error case knowledge base is its automatic update mechanism. It will regularly analyze newly emerging error cases, extract the patterns and rules from them, and add them to the knowledge base, enabling the knowledge base to continuously expand and improve to adapt to newly emerging error types. This dynamic update mechanism ensures that various error situations can be continuously and effectively processed.

[0132] Finally, the second preliminary calibrated recognition results are output, which have corrected most of the complex error patterns and greatly improved the recognition accuracy. By combining the common error mapping table and the error case knowledge base, comprehensive coverage and accurate correction of various error types are achieved.

[0133] S3.3: According to the second preliminary calibrated recognition results, use the dynamic feedback algorithm to calculate the correct rate index of the current batch, and obtain the calibrated recognition results with confidence scores.

[0134] After two rounds of calibration are completed, it is necessary to quantitatively evaluate the quality of the calibration results for subsequent processing and optimization. This step is achieved through the dynamic feedback algorithm, which can calculate the correct rate index of the current batch and generate a confidence score for each recognition result.

[0135] The dynamic feedback algorithm is an adaptive evaluation mechanism that evaluates performance by analyzing the relationship between input and output. The algorithm first collects the second preliminary calibrated recognition results, and then compares and analyzes them with the original input data to calculate the accuracy and reliability indexes of the processing.

[0136] S3.3.1: According to the second preliminary calibrated recognition results, use the logarithmic function to process the quantity ratio of the input yellow page number data to the output recognition results, and obtain the preliminary performance evaluation score, where the preliminary performance evaluation score reflects the data processing effect of the second preliminary calibrated recognition results.

[0137] In this step, first, a quantitative evaluation of the second preliminary calibrated recognition results is carried out. Specifically, the logarithmic function is used to process the quantity ratio of the input yellow page number data to the output recognition results, and this method can effectively smooth the influence of the data volume difference on the evaluation results.

[0138] The following formula is applied to calculate the preliminary performance evaluation score:

[0139]

[0140] I is the amount of input data;

[0141] O is the amount of data finally returned by the model;

[0142] Pm is the correct rate returned by the model (expressed as a percentage, e.g., 0.8 represents 80%);

[0143] Ph is the correct rate statistically counted manually (expressed as a percentage, e.g., 0.8 represents 80%);

[0144] W1 is the weight of the input data volume;

[0145] W2 is the weight of the output data volume;

[0146] W3 is the weight of the model accuracy;

[0147] W4 is the weight of manual accuracy;

[0148] C is other parameters (which can be set according to actual conditions, such as model running time, data processing rate, etc.);

[0149] Use logarithmic functions to process input and output to avoid excessive impact of large numbers on the results, while retaining the trend of data volume changes; if F is high, it means that the model performance is good, and the current model can be maintained or slightly optimized; if F is low, it is necessary to count various parameters to accurately find the cause of the model error, so as to specifically analyze and optimize the model.

[0150] Based on the overall score, determine whether the results need to be optimized, and based on the scores of each dimension, determine in which aspects the output format needs to be optimized and modified.

[0151] In practice, the system calculates the overall accuracy (F-value) and the scores for various dimensions (such as number format, address format, and merchant name). A high F-value indicates good overall model performance; a low F-value indicates further analysis of the scores for each dimension to identify areas of poor performance and provide guidance for subsequent optimization.

[0152] In addition, the system calculates a confidence score for each recognition result, reflecting the system's confidence in the correctness of the result. The confidence calculation takes into account multiple factors, including the output confidence of the AI model and the matching degree during the calibration process.

[0153] Ultimately, this step outputs calibration recognition results with confidence scores. These results not only contain the calibrated data, but also come with confidence information and performance indicators, providing an important reference for subsequent standardized output and system optimization.

[0154] This calculation takes into account two key factors: processing completeness (i.e., what proportion of the input data was processed) and processing accuracy (i.e., how much of the processed results were correct). These two factors together determine the overall performance.

[0155] The preliminary performance evaluation score directly reflects the effectiveness of the data processing of the second preliminary calibration recognition results. A high score indicates that the system can process most input data with high accuracy. A low score indicates that there may be issues with incomplete data processing or insufficient accuracy, requiring further analysis of the specific causes.

[0156] In this way, the overall quality of the calibration results can be preliminarily evaluated, providing a basis for subsequent multi-dimensional evaluation. At the same time, this evaluation method can also adapt to processing scenarios with different data volumes and types, providing consistent and comparable performance indicators.

[0157] S3.3.2: According to the preliminary performance evaluation score, use a dynamic weight adjustment algorithm to calculate the evaluation indicators for each dimension, and generate a multi-dimensional performance score, where the multi-dimensional performance score includes the correct rate of number format, the correct rate of address format, and the correct rate of merchant information.

[0158] After obtaining the preliminary performance evaluation score, it is necessary to further analyze the performance of each dimension in depth to accurately identify possible problem areas. This step is achieved through a dynamic weight adjustment algorithm, which can dynamically calculate the evaluation indicators for each dimension according to the importance and historical performance of different dimensions.

[0159] First, several key evaluation dimensions are defined, mainly including the correct rate of number format, the correct rate of address format, and the correct rate of merchant information. These dimensions respectively reflect the performance in processing different types of data. For example, the correct rate of number format measures the ability to identify and calibrate the telephone number format; the correct rate of address format evaluates the effect of processing complex address information; and the correct rate of merchant information reflects the processing accuracy of information such as merchant names and descriptions.

[0160] The dynamic weight adjustment algorithm will assign appropriate weights to each dimension according to the importance of these dimensions in the current task and their historical performance. For example, if it is found that number format errors are the most common problems in historical processing, then a higher weight may be given to the correct rate of number format; or if the current task particularly focuses on the accuracy of address information, then the weight of the correct rate of address format may be increased.

[0161] Based on these weights, calculate the evaluation indicator pi for each dimension to form a multi-dimensional performance score. This multi-dimensional score not only provides a comprehensive view of the overall performance but also can accurately locate possible problem areas. For example, it may be found that the overall performance is good, but the correct rate of address format is relatively low, which indicates the direction that needs to be improved first.

[0162] Through this multi-dimensional evaluation, the performance characteristics can be understood more precisely, providing a clear direction and basis for subsequent optimization. At the same time, this evaluation method can also adapt to different types of task requirements, providing task-related performance indicators.

[0163] S3.3.3: Based on the multi-dimensional performance score, the reliability of the recognition result is comprehensively evaluated using a preset confidence calculation model. When the score of any dimension in the multi-dimensional performance is lower than the first preset threshold, the recognition result of the corresponding dimension is recalibrated and the calibrated recognition result with the confidence score is output.

[0164] After obtaining multi-dimensional performance scores, the final evaluation and calibration phase begins. This step uses a preset confidence calculation model to comprehensively evaluate the reliability of the recognition results and, if necessary, perform additional calibration.

[0165] The confidence calculation model is a complex assessment that comprehensively considers multiple factors to evaluate the reliability of each recognition result. These factors include the confidence of the AI model's original output, the degree of match during the calibration process, and performance scores in various dimensions. The model calculates a confidence score for each recognition result, typically expressed as a floating-point number between 0 and 1, reflecting the degree of confidence in the correctness of the result.

[0166] We focus on the scores of each dimension within the multi-dimensional performance. If the score for any dimension falls below a preset threshold, we deem the recognition results for that dimension problematic and require additional calibration. For example, if the address format accuracy rate falls below a threshold, all address information will be recalibrated; or if the number format accuracy rate is insufficient, phone number information will undergo additional verification and correction.

[0167] The recalibration process employs strategies tailored to the specific problem area. For example, for a number format issue, stricter format validation rules might be applied; for an address issue, a geographic database reference check might be added. This targeted calibration effectively addresses accuracy issues in specific areas and improves overall recognition quality.

[0168] Finally, the system outputs calibration recognition results with confidence scores. These results not only include the data from multiple rounds of calibration but also include a score assessing the reliability of each result. This confidence-based output format allows downstream applications to flexibly handle different results based on their reliability levels. For example, high-confidence results can be used directly, while low-confidence results may require manual review or additional verification.

[0169] Through this comprehensive evaluation and targeted calibration mechanism, the accuracy and reliability of recognition results can be significantly improved, while also providing detailed performance indicators and problem location for subsequent optimization.

[0170] Before intelligently correcting the recognition results by matching common error mapping tables and verifying error case knowledge bases, the process also includes establishing an unbounded Gaussian difference private sampling mechanism, including:

[0171] C1: Receive the preliminary standardized AI recognition result, adopt the unbounded Gaussian mechanism and differential privacy protection mechanism, add Gaussian-distributed noise to the yellow page number data, and obtain the data with differential privacy protection.

[0172] Before performing intelligent correction, an unbounded Gaussian differential private sampling mechanism is first implemented, which is an advanced technology aiming to enhance the calibration effect while protecting data privacy. This step first receives the preliminary standardized AI recognition result output from the AI model enhancement layer, and then processes the data through a carefully designed privacy protection mechanism.

[0173] Differential privacy is a widely recognized data protection technology. Its core idea is to ensure that even if an attacker has all information except the target data, it is impossible to determine whether the target data is included in the dataset by adding carefully designed noise to the data or query results. When processing yellow page number data that may contain sensitive information, this protection is particularly important and can effectively prevent the leakage and abuse of personal information.

[0174] Adopt the unbounded Gaussian mechanism as a specific implementation of differential privacy. Compared with the traditional bounded differential privacy mechanism, the unbounded Gaussian mechanism can handle data with an unbounded value range, which makes it particularly suitable for processing various numerical and text information in yellow page number data. This mechanism adds noise obeying the Gaussian distribution to the data and can theoretically provide optimal privacy protection performance while maintaining data availability.

[0175] In specific operations, the privacy budget (i.e., the intensity of privacy protection) will be dynamically adjusted according to the sensitivity of the data and the purpose of use. For example, for data fields containing sensitive information such as personal phone numbers, a higher privacy protection intensity will be applied, and relatively large noise will be added; while for public business information, a relatively low protection intensity may be adopted to obtain better data utility.

[0176] Through this mechanism, the data with differential privacy protection is output. These data not only protect privacy but also maintain sufficient information for subsequent processing. This balance is particularly important for yellow page number processing because it needs to find the best balance between protecting personal privacy and providing useful services.

[0177] C2: According to the data with differential privacy protection, use the gradient descent method or reinforcement learning technology to dynamically adjust the parameters of differential private sampling, optimize the sampling strategy, and output the optimized unbounded Gaussian differential private sampling parameters.

[0178] After obtaining data with differential privacy protection, it is necessary to further optimize the parameters of differential private sampling to achieve the best balance between privacy protection and data utility. This step realizes the dynamic optimization of sampling parameters through gradient descent method or reinforcement learning technology.

[0179] An adaptive algorithm is designed to dynamically adjust the differential private sampling parameters according to data characteristics and usage purposes. This algorithm considers multiple factors, such as data sensitivity, required accuracy, privacy risk, etc., enabling it to find the best balance point between privacy protection and data utility in different situations.

[0180] When using the gradient descent method for parameter optimization, a target function is first defined, which comprehensively measures the privacy protection level and data utility. Then, by iteratively adjusting sampling parameters (such as noise amplitude, distribution parameters, etc.), the parameter combination that makes the target function reach the optimum is gradually found. This gradient-based optimization method can efficiently find the optimal solution in a large parameter space and is suitable for dealing with complex optimization problems.

[0181] Alternatively, reinforcement learning technology may also be used for parameter optimization. In this case, parameter adjustment is regarded as a decision-making process, and the optimal parameter adjustment strategy is learned through interaction with the environment (i.e., data processing). By experimenting with different parameter settings, observing their effects, and adjusting the strategy according to the effects, a strategy that can maximize the long-term reward (i.e., the comprehensive index of privacy protection and data utility) is gradually learned.

[0182] Through these optimization methods, the sampling strategy can be continuously adjusted and improved, so that the added noise can effectively protect privacy without overly affecting the data availability. This optimization is a continuous process that will continuously adjust the strategy according to the feedback of the calibration results to adapt to different types of data and usage scenarios.

[0183] Finally, the optimized unbounded Gaussian differential private sampling parameters are output, which can guide the addition of the most appropriate amount of noise when processing yellow page number data to achieve the best balance between privacy protection and data utility. This carefully adjusted parameter setting is the key foundation for providing high-quality and secure services.

[0184] C3: According to the optimized unbounded Gaussian differential private sampling parameters, track the privacy consumption in each processing step through a privacy budget management mechanism, and automatically adjust the processing strategy of the unbounded Gaussian differential private sampling mechanism when it is detected that the privacy budget is lower than the preset threshold, generating a final calibration processing strategy, and the calibration processing strategy is used to guide the subsequent intelligent correction process.

[0185] After obtaining the optimized unbounded Gaussian differential private sampling parameters, a privacy budget management mechanism is established, which is a key control to ensure the maintenance of the expected privacy protection level during long-term operation. The privacy budget is a finite resource that is gradually consumed as data queries and processing proceed, so it needs to be carefully managed.

[0186] A mathematical model is established to theoretically analyze the privacy protection performance of the entire calibration process, including the quantitative calculation of privacy loss, the application of the privacy protection combination theorem, and the impact of different operations on the privacy budget, etc. These theoretical analyses provide a scientific basis for the management of the privacy budget.

[0187] Based on these theoretical analyses, a privacy budget management mechanism is designed, which can track and control the privacy consumption in each processing step. The amount of noise used in each processing operation and the corresponding privacy loss will be recorded, the total privacy consumption will be calculated cumulatively, and compared with the preset upper limit of the privacy budget.

[0188] When it is detected that the privacy budget has approached or fallen below the preset threshold, an automatic adjustment mechanism will be triggered. Such adjustments may include increasing the amount of noise to enhance privacy protection, restricting the frequency of sensitive operations, adjusting data access permissions, etc. For example, more noise protection may be added to particularly sensitive fields (such as personal phone numbers), or the access frequency to these fields may be reduced.

[0189] Through this dynamic adjustment, while maintaining the basic functions and performance, it can ensure that the privacy protection level is not lower than the preset standard. When the privacy budget is about to be exhausted, it may prompt the need to reset the privacy budget, or switch to a more conservative processing mode to prevent privacy leakage.

[0190] Finally, a final calibration processing strategy is generated, which comprehensively considers the privacy protection requirements and data processing effects, and provides clear guidelines for the subsequent intelligent correction process. This strategy ensures that during the calibration process, both the data quality can be effectively improved and a sufficient privacy protection level can be maintained, achieving the dual goals of security and utility.

[0191] This privacy budget management mechanism and dynamic adjustment strategy are important guarantees for the secure processing of sensitive yellow page data, which enables providing provable privacy protection guarantees and enhances compliance and security when processing sensitive data.

[0192] S4: According to the final consistent calibration and standardized number recognition results, use the performance index calculation model to calculate the overall recognition accuracy rate and scores of each dimension, identify high-frequency error patterns to generate an error pattern classification report, and update the error correction rules and the general AI large model optimization strategy.

[0193] S4 includes:

[0194] S4.1: According to the finally consistency-calibrated and standardized number recognition results, use the performance index calculation model to calculate the overall recognition accuracy rate and the scores of each dimension, and obtain a model performance evaluation report.

[0195] After completing the processing and calibration of the yellow page number data, it enters the first step of the dynamic optimization stage: performance evaluation. This step comprehensively evaluates the processing performance of the model through standardized index calculation and analysis, providing data support for subsequent optimization.

[0196] First, receive the finally consistency-calibrated and standardized number recognition results as the basic data for evaluation. These results have undergone comprehensive processing and calibration, representing the current best output quality. Use the preset performance index calculation model to analyze these results, which includes a series of indicators and algorithms for quantitatively evaluating performance.

[0197] In terms of overall performance evaluation, the overall recognition accuracy rate is calculated, which is the core indicator for measuring the overall processing quality. The overall recognition accuracy rate is obtained by comparing the processing results with the standard answers (or the correct results confirmed through other verification methods) and calculating the proportion of correct recognition. This indicator directly reflects the overall accuracy of processing the yellow page number data.

[0198] In addition to the overall accuracy rate, the scores of each key dimension are also calculated. These dimensions include but are not limited to the number format accuracy rate, address information accuracy, merchant data integrity, etc. Each dimension is quantitatively analyzed through a specially designed evaluation algorithm to generate the corresponding performance scores. This multi-dimensional evaluation enables precise positioning of performance strengths and weaknesses, providing directions for targeted optimization.

[0199] In addition, some auxiliary performance indicators, such as processing speed, resource consumption, stability, etc., are also calculated. These indicators reflect the operating conditions and efficiency of the system from different perspectives.

[0200] By integrating the above indicators and analysis results, a comprehensive model performance evaluation report is generated. This report not only includes the specific values of each indicator, but also includes in-depth content such as historical trend analysis, comparison with peers, and performance fluctuation analysis. Such a comprehensive evaluation report provides a clear performance view for managers and developers, helping them understand the current state and potential problems of the system.

[0201] Most importantly, this performance evaluation report provides a solid data foundation for subsequent error analysis and optimization strategy formulation, ensuring that the optimization work can be carried out based on objective data rather than subjective judgment, improving the accuracy and effectiveness of optimization.

[0202] S4.2: Using the error pattern analysis algorithm, identify the high-frequency error patterns and abnormal situations in the model performance evaluation report, and output the error pattern classification report.

[0203] After obtaining the model performance evaluation report, it enters a deeper error analysis stage. This step uses a dedicated error pattern analysis algorithm to analyze and classify the problems found in the performance evaluation, helping to understand the nature and patterns of the problems.

[0204] First, extract error cases and abnormal situations from the performance evaluation report. These cases represent situations where processing fails or is not ideal enough, and are the key focus for optimization. Then, apply the error pattern analysis algorithm to deeply analyze these cases.

[0205] The error pattern analysis algorithm uses a variety of advanced pattern recognition and clustering techniques, which can identify potential common patterns and rules from seemingly different error cases. For example, the algorithm may find that multiple error cases are related to the processing of specific types of area code formats, or all occur in the processing of address information in a specific context. This pattern recognition ability enables the abstraction of more general error types from specific error cases.

[0206] During the process of identifying patterns, the algorithm pays special attention to high-frequency error patterns, that is, those error types that occur repeatedly. These high-frequency error patterns usually represent relatively common problems in the system, and solving these problems may bring more significant performance improvements. For example, if the number format with special characters is frequently processed incorrectly, then optimizing this specific processing logic may greatly improve the overall accuracy rate.

[0207] In addition to high-frequency errors, the algorithm also identifies abnormal situations, that is, those error types that are not common but may have a significant impact. These abnormal situations may represent potential risks or newly emerging challenges under specific conditions and require special attention.

[0208] Based on the results of pattern recognition and analysis, a detailed error pattern classification report is generated. This report classifies and sorts the identified error patterns according to dimensions such as type, frequency, severity, etc., and provides representative cases and analyses for each type of error. The report also includes content such as the time distribution of error patterns, correlation analysis, and speculation on possible root causes, providing a comprehensive reference for subsequent rule updates and model optimization.

[0209] Through this systematic error analysis, key problems and optimization directions can be refined from a large number of specific errors, making the optimization work more targeted, and improving resource utilization efficiency and optimization effects.

[0210] S4.3: According to the error pattern classification report, adopt a rule library update mechanism to update error correction rules and model optimization strategies.

[0211] After completing the analysis and classification of error patterns, it enters the last step of the dynamic optimization phase: rule update and optimization strategy formulation. This step is a key link in the closed-loop feedback, converting the conclusions obtained from the analysis into specific optimization actions to achieve continuous progress.

[0212] First, based on the error pattern classification report, determine the rules and strategies that need to be updated with priority. This determination of priority takes into account multiple factors, such as the frequency of errors, impact degree, difficulty of solution, and resource limitations, etc. Through this comprehensive consideration, the most valuable optimization direction can be selected under limited resource conditions.

[0213] In terms of the update of error correction rules, a rule library update mechanism is adopted. This mechanism can automatically generate or modify corresponding error correction rules based on the identified error patterns. For example, if it is analyzed that errors frequently occur when processing area codes in a specific format, the relevant area code recognition and verification rules will be updated; or if it is found that a certain type of special character often causes confusion, the character replacement rules will be updated.

[0214] For more complex error patterns, it may be necessary to add new rules or modify the combined logic of existing rules. For example, if it is found that the performance is poor when processing multi-level nested address information, new address parsing rules may need to be designed, or the priority and application order of existing rules may need to be adjusted.

[0215] These rule updates are not limited to the post-processing calibration stage, but may also involve relevant rules in the pre-processing and AI model processing stages, forming a coordinated optimization throughout the whole process.

[0216] In terms of model optimization strategies, targeted model improvement plans will be formulated according to the analysis results. These plans may include various methods such as model retraining, parameter adjustment, and structure modification. For example, if it is found that the model performs poorly on a specific type of data, the training samples of this type of data may need to be increased; or if it is found that certain features have a significant impact on performance, the way the model processes these features may need to be adjusted.

[0217] Detailed optimization suggestions will be generated, including the specific content of optimization, expected effects, implementation plans, etc., for managers and developers to refer to. After these suggestions are reviewed and confirmed, they will be applied to the next iteration.

[0218] Through this closed-loop feedback mechanism, it is possible to continuously learn and improve from actual operations, and continuously enhance its ability and efficiency in processing yellow page number data. This continuous optimization mechanism is the key guarantee for maintaining high performance and adaptability in the long term, enabling it to cope with changing data characteristics and business requirements.

[0219] As Figure 5 shown, the present application also provides a system for consistency calibration and standardized output of AI model recognition results, including a data preprocessing module, an AI model enhancement processing module, a post-processing calibration module, and a dynamic optimization module connected in sequence, where:

[0220] The data preprocessing module is used to receive yellow page number data in a preset format set, and use a format recognition algorithm to identify the source type and format characteristics of the yellow page number data, and use an input format unification algorithm to standardize keyword fields, and output yellow page number data in a standardized format to the AI model enhancement processing module;

[0221] The AI model enhancement processing module is used to generate input instructions for a structured AI model using a stored algorithm based on a prompt template, perform yellow page number data recognition and verification processing through a fine-tuned general AI large model, and convert the processing result into a preliminary standardized AI recognition result that conforms to a preset specification according to a fixed field output format rule, and transmit the preliminary standardized AI recognition result to the post-processing calibration module;

[0222] The post-processing calibration module is used to intelligently correct the recognition result according to the preliminary standardized AI recognition result by using common error mapping table matching and error case knowledge base verification, calculate the correct rate index of the current batch, and obtain a calibrated recognition result with a confidence score, and use a standardized output format conversion to generate a final consistency calibration and standardized number recognition result, and send the final consistency calibration and standardized number recognition result to the dynamic optimization module;

[0223] The dynamic optimization module is used to calculate the overall recognition correct rate and scores of each dimension according to the final consistency calibration and standardized number recognition result by using a performance index calculation model, identify high-frequency error patterns to generate an error pattern classification report, and update error correction rules and general AI large model optimization strategies. An embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the method for consistency calibration and standardized output of AI model recognition results described in the above method embodiments. Among them, the storage medium can be a volatile or non-volatile computer-readable storage medium.

[0224] In addition, an embodiment of the present disclosure further provides a computer program product, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the method for consistency calibration and standardized output of the AI model recognition result provided in any of the above embodiments of the present disclosure. For details, reference may be made to the above method embodiments and will not be elaborated herein.

[0225] Among them, the above computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium, which can be a volatile or non-volatile computer-readable storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0226] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and systems can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In the several embodiments provided by the present disclosure, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the systems or units can be in electrical, mechanical, or other forms.

[0227] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0228] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0229] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0230] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions described in the foregoing embodiments or easily conceive of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method for consistency calibration and standardized output of AI model recognition results, characterized in that: include: Receiving yellow page number data in a preset format set, and using a format recognition algorithm to identify the source type and format characteristics of the yellow page number data, and using an input format unification algorithm to normalize key fields to obtain yellow page number data in a standardized format; Utilize the stored algorithm based on the prompt word template to generate the input instructions of the structured AI model, perform yellow page number data recognition and verification processing through the fine-tuned general AI large model, and convert the processing results into preliminary standardized AI recognition results that meet the preset specifications according to the fixed field output format rules; Based on the preliminary standardized AI recognition results, intelligently correct the recognition results using common error mapping table matching and error case knowledge base verification, calculate the accuracy index of the current batch, output the calibrated recognition results with confidence scores, and generate the final consistent calibrated and standardized number recognition results using standardized output format conversion; Based on the final consistency calibration and standardized number recognition results, the performance indicator calculation model is used to calculate the overall recognition accuracy and the scores of each dimension, identify high-frequency error patterns, generate error pattern classification reports, and update error correction rules and general AI large model optimization strategies; Among them, the recognition results are intelligently corrected by matching common error mapping tables and verifying error case knowledge bases, calculating the accuracy index of the current batch, and outputting the calibrated recognition results with confidence scores, including: Using a preset common error mapping table, the preliminary standardized AI recognition results are matched to identify and correct character confusion and format errors to obtain a first preliminary calibrated recognition result; Based on the recognition result of the first preliminary calibration, perform in-depth verification and correction using the error case knowledge base, update the error case knowledge base according to the verification and correction results, and output the recognition result of the second preliminary calibration; Calculating the accuracy index of the current batch using a dynamic feedback algorithm based on the recognition result of the second preliminary calibration to obtain the calibration recognition result with a confidence score; The method of calculating the accuracy index of the current batch using a dynamic feedback algorithm to obtain the calibration recognition result with a confidence score includes: Based on the recognition result of the second preliminary calibration, processing the ratio of the input yellow pages number data to the output recognition result using a logarithmic function to obtain a preliminary performance evaluation score, wherein the preliminary performance evaluation score reflects the data processing effect of the recognition result of the second preliminary calibration; Based on the preliminary performance evaluation score, a dynamic weight adjustment algorithm is used to calculate the evaluation indicators of each dimension to generate a multi-dimensional performance score, wherein the multi-dimensional performance score includes the number format accuracy, address format accuracy, and merchant information accuracy; Based on the multi-dimensional performance score, the reliability of the recognition result is comprehensively evaluated using a preset confidence calculation model. When the score of any dimension in the multi-dimensional performance is lower than a first preset threshold, the recognition result of the corresponding dimension is recalibrated and the calibrated recognition result with a confidence score is output.

2. The method according to claim 1, characterized in that The method of utilizing the input format unification algorithm to normalize key fields and obtain yellow page number data in a standardized format includes: Based on the source type and format characteristics of the yellow page number data, use the number type standard library to match and verify the numbers in the yellow page number data to obtain structured data with type identification; According to the structured data with type identification, an automated format input template generation tool and a preset format template architecture are used to convert and generate yellow page number data in a standardized format.

3. The method according to claim 1, characterized in that Yellow Pages number data recognition and verification are performed using a fine-tuned general AI model. The results are then converted into preliminary standardized AI recognition results that meet preset specifications based on fixed field output format rules, including: Using the general AI large model to process the yellow pages number data in the standardized format based on the structured AI model input instructions to obtain an original AI recognition result; Based on the original AI recognition result, the preset fixed field output format rules are adopted to perform format standardization processing on the recognition results one by one, and the preliminary standardized AI recognition result is output.

4. The method according to claim 1, wherein The performance indicator calculation model is used to calculate the overall recognition accuracy and the scores of each dimension, identify high-frequency error patterns to generate error pattern classification reports, and update error correction rules and model optimization strategies, including: Based on the final consistency calibration and standardized number recognition results, the performance indicator calculation model is used to calculate the overall recognition accuracy and the scores of each dimension to obtain a model performance evaluation report; Using an error pattern analysis algorithm, identifying high-frequency error patterns and abnormal conditions in the model performance evaluation report, and outputting the error pattern classification report; According to the error pattern classification report, a rule base update mechanism is adopted to update the error correction rules and model optimization strategy.

5. The method according to claim 1, characterized in that Before the yellow pages number data recognition and verification process is performed through a fine-tuned general AI model, it also includes: Build a linear-nonlinear hybrid slot machine learning framework, including: Based on the yellow page number data in the standardized format, a hybrid machine learning framework including a linear model and a nonlinear model is constructed to generate a hybrid processing framework; wherein the linear model is used to process data items that conform to a first preset structural rule, and the nonlinear model is used to process data items that require context understanding; Based on the hybrid processing framework, a temporal attention mechanism is introduced. A temporal feature extractor is used to extract time tags and temporal patterns. An attention calculation unit is used to assign weights to data in different time periods to obtain yellow page number data features with time weights. According to the yellow page number data characteristics with time weight, the UpperConfidenceBound algorithm is used to dynamically adjust the balance between exploring new strategies and using known strategies, and output an optimized processing strategy for yellow page number data recognition.

6. The method according to claim 1, characterized in that Before normalizing the key fields using the input format normalization algorithm, it also includes: Construct an equivariant network preprocessing framework for Markov data, including: Receiving yellow page number data in the preset format set, constructing the data into a Markov process, and generating an equivariant network structure including a Markov transition modeler, an equivariant feature extractor, and a structure consistency verifier using an equivariant network construction algorithm; Based on the equivariant network structure, the generalization boundary of the equivariant network structure when processing yellow page number data is analyzed, and a targeted data enhancement strategy is generated using theoretical generalization error upper bound calculations and empirical test results; According to the data enhancement strategy, the preset complexity index and dependency structure of the yellow page number data are analyzed, the order of the Markov process and the network structure parameters are dynamically adjusted according to the analysis results, and the optimized equivariant network preprocessing model is output.

7. The method according to claim 1, characterized in that Before intelligently correcting the recognition results using common error mapping table matching and error case knowledge base verification, it also includes: Establish an unbounded Gaussian difference private sampling mechanism, including: Receiving the preliminary standardized AI recognition results, using an unbounded Gaussian mechanism and a differential privacy protection mechanism to add Gaussian distributed noise to the yellow page number data to obtain data with differential privacy protection; According to the differentially private data, dynamically adjust the parameters of the differentially private sampling using a gradient descent method or reinforcement learning technology, optimize the sampling strategy, and output the optimized unbounded Gaussian differentially private sampling parameters; Based on the optimized unbounded Gaussian difference private sampling parameters, the privacy consumption in each processing step is tracked through a privacy budget management mechanism. When it is detected that the privacy budget is lower than a preset threshold, the processing strategy of the unbounded Gaussian difference private sampling mechanism is automatically adjusted to generate a final calibration processing strategy. The calibration processing strategy is used to guide the subsequent intelligent correction process.

8. A system for consistency calibration and standardized output of AI model recognition results, characterized in that: It includes a data preprocessing module, an AI model enhancement processing module, a post-processing calibration module, and a dynamic optimization module, which are connected in sequence. The data preprocessing module is used to receive yellow page number data in a preset format set, and use a format recognition algorithm to identify the source type and format characteristics of the yellow page number data, and use an input format unification algorithm to normalize key fields, and output the yellow page number data in a standardized format to the AI model enhancement processing module; The AI model enhancement processing module is used to generate input instructions for a structured AI model using a stored algorithm based on a prompt word template, perform yellow page number data recognition and verification processing using a fine-tuned general AI large model, convert the processing results into preliminary standardized AI recognition results that meet preset specifications based on fixed field output format rules, and transmit the preliminary standardized AI recognition results to the post-processing calibration module; The post-processing calibration module is used to intelligently correct the recognition results based on the preliminary standardized AI recognition results by matching the common error mapping table and verifying the error case knowledge base, calculate the accuracy index of the current batch, and obtain the calibration recognition results with confidence scores, and generate the final consistency calibration and standardized number recognition results using the standardized output format conversion, and send the final consistency calibration and standardized number recognition results to the dynamic optimization module; The dynamic optimization module is used to calculate the overall recognition accuracy and the scores of each dimension based on the final consistency calibration and standardized number recognition results using the performance indicator calculation model, identify high-frequency error patterns to generate error pattern classification reports, and update error correction rules and general AI large model optimization strategies; Among them, the recognition results are intelligently corrected by matching common error mapping tables and verifying error case knowledge bases, calculating the accuracy index of the current batch, and outputting the calibrated recognition results with confidence scores, including: Using a preset common error mapping table, the preliminary standardized AI recognition results are matched to identify and correct character confusion and format errors to obtain a first preliminary calibrated recognition result; Based on the recognition result of the first preliminary calibration, perform in-depth verification and correction using the error case knowledge base, update the error case knowledge base according to the verification and correction results, and output the recognition result of the second preliminary calibration; Calculating the accuracy index of the current batch using a dynamic feedback algorithm based on the recognition result of the second preliminary calibration to obtain the calibration recognition result with a confidence score; The method of calculating the accuracy index of the current batch using a dynamic feedback algorithm to obtain the calibration recognition result with a confidence score includes: Based on the recognition result of the second preliminary calibration, processing the ratio of the input yellow pages number data to the output recognition result using a logarithmic function to obtain a preliminary performance evaluation score, wherein the preliminary performance evaluation score reflects the data processing effect of the recognition result of the second preliminary calibration; Based on the preliminary performance evaluation score, a dynamic weight adjustment algorithm is used to calculate the evaluation indicators of each dimension to generate a multi-dimensional performance score, wherein the multi-dimensional performance score includes the number format accuracy, address format accuracy, and merchant information accuracy; Based on the multi-dimensional performance score, the reliability of the recognition result is comprehensively evaluated using a preset confidence calculation model. When the score of any dimension in the multi-dimensional performance is lower than a first preset threshold, the recognition result of the corresponding dimension is recalibrated and the calibrated recognition result with a confidence score is output.

Citation Information

Patent Citations

  • Data standardization method based on large model

    CN119003583A