Digital Quality Management Method for Real Estate Archives Based on End-to-End Intelligent Verification

By performing intelligent damage analysis and adaptive enhancement processing on real estate archives, the problem of poor quality during the digitization process of archives has been solved, and high-quality and secure digitization of archives has been achieved.

CN120953133BActive Publication Date: 2026-05-26GUANGZHOU AISIKAI INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU AISIKAI INFORMATION TECH CO LTD
Filing Date
2025-09-26
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, the lack of a full-process intelligent verification mechanism in the digitization of real estate archives leads to inconsistent archive quality and makes it difficult to guarantee the authenticity and integrity of digitized archives.

Method used

By loading PDF files labeled with business type to be digitized, damage analysis is performed to identify damage areas and types, adaptive enhancement algorithms are matched, an image enhancement sample set is collected and the proportion of tampered samples is statistically analyzed, and quality enhancement processing is carried out.

Benefits of technology

It significantly improves the processing quality and security of digitized real estate archives, enables accurate identification of damage and enhances adaptability, and ensures the authenticity and integrity of the archives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953133B_ABST
    Figure CN120953133B_ABST
Patent Text Reader

Abstract

This invention discloses a method for quality management of digitized real estate archives based on end-to-end intelligent verification, belonging to the field of digital archive management technology. The method includes: loading the PDF of the archive to be digitized; performing damage analysis on the PDF to obtain damaged areas and damage types; extracting archive data attributes of the damaged areas based on business type labeling; matching named entities of image enhancement algorithms from an enhancement algorithm configuration library based on the damage type; collecting an image enhancement sample set with the archive data attributes and image enhancement algorithm named entities as constraints, and statistically analyzing the proportion of content tampering samples; when the proportion of content tampering samples is less than a tampering sample proportion threshold, performing quality enhancement processing on the damaged areas of the archive according to the image enhancement algorithm named entities. This invention solves the technical problem of poor damage processing effect in the digitization process of real estate archives in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital archives management technology, specifically to a method for digital quality management of real estate archives based on full-process intelligent verification. Background Technology

[0002] With the development of information technology, the digitization of real estate archives has become an important means to improve the efficiency and service level of archive management. Current technologies for digitizing real estate archives typically employ conventional scanning, image processing, and storage processes, lacking a comprehensive intelligent verification mechanism for archive quality. In practice, historical archives may suffer from various types of damage, such as stains, creases, and faded writing. Traditional image enhancement methods often use fixed processing algorithms, failing to adapt to the specific business type and damage characteristics of the archives. This results in inconsistent digitization quality, making it difficult to guarantee the authenticity and integrity of the digitized archives. These shortcomings not only affect the quality of digitized real estate archives but also pose potential risks to the long-term preservation and utilization of the archives. Summary of the Invention

[0003] This application provides a quality management method for the digitization of real estate archives based on full-process intelligent verification, which is used to address the technical problem of poor damage handling in the digitization process of real estate archives in the prior art.

[0004] In view of the above problems, this application provides a method for digital quality management of real estate archives based on end-to-end intelligent verification, the method comprising:

[0005] Load the PDF of the document to be digitized, wherein the PDF of the document to be digitized has a business type label;

[0006] Perform damage analysis on the PDF archive to be digitized to obtain the damaged areas and types of damage.

[0007] Based on the business type labeling, extract the archive data attributes of the damaged archive area;

[0008] Based on the archive damage type, match named entities of image enhancement algorithms from the enhancement algorithm configuration library;

[0009] Using the archive data attributes and the named entities of the image enhancement algorithm as constraints, an image enhancement sample set is collected, and the proportion of samples with content tampering is statistically analyzed.

[0010] When the proportion of tampered samples is less than the threshold for the proportion of tampered samples, the damaged area of ​​the archive is subjected to quality enhancement processing according to the named entity of the image enhancement algorithm.

[0011] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0012] This application proposes a quality management method for digitized real estate archives based on end-to-end intelligent verification. By performing intelligent damage analysis on PDF files to be digitized and accurately identifying damaged areas and types, and then applying corresponding enhancement algorithms for adaptive improvement, the processing quality and security of digitized real estate archives are significantly improved. Compared with traditional methods, the technical solution provided in this application significantly improves the accuracy and adaptability of damage handling, achieving the technical effect of enhancing the security and reliability of the archive processing process. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart illustrating the digital quality management method for real estate archives based on full-process intelligent verification, as provided in an embodiment of this application.

[0015] Figure 2 This is a schematic diagram illustrating the process of performing damage analysis on PDF files to be digitized in the real estate archive digitization quality management method based on full-process intelligent verification provided in the embodiments of this application. Detailed Implementation

[0016] This application provides a quality management method for the digitization of real estate archives based on full-process intelligent verification, which is used to address the technical problem of poor damage handling in the digitization process of real estate archives in the prior art.

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0018] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.

[0019] Example 1, as Figure 1As shown, this application provides a method for digital quality management of real estate archives based on end-to-end intelligent verification, wherein the method includes:

[0020] S10: Load the PDF of the document to be digitized, wherein the PDF of the document to be digitized has a business type label.

[0021] In this embodiment, the first step is to load the PDF file of the solution to be digitized. This PDF file contains a business type label, such as a mortgage contract or a survey report. By loading the PDF file with the business type label at the beginning of the process, this step provides crucial structured context information for all subsequent processing stages.

[0022] S20: Perform damage analysis on the PDF archive to be digitized to obtain the damaged areas and types of damage.

[0023] Traditional real estate record digitization processes often rely on manual inspection to identify and assess physical damage to the records. This method suffers from numerous drawbacks, including inefficiency, subjectivity, and a tendency to miss details. Due to the lack of automated damage detection methods, it is impossible to accurately locate damaged areas on the records, let alone accurately classify the types of damage.

[0024] Step S20 in the method provided in the embodiments of this application, such as Figure 2 As shown, it includes:

[0025] Load a set of damaged real estate archive PDFs, wherein any one of the real estate archive PDFs in the set has a number of labels that identify the damage type and a number of labels that identify the damage area.

[0026] Construct a damage region assessment loss function, wherein the damage region assessment loss function is used to calculate the mean intersection-union ratio of the areas of the segmented damage regions and the labels that identify the damage regions;

[0027] Construct a damage type assessment loss function, wherein the damage type assessment loss function is used to calculate the ratio of the number of accurately assessed damage types to the number of identified damage types;

[0028] Retrieve the PDF set of real estate files, evaluate the loss function based on the damage area and the damage type, and train the damage analysis model using machine learning;

[0029] Damage analysis is performed on the PDF archive to be digitized based on the damage analysis model to obtain the damaged areas and damage types of the archive.

[0030] In this embodiment, a set of damaged real estate archive PDFs is loaded. Each real estate archive PDF in the set has a number of labels that identify the type of damage and a number of labels that identify the damaged area. For example, the original real estate archive may be stained with ink or the text may have faded, making the text unclear. Therefore, the set of damaged real estate archive PDFs is labeled. The labels that identify the type of damage include the type of damage, such as ink stains or faded text, and the labels that identify the damaged area indicate the coordinates of the damaged area.

[0031] A damage region assessment loss function is constructed, which is used to calculate the mean intersection-union ratio (IUU) of the areas of segmented damage regions and the labels identifying the damage regions, thus characterizing the accuracy of damage region assessment. For each damage region, the intersection area of ​​the segmented damage region and the labeled area is calculated, and then divided by their union area, which is the value of the damage region assessment loss function. The closer the value is to 1, the more accurate the damage region segmentation.

[0032] A damage type assessment loss function is constructed. This function calculates the ratio of the number of accurately assessed damage types to the number of identified damage types, thus characterizing the accuracy of the damage type assessment. First, the number of correctly assessed damage types (i.e., the number of damage types whose model assessment matches the actual labeled types) is counted. Then, this number of correctly predicted damage types is divided by the total number of actually existing, correctly labeled damage types in the archive. The resulting ratio is the value of the damage type assessment loss function; a value closer to 1 indicates a more accurate damage type assessment.

[0033] Using a set of real estate archive PDFs, a damage analysis model is trained based on damage area assessment loss functions and damage type assessment loss functions, utilizing machine learning. Specifically, the damage analysis model includes two task-specific output heads: one for damage area segmentation and the other for damage type classification. The weighted sum of the two loss functions is simultaneously optimized using the backpropagation algorithm. The network parameters are continuously adjusted using backpropagation and gradient descent optimization methods until the model converges; for example, training is considered complete when the weighted sum of the two loss functions is greater than or equal to 1.85.

[0034] Using the trained damage analysis model, damage analysis is performed on the PDF archives to be digitized, obtaining the damaged areas and types of damage.

[0035] Through automated damage analysis, the system achieves precise perception and quantitative assessment of the physical condition of archives. It intelligently and efficiently identifies the specific locations and areas of damage within the archives and further accurately determines the specific type of each damage. This transforms archive quality issues into several clearly locatable specific damage problems, providing precise input and clear objectives for the subsequent selection and application of targeted enhancement and repair algorithms.

[0036] S30: Based on the business type label, extract the archive data attributes of the damaged archive area.

[0037] After identifying damaged areas in archives, existing technologies typically focus only on the characteristics of the damage itself, neglecting the data attributes and business importance of the content covered by the damaged area. The value and sensitivity of different data areas, such as signature areas and dates, vary drastically depending on the type of archive. This lack of understanding of the data attributes of damaged areas makes it impossible to distinguish between "critical information damage" and "non-critical information damage" during enhancement, potentially leading to ineffective enhancement of critical information and causing archive security issues.

[0038] In this embodiment, archival data attributes of the damaged area are extracted based on business type labeling. For example, business type labeling may include mortgage contracts, surveying reports, etc., and according to business type standards, the archival data attributes of the damaged area may be contract terms, contract signatures, surveying images, etc.

[0039] By combining business type annotation to extract archival data attributes of the damaged area, a leap from image damage to information value damage is achieved, providing a data foundation for accurately judging the business impact of the damage.

[0040] S40: Based on the file damage type, match image enhancement algorithm named entities from the enhancement algorithm configuration library.

[0041] Existing archival image enhancement processing typically employs fixed, general-purpose algorithms, failing to dynamically select the most suitable specialized processing algorithm based on the type of damage. Furthermore, the sheer number of algorithm options makes it difficult to make the correct selection.

[0042] Step S40 in the method provided in this application embodiment includes:

[0043] The steps for building the enhanced algorithm configuration library include:

[0044] Load the archive PDF enhancement log, wherein the archive PDF enhancement log has a one-to-one corresponding enhancement algorithm named entity and records the archive damage type;

[0045] The enhanced log for loading PDF archives includes:

[0046] Obtain the name entity set of the enhancement algorithm for the initial archive PDF enhancement log, wherein the name entity set of the enhancement algorithm has a one-to-one corresponding name entity set of the mathematical principle of the enhancement algorithm;

[0047] Based on the named entity set of the mathematical principles of the enhancement algorithm, the named entity set of the enhancement algorithm is clustered to obtain several clusters of enhanced algorithm named entities, wherein the several clusters of enhanced algorithm named entities correspond to several named entities of the mathematical principles of the enhancement algorithm;

[0048] Specifically, based on the mathematical principles of the enhancement algorithm, the named entity set is clustered to obtain several clusters of enhanced algorithm named entities, including:

[0049] Extract the first mathematical principle sub-named entity sequence and the second mathematical principle sub-named entity sequence from the mathematical principle named entity set of the enhancement algorithm;

[0050] Based on the first mathematical principle sub-named entity sequence, extract the first enhanced algorithm named entity from the enhanced algorithm named entity set;

[0051] Based on the second mathematical principle sub-named entity sequence, extract the second enhanced algorithm named entity from the enhanced algorithm named entity set;

[0052] When the sequence similarity between the first mathematical principle sub-named entity sequence and the second mathematical principle sub-named entity sequence is greater than or equal to the sequence similarity threshold, the first enhancement algorithm named entity and the second enhancement algorithm named entity are added to the same cluster.

[0053] Otherwise, add the first and second augmentation algorithm named entities to the heterogeneous cluster;

[0054] Based on the mathematical principles of the aforementioned enhancement algorithms, the clusters of enhancement algorithm named entities are reconstructed in the initial archive PDF enhancement log to obtain the archive PDF enhancement log.

[0055] Based on the named entities of the enhancement algorithm and the record file damage type, the support degree between the first record enhancement algorithm and the first record file damage type is set according to the ratio of the number of logs in the record PDF enhancement log statistics and the number of logs in which the first record file damage type and the first record enhancement algorithm appear simultaneously to the total number of logs.

[0056] When the support is greater than or equal to the support threshold, the first record enhancement algorithm is associated with the first record file damage type and stored, and added to the enhancement algorithm configuration library.

[0057] In this embodiment, firstly, an enhancement algorithm configuration library is constructed. The enhancement algorithm configuration library is a database for adaptively configuring enhancement algorithms for different damage types.

[0058] The steps for building the enhanced algorithm configuration library include:

[0059] Load the archive PDF enhancement log, which is a log of historical archive PDF enhancement information, with one-to-one corresponding enhancement algorithm named entities and record archive damage types. For example, for "faded handwriting", the corresponding enhancement algorithm is "histogram equalization".

[0060] Based on the named entity set of the mathematical principles of the enhancement algorithm, the named entity set of the enhancement algorithm is clustered to obtain several clusters of enhanced algorithm named entities, wherein the several clusters of enhanced algorithm named entities correspond to several named entities of the mathematical principles of the enhancement algorithm.

[0061] Specifically, from the named entity set of mathematical principles of the enhancement algorithms, we extract the first mathematical principle sub-named entity sequence and the second mathematical principle sub-named entity sequence. Enhancement algorithms may have different names or originate from different papers, but their core ideas, mathematical foundations, or final effects may indeed exhibit high similarity or equivalence. Each enhancement algorithm is a sequence of multiple mathematical principles combined in a certain order; therefore, the algorithm name is reconstructed based on the combined sequence of its mathematical principles to improve the accuracy of the analysis results. Specifically, the first mathematical principle sub-named entity sequence is a sequence of multiple sub-mathematical principles contained in the first mathematical principle combined in order, and the second mathematical principle sub-named entity sequence is a sequence of multiple sub-mathematical principles contained in the second mathematical principle combined in order.

[0062] Based on the first mathematical principle of the named entity sequence, the first enhancement algorithm named entity is extracted from the enhancement algorithm named entity set. The first enhancement algorithm named entity is an enhancement algorithm randomly extracted from the enhancement algorithm named entity set and is used as the current processing object.

[0063] Similarly, based on the second mathematical principle sub-name entity sequence, the second enhancement algorithm name entity is extracted from the enhancement algorithm name entity set.

[0064] Specifically, the sequence similarity is calculated as the ratio of the number of entities with the same name and index in the first and second named entity sequences to the sequence length. For example, if the sequence length is 20, and there are 8 named entities with the same name and index in both the first and second named entity sequences, then the sequence similarity is 8 ÷ 20 = 0.4. A higher sequence similarity indicates greater similarity between the first and second named entities.

[0065] When the sequence similarity between the first mathematical principle sub-name entity sequence and the second mathematical principle sub-name entity sequence is greater than or equal to a sequence similarity threshold, the first and second augmented algorithm named entities are added to the same cluster. For example, the similarity threshold can be set to 0.6 to filter out similar augmented algorithm named entities.

[0066] Otherwise, add the first and second enhanced algorithm named entities to the heterogeneous cluster.

[0067] Based on several named entities representing the mathematical principles of enhancement algorithms, the named entities of several clusters of enhancement algorithms are reconstructed in the initial archive PDF enhancement log to obtain the archive PDF enhancement log. Specifically, for each record in the initial archive PDF enhancement log, the corresponding algorithm cluster is found based on its original enhancement algorithm named entity, and then the original specific algorithm name is replaced with the general mathematical principle named entity represented by that cluster. For example, "Adaptive Histogram Equalization Algorithm" and "Contrast-Limited Adaptive Histogram Equalization Algorithm" are all uniformly reconstructed into "Histogram Equalization Algorithm" for easier retrieval and use.

[0068] Based on the named entities of the enhancement algorithm and the record archive damage types, the support of the first record enhancement algorithm for the first record archive damage type is defined as the ratio of the number of log entries where the first record archive damage type and the first record enhancement algorithm co-occur in the archive PDF enhancement log statistics to the total number of log entries. For example, the number of times a specific damage type, such as "blurred handwriting," and a specific enhancement algorithm's mathematical principle named entity, such as "histogram equalization algorithm," co-occur in the same record is counted. This number of co-occurrences is then divided by the total number of records in the log; the resulting ratio is the support of the enhancement algorithm's mathematical principle named entity for that damage type. A higher support indicates that the algorithm has been used more frequently in the past to handle this type of damage.

[0069] When the support is greater than or equal to the support threshold, it indicates that the algorithm is suitable for handling this type of damage. The first record augmentation algorithm is associated with the damage type of the first record file and stored, and added to the augmentation algorithm configuration library.

[0070] By accurately matching the corresponding image enhancement algorithm named entities from a pre-built enhancement algorithm configuration library based on the specific damage type identified, the limitations of fixed algorithms are broken. It can automatically match and call verified specific enhancement algorithms that are most effective for different types of damage, thereby improving the professionalism and targeting of image enhancement processing and ensuring that all kinds of damage can be properly processed from the source.

[0071] S50: Using the archive data attributes and the named entities of the image enhancement algorithm as constraints, collect an image enhancement sample set and count the percentage of samples with content tampering.

[0072] Existing technologies lack verification mechanisms for the security and reliability of selected enhancement algorithms under specific business data attributes. Different enhancement algorithms may pose potential risks when processing business data of varying sensitivity, such as signatures, and certain algorithm parameters may unintentionally alter the authenticity of the original information.

[0073] Step S50 in the method provided in this application embodiment includes:

[0074] Using the named entities of the image enhancement algorithm as constraints, image enhancement samples to be analyzed are collected. The image enhancement samples to be analyzed have the location of the archive damage detection record, the sample archive template, and the enhanced quality identifier, which includes a qualified identifier and an unqualified identifier.

[0075] Based on the location of the detected damage record in the archive, extract the archive data attributes of the damaged area from the sample archive template;

[0076] When the archive data attributes of the damaged area of ​​the sample are the same as the archive data attributes, the image enhancement sample to be analyzed is added to the image enhancement sample set.

[0077] When the number of the image enhancement sample set is greater than or equal to the fitting quantity threshold, the proportion of the number of enhanced quality indicators that are unqualified is calculated and set as the content tampering sample proportion.

[0078] In this embodiment, image enhancement samples are collected based on named entities of the image enhancement algorithm. Each image enhancement sample includes the location of the archive damage detection record, a sample archive template, and an enhanced quality identifier, which includes acceptable and unacceptable identifiers. Specifically, the collected image enhancement samples have already been enhanced by the algorithm, and their enhanced quality identifiers are manually determined. For example, "excellent enhancement quality" corresponds to complete repair of archive damage after enhancement, while "extremely poor enhancement quality" corresponds to unrepaired archive content after enhancement.

[0079] Based on the location of the detected damage in the archives, extract the archive data attributes of the damaged area from the sample archive template.

[0080] When the attributes of the damaged area of ​​the sample are the same as those of the archive data, the image enhancement sample to be analyzed is added to the image enhancement sample set.

[0081] When the number of image enhancement sample sets is greater than or equal to the fitted quantity threshold, the proportion of the number of enhanced samples with unqualified quality is set as the content tampering sample proportion, which is used to assess the probability of content tampering that may occur when using the enhancement algorithm to process such data attribute damage.

[0082] By using specific archival data attributes and image enhancement algorithms as dual constraints, collecting historical sample sets and statistically analyzing the proportion of samples whose content has been tampered with, it is possible to quantitatively assess the probability of content tampering risks that may arise when "a certain algorithm is used to process a certain type of sensitive data," thus setting a forward-looking security risk warning for subsequent enhancement operations.

[0083] S60: When the proportion of the content tampered samples is less than the threshold of the proportion of tampered samples, the damaged area of ​​the archive is subjected to quality enhancement processing according to the named entity of the image enhancement algorithm.

[0084] Step S60 in the method provided in this application embodiment further includes:

[0085] When the proportion of the content tampering samples is greater than or equal to the threshold of the tampering sample proportion, based on the business type label and the archive data attributes, the associated archives with the archive data attributes are retrieved, and the archive data attribute feature values ​​are extracted to mark the damaged areas of the archives.

[0086] When the associated file is empty, a manual completion prompt is generated based on the damaged area of ​​the file and the file data attributes;

[0087] The steps for constructing the associated archives include:

[0088] Based on the preset business type, the real estate archive set is matched through the archive group configuration library;

[0089] Cluster the real estate archives with the same attribute to obtain real estate archives with the first data attribute up to the Nth data attribute, store them as related archives of a preset business type, and add them to the related archives library.

[0090] In this embodiment of the application, when the proportion of content tampering samples is less than the threshold for the proportion of tampering samples, for example, when the threshold for the proportion of tampering samples is set to 0.7, it indicates that the risk of the current damaged area archive data attributes is low and the enhancement effect may be good. The risk of the archive content being tampered with after enhancement is low. The archive damaged area is then subjected to quality enhancement processing according to the named entity of the image enhancement algorithm.

[0091] When the proportion of tampered samples is greater than or equal to the threshold for the proportion of tampered samples, the associated archives with archive data attributes are retrieved based on the business type label and archive data attributes.

[0092] The steps for constructing associated files include:

[0093] Based on preset business types, such as mortgage registration and property management, the real estate archive set is matched according to the preset business type through the archive group configuration library.

[0094] Cluster the real estate archives with the same attribute, and obtain the real estate archives with the first data attribute up to the Nth data attribute real estate archives. Store them as related archives of preset business type and add them to the related archives library.

[0095] Furthermore, by retrieving associated archives with archival data attributes and extracting archival data attribute feature values, damaged areas of the archives are marked with warnings. For example, the warning text might include phrases like "Important content, enhanced based on associated content, please verify accuracy," to indicate potential processing risks.

[0096] When the associated file is empty, a manual completion prompt is generated based on the damaged area of ​​the file and the file data attributes, such as "No similar file reference, there is a risk of processing, please process manually".

[0097] By comparing the proportion of content tampering samples with a security threshold, and only automatically executing predetermined enhancement algorithms when the risk is controllable, while introducing manual intervention when the risk is high, a balance between security and image quality is ultimately achieved, ensuring that the final output of high-quality digital images is not only visually clear, but also content-based and reliable.

[0098] In summary, the embodiments of this application have at least the following technical effects:

[0099] This application proposes a quality management method for digitized real estate archives based on end-to-end intelligent verification. By performing intelligent damage analysis and accurately identifying damaged areas and types, and then applying corresponding enhancement algorithms for adaptive enhancement, the processing quality and security of digitized real estate archives are significantly improved. Specifically, by performing intelligent damage analysis on the PDF of the archives to be digitized and accurately identifying damaged areas and types, and extracting archive data attributes by combining business type annotations, the optimal image enhancement algorithm is matched from the enhancement algorithm configuration library. Before implementing enhancement, security verification is performed by statistically analyzing the proportion of content tampering samples. This allows for the automatic selection of the most suitable enhancement algorithm based on different business types and damage characteristics. Simultaneously, a tampering risk monitoring mechanism effectively prevents the risk of content tampering during the digitization process. Compared with traditional methods, the technical solution provided in this application significantly improves the accuracy and adaptability of damage processing, achieving the technical effect of enhancing the security and reliability of the archive processing process.

[0100] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0101] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0102] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. A method for digital quality management of real estate archives based on end-to-end intelligent verification, characterized in that: include: Load the PDF of the document to be digitized, wherein the PDF of the document to be digitized has a business type label; Perform damage analysis on the PDF archive to be digitized to obtain the damaged areas and types of damage. Based on the business type labeling, extract the archive data attributes of the damaged archive area; Based on the archive damage type, match named entities of image enhancement algorithms from the enhancement algorithm configuration library; Using the archive data attributes and the named entities of the image enhancement algorithm as constraints, an image enhancement sample set is collected, and the proportion of samples with content tampering is statistically analyzed. When the proportion of the content tampered samples is less than the threshold of the proportion of tampered samples, the damaged area of ​​the archive is subjected to quality enhancement processing according to the named entity of the image enhancement algorithm. When the proportion of the content tampering samples is greater than or equal to the threshold of the tampering sample proportion, based on the business type label and the archive data attributes, the associated archives with the archive data attributes are retrieved, and the archive data attribute feature values ​​are extracted to mark the damaged areas of the archives. When the associated file is empty, a manual completion prompt is generated based on the damaged area of ​​the file and the file data attributes; Perform damage analysis on the PDF archive to be digitized to obtain the damaged areas and types of damage, including: Load a set of damaged real estate archive PDFs, wherein any one of the real estate archive PDFs in the set has a number of labels that identify the damage type and a number of labels that identify the damage area. Construct a damage region assessment loss function, wherein the damage region assessment loss function is used to calculate the mean intersection-union ratio of the areas of the segmented damage regions and the labels that identify the damage regions; Construct a damage type assessment loss function, wherein the damage type assessment loss function is used to calculate the ratio of the number of accurately assessed damage types to the number of identified damage types; Retrieve the PDF set of real estate files, evaluate the loss function based on the damage area and the damage type, and train the damage analysis model using machine learning; Damage analysis is performed on the PDF archive to be digitized based on the damage analysis model to obtain the damaged areas and damage types of the archive.

2. The method as described in claim 1, characterized in that, The steps for constructing the associated files include: Based on the preset business type, the real estate archive set is matched through the archive group configuration library; Cluster the real estate archives with the same attribute to obtain real estate archives with the first data attribute up to the Nth data attribute, store them as related archives of a preset business type, and add them to the related archives library.

3. The method as described in claim 1, characterized in that, The steps for building the enhanced algorithm configuration library include: Load the archive PDF enhancement log, wherein the archive PDF enhancement log has a one-to-one corresponding enhancement algorithm named entity and records the archive damage type; Based on the named entity of the enhancement algorithm and the record file damage type, the support degree between the first record enhancement algorithm and the first record file damage type is set according to the ratio of the number of logs in the record PDF enhancement log statistics and the number of logs in which the first record file damage type and the first record enhancement algorithm appear simultaneously to the total number of logs. When the support is greater than or equal to the support threshold, the first record enhancement algorithm is associated with the first record file damage type and stored, and added to the enhancement algorithm configuration library.

4. The method as described in claim 3, characterized in that, Load archive PDF enhanced log, including: Obtain the name entity set of the enhancement algorithm for the initial archive PDF enhancement log, wherein the name entity set of the enhancement algorithm has a one-to-one corresponding name entity set of the mathematical principle of the enhancement algorithm; Based on the named entity set of the mathematical principles of the enhancement algorithm, the named entity set of the enhancement algorithm is clustered to obtain several clusters of enhanced algorithm named entities, wherein the several clusters of enhanced algorithm named entities correspond to several named entities of the mathematical principles of the enhancement algorithm; Based on the mathematical principles of the several enhancement algorithms, the several clusters of enhancement algorithm named entities are reconstructed in the initial archive PDF enhancement log to obtain the archive PDF enhancement log.

5. The method as described in claim 4, characterized in that, Based on the mathematical principles of the enhancement algorithm, the named entity set is clustered to obtain several clusters of enhanced algorithm named entities, including: Extract the first mathematical principle sub-named entity sequence and the second mathematical principle sub-named entity sequence from the mathematical principle named entity set of the enhancement algorithm; Based on the first mathematical principle sub-named entity sequence, extract the first enhanced algorithm named entity from the enhanced algorithm named entity set; Based on the second mathematical principle sub-named entity sequence, extract the second enhanced algorithm named entity from the enhanced algorithm named entity set; When the sequence similarity between the first mathematical principle sub-named entity sequence and the second mathematical principle sub-named entity sequence is greater than or equal to the sequence similarity threshold, the first enhanced algorithm named entity and the second enhanced algorithm named entity are added to the same cluster. Otherwise, add the first and second enhanced algorithm named entities to the heterogeneous cluster.

6. The method as described in claim 1, characterized in that, Using the archive data attributes and the named entities of the image enhancement algorithm as constraints, an image enhancement sample set is collected, and the proportion of samples with content tampering is statistically analyzed, including: Using the named entities of the image enhancement algorithm as constraints, image enhancement samples to be analyzed are collected. The image enhancement samples to be analyzed have the location of the archive damage detection record, the sample archive template, and the enhanced quality identifier, which includes a qualified identifier and an unqualified identifier. Based on the location of the detected damage in the archives, extract the archive data attributes of the damaged area from the sample archive template; When the archive data attributes of the damaged area of ​​the sample are the same as the archive data attributes, the image enhancement sample to be analyzed is added to the image enhancement sample set; When the number of the image enhancement sample set is greater than or equal to the fitting quantity threshold, the proportion of the number of enhanced quality indicators that are unqualified is calculated and set as the content tampering sample proportion.