Intellectual property data statistical method and device
By collecting and extracting multidimensional data features of intellectual property, a rule learning model is constructed for multi-level detection, which solves the problems of address indexing errors and caliber mismatch in intellectual property data statistics, and achieves efficient and reliable data correction and statistics.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZHONGZHI SMART TECH CO LTD
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-28
AI Technical Summary
Existing intellectual property data statistical technologies are inadequate in terms of address indexing error correction and caliber adaptation, as well as in terms of applicant statistical flexibility and accuracy, resulting in low statistical efficiency and an inability to meet actual business needs.
By collecting statistical documents, extracting multidimensional data features, constructing a rule learning model, and using a multi-level detection mechanism and detection strategy to reinforce the learning model, the detection process is optimized to achieve data correction and standard adaptation.
It improved the accuracy and efficiency of intellectual property data statistics, solved the problems of address indexing errors and statistical caliber mismatch, ensured the reliability and consistency of data, and supported the accurate business statistical needs.
Smart Images

Figure CN121935653A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data statistical analysis technology, and in particular to a method and apparatus for statistical analysis of intellectual property data. Background Technology
[0002] In intellectual property data statistics, accurate regional attribution and applicant statistics are core requirements supporting actual business operations. Currently, the industry mainly relies on address recognition technology and applicant name normalization processing to achieve relevant statistical goals, but in practical applications, it still faces many technical bottlenecks, as follows:
[0003] I. Address indexing
[0004] Current regional attribution statistics for intellectual property data primarily rely on third-party interfaces (such as map interfaces) to identify applicant addresses, obtain identification results at various regional levels, and then complete aggregate statistics. While address identification technology has matured, two key issues remain in practical business scenarios:
[0005] Address indexing errors and statistical mismatch issues coexist, making traditional processing methods inefficient. On the one hand, address identification results contain explicit errors (such as missing levels, disordered order, non-standard format, redundancy) and implicit errors (such as conflicting administrative affiliations, inconsistent geographic locations, and fictitious addresses), requiring correction. On the other hand, statistical standards are dynamically changing. For example, some scenarios require the address after patent transfer to be used as the basis for attribution, rather than the original applicant's address; or adjustments such as the merger, splitting, or delineation of administrative divisions necessitate re-indexing the original address according to the latest administrative boundaries, causing the original identification results to no longer meet statistical requirements. Traditional methods separate "error correction" and "re-indexing for statistical mismatch" into independent processes, resulting in complex system logic, high maintenance costs, and poor statistical consistency.
[0006] Manual review and correction are inefficient and cannot meet the needs of batch processing. For the aforementioned address indexing issue, the existing solution still relies heavily on manual review and correction, which is not only inefficient but also fails to cover hidden errors in massive amounts of data, thus failing to guarantee statistical accuracy.
[0007] II. Applicant Statistics
[0008] Existing techniques for statistical analysis of applicants primarily focus on normalizing the applicant's name, i.e., unifying different names for the same entity to avoid duplicate or omissions in the statistics. However, this approach has significant limitations:
[0009] The indexing type lacks flexibility and cannot adapt to diverse statistical standards. In actual business operations, statistical standards often need to be adjusted according to the analysis scenario, such as requiring consolidated statistics for group companies and their subsidiaries with equity control relationships, joint statistics for universities and university-run enterprises, or separate statistics based on independent legal entities. Existing normalization processing can only achieve name uniformity and cannot meet the above statistical requirements.
[0010] The lack of a structured relationship management mechanism means that when statistical standards are switched, data needs to be reprocessed manually, which further reduces statistical efficiency and accuracy and makes it difficult to meet dynamically changing business needs.
[0011] In summary, existing intellectual property data statistical technologies have significant shortcomings in address indexing error correction and caliber adaptation, applicant relationship identification and flexible statistics, making it difficult to balance statistical accuracy and processing efficiency, and failing to fully support the accurate statistical needs of actual business. There is an urgent need for a technical solution that can integrate statistical caliber and automatically process address and applicant data. Summary of the Invention
[0012] In a first aspect, embodiments of the present invention provide an intellectual property data statistics method that can integrate statistical standards and automatically process address and applicant data. The method includes:
[0013] Collect statistical documents, determine statistical objectives and business scope, and collect multidimensional intellectual property data based on statistical objectives and business scope;
[0014] Feature extraction is performed on multidimensional intellectual property data to obtain multidimensional intellectual property features;
[0015] Multi-level detection is performed on the extracted multi-dimensional features of intellectual property to obtain detection results;
[0016] A training dataset is constructed based on the multidimensional features of intellectual property rights and the detection results. Based on the training dataset, a rule learning model is generated. The rule learning model includes a rule engine, optimized threshold parameters, and a detection strategy reinforcement learning model.
[0017] After constructing an initial state space based on the multidimensional features of intellectual property, it is input into the detection strategy reinforcement learning model to obtain the optimal detection strategy. According to the optimal detection strategy, the rule engine and optimized threshold parameters are called to perform multi-level detection to obtain optimized detection results. The optimal detection strategy includes the execution order of detection levels, the logical rules used for each level of detection and the corresponding rule execution order, and the optimized threshold parameters used.
[0018] Based on the detection results and optimized detection results, the multidimensional data of intellectual property rights are corrected.
[0019] Secondly, embodiments of the present invention also provide an intellectual property data statistics device capable of integrating statistical standards and automatically processing address and applicant data. The device includes:
[0020] The data acquisition module is used to collect statistical documents, determine statistical objectives and business scope, and collect multidimensional intellectual property data based on statistical objectives and business scope.
[0021] The feature extraction module is used to extract features from multidimensional intellectual property data to obtain multidimensional intellectual property features;
[0022] The multi-level detection module is used to perform multi-level detection on the extracted multi-dimensional features of intellectual property rights to obtain detection results;
[0023] The rule learning module is used to construct a training dataset based on the multidimensional features of intellectual property rights and detection results, and generate a rule learning model based on the training dataset. The rule learning model includes a rule engine, optimized threshold parameters, and a detection strategy reinforcement learning model.
[0024] The optimization detection module is used to construct an initial state space based on the multidimensional features of intellectual property rights and then input it into the detection strategy reinforcement learning model to obtain the optimal detection strategy. According to the optimal detection strategy, the rule engine and optimized threshold parameters are called to perform multi-level detection to obtain the optimized detection result. The optimal detection strategy includes the execution order of detection levels, the logical rules used for each level of detection and the corresponding rule execution order, and the optimized threshold parameters used.
[0025] The calibration module is used to calibrate multidimensional intellectual property data based on the detection results and optimized detection results.
[0026] Thirdly, embodiments of the present invention also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described intellectual property data statistics method.
[0027] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described intellectual property data statistics method.
[0028] Fifthly, embodiments of the present invention also provide a computer program product, the computer program product comprising a computer program, which, when executed by a processor, implements the above-described intellectual property data statistics method.
[0029] The method proposed in this invention first clarifies the statistical objectives and business scope, then collects multidimensional intellectual property data in a targeted manner. This ensures the accuracy and relevance of data collection, avoids interference from irrelevant data, lays a high-quality data foundation for subsequent statistical analysis, and adapts to the needs of different statistical scenarios. Feature extraction is performed on the multidimensional intellectual property data to comprehensively mine features such as address quality, statistics and behavior, text and spatiotemporal relationships, and cross-dimensional correlations. This provides rich and effective judgment criteria for subsequent anomaly detection, improving the comprehensiveness of anomaly identification. A multi-level detection mechanism is used to detect the extracted features, enabling progressive identification of explicit and implicit anomalies in the data, improving the accuracy of anomaly detection, and reducing anomaly omissions or misjudgments. A training dataset is constructed based on the multidimensional features and detection results, generating a rule learning model that includes a rule engine, optimized threshold parameters, and a detection strategy reinforcement learning model. This enables dynamic optimization of detection rules and threshold parameters, enhancing the model's adaptability to complex data scenarios. The optimal detection strategy is output using the detection strategy reinforcement learning model. Multi-level detection is then executed by calling the rule engine and optimized thresholds according to the optimal strategy, optimizing the detection process, improving detection efficiency, and ensuring the consistency and reliability of detection results. By combining the test results with optimized test results to correct the data, the system can accurately correct data errors and solve problems such as address indexing errors, statistical mismatches, and insufficient identification of subject relationships. It balances statistical accuracy and processing efficiency, providing accurate and reliable data support for intellectual property data statistics and fully supporting the accurate statistical needs of actual business operations. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0031] Figure 1 This is a flowchart of the intellectual property data statistics method in an embodiment of the present invention;
[0032] Figure 2 This is a flowchart of the rule-learning model generated in an embodiment of the present invention;
[0033] Figure 3 This is a flowchart illustrating the correction of multidimensional intellectual property data in an embodiment of the present invention;
[0034] Figure 4 This is a schematic diagram of the intellectual property data statistics device in an embodiment of the present invention;
[0035] Figure 5This is a schematic diagram of a computer device in an embodiment of the present invention. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0037] Figure 1 This is a task diagram of the intellectual property data statistics method in an embodiment of the present invention. The method includes:
[0038] Step 101: Collect statistical data, determine statistical objectives and business scope, and collect multidimensional intellectual property data based on statistical objectives and business scope;
[0039] Step 102: Extract features from the multidimensional intellectual property data to obtain multidimensional intellectual property features;
[0040] Step 103: Perform multi-level detection on the extracted intellectual property multidimensional features to obtain the detection results;
[0041] Step 104: Construct a training dataset based on the multidimensional features of intellectual property and the detection results. Based on the training dataset, generate a rule learning model, which includes a rule engine, optimized threshold parameters, and a detection strategy reinforcement learning model.
[0042] Step 105: After constructing an initial state space based on the multidimensional features of intellectual property, input it into the detection strategy reinforcement learning model to obtain the optimal detection strategy. According to the optimal detection strategy, call the rule engine and the optimized threshold parameters to perform multi-level detection to obtain the optimized detection result. The optimal detection strategy includes the execution order of detection levels, the logical rules used for each level of detection and the corresponding rule execution order, and the optimized threshold parameters used.
[0043] Step 106: Based on the detection results and optimized detection results, correct the multidimensional data of intellectual property rights.
[0044] In step 101, statistical caliber documents are collected to determine statistical objectives and business scope, and multidimensional intellectual property data are collected based on the statistical objectives and business scope.
[0045] In this embodiment of the invention, the statistical documents include, but are not limited to, statistical yearbooks and various types of intellectual property office statistical bulletins. The statistical objectives can be predefined regional independent patent statistics, or statistics based on merged administrative divisions after a preset year. The boundaries of the business scope can be derived based on the statistical objectives, specifically including:
[0046] Geographic scope: Determine the administrative / economic regions corresponding to the statistics, and adapt to historical versions of administrative divisions;
[0047] Time range: Define the time interval for statistics and clarify the corresponding patent application / alteration timestamp requirements for that interval;
[0048] Data type range: limited to the types of intellectual property rights (such as patents and trademarks) and related data types (such as address data and entity relationship data).
[0049] Intellectual property multidimensional data includes, but is not limited to, basic address data, spatiotemporal and administrative data, equity relationship data, address verification data, supplementary data on subject characteristics, dynamic caliber-related data, and cross-domain verification data.
[0050] Basic address data includes, but is not limited to: the original address of the patent applicant (the initial address of all applicants within the business scope without processing), the multi-level address identification results generated from preliminary processing, and the address change data after the patent legal status is split. The address change data includes the changed address in scenarios such as transfer, licensing, etc., and the corresponding multi-level indexing results.
[0051] The collection of spatiotemporal and administrative data can support address attribution and spatiotemporal adaptation verification. Specifically, it includes collecting historical versions of administrative divisions, address geographic coordinate (GIS) data (for geospatial consistency verification), and patent change records with timestamps (providing time dimension support for spatiotemporal statistical analysis).
[0052] Collecting equity relationship data can support the identification of applicants' related relationships, including: corporate equity structure data, corporate holding relationship data, and related data of university-run enterprises (providing support for statistical methods such as university-enterprise joint models).
[0053] Collect external verification data, including address verification data (enterprise business registration address database, express logistics address filing database, used for multi-source verification of address authenticity) and auxiliary verification data (patent agency address filing information, to assist in verifying the applicant's address relevance).
[0054] Supplementary data on key characteristics: Applicant industry classification, enterprise size level, historical trend data of patent applications, and cross-regional business registration address, providing multi-dimensional basis for identifying applicant relationships;
[0055] Dynamic statistical caliber-related data: historical records of changes in statistical calibers, comparative data of statistical results under different calibers, and a label library for applicable scenarios of calibers (such as regional development assessment), to achieve accurate mapping between calibers and data;
[0056] Cross-domain verification data: Addresses mentioned in third-party information (such as intellectual property pledge registration addresses) further enhance the comprehensiveness of address authenticity verification.
[0057] After the collection of multidimensional intellectual property data, preprocessing can be performed. The preprocessing steps include:
[0058] Unified data format: Basic address data is organized into a multi-level hierarchical structure, time data is standardized in a unified format, and equity relationship data is stored in a structured manner according to "controlling party - controlled party - shareholding ratio - effective time";
[0059] Basic quality screening: Filter out null values and duplicate data, and mark data with abnormal format (such as missing levels or abnormal equity ratios).
[0060] In step 102, feature extraction is performed on the multidimensional intellectual property data to obtain multidimensional intellectual property features;
[0061] In one embodiment, the multidimensional features of intellectual property rights include at least one of the following: address quality features, statistical and behavioral features, textual and spatiotemporal features, and cross-dimensional correlation features;
[0062] The address quality features include attribution conflict classification features, geographic logical conflict features, and format features; geographic logical conflict features include enclave detection markers and economic zone adaptability; format features include hierarchical integrity coefficients, sequence compliance markers, and redundant character ratios.
[0063] The statistical and behavioral features include a hierarchical combination frequency matrix, a rarity index, a main pattern feature vector, and a cross-cycle stability coefficient.
[0064] The text and spatiotemporal features include text similarity features and spatiotemporal dynamic weight features;
[0065] The cross-dimensional correlation features include subject correlation features, caliber adaptation features, and cross-data source verification features; the subject correlation features include the correlation degree of joint agency and the overlap of patent technology fields; the caliber adaptation features include the applicable time range of statistical caliber and the adaptation requirements of statistical caliber level; the cross-data source verification features include the consistency score of different verification library addresses and the equity relationship matching degree.
[0066] (1) When extracting address quality features, an ERNIE-DPCNN hybrid model is used. The address text (including street, POI, and other information) is input into the model. The semantic association of the address is captured through the knowledge enhancement pre-training capability of ERNIE. Then, the high-dimensional word vectors are compressed by the DPCNN deep convolutional network to generate low-dimensional, highly recognizable address feature vectors (the dimension is compressed from the traditional 512 dimensions to 128 dimensions, while retaining the core semantic information). The feature dimensions include:
[0067] Classification characteristics of conflict attribution (0 - no conflict / 1 - provincial / municipal level conflict / 2 - municipal / district level conflict);
[0068] Geographical logical contradiction characteristics: enclave detection markers (0 - non-enclave / 1 - enclave), economic zone fit (used to represent the overlap rate between economic zone and administrative zone boundaries);
[0069] Format characteristics: Hierarchical integrity coefficient (actual number of levels / standard 4 levels, value 0.25-1), sequence compliance flag (0 - compliant / 1 - non-compliant), redundancy character ratio (number of redundancy characters / total address length).
[0070] (2) When extracting statistical and behavioral features, it includes two parts: address level and applicant level. The feature dimensions of the address level include the hierarchical combination frequency matrix and the rarity index; the feature dimensions of the applicant level include the subject pattern feature vector and the cross-cycle stability coefficient.
[0071] Hierarchical combination frequency matrix: Construct a time series frequency matrix of multi-level address combinations, recording the average monthly frequency of each combination over the past 24 months;
[0072] Rarity index: Calculated using the formula: R = 1 - (frequency of target combination / frequency of highest combination in all data), with a value of 0-1 (the closer to 1, the rarer it is);
[0073] Main pattern feature vector: The three patterns of independent entity, group merger, and school-enterprise cooperation are converted into One-Hot encoding vectors, and the pattern adaptation score is calculated by combining equity relationship data;
[0074] Cross-cycle stability coefficient: The coefficient of variation formula CV = (standard deviation / mean) × 100% is used to measure the degree of fluctuation of the applicant's statistical pattern over the past 3 years (CV < 20% is considered stable).
[0075] (3) When extracting text and spatiotemporal features, a spatiotemporal attention mechanism is introduced. Text and spatiotemporal features include text similarity features and spatiotemporal dynamic features.
[0076] Text similarity features: Semantic similarity scores between addresses and a common error pattern library are calculated based on the Siamese-CNN network (0-1, ≥0.7 is considered high similarity);
[0077] Spatiotemporal dynamic weighting features: Construct a spatiotemporal attention model with a sliding time window, the formula is: W_t = α×e(-β×(t0-t)), where t0 is the base time (latest application date), t is the target data time, α=1.2 and β=0.08 are optimization coefficients, and the hierarchical weighted frequency of each time slice is calculated through this weight.
[0078] (4) Cross-dimensional association features include subject association features (association degree of joint agency, overlap of patent technology fields), caliber adaptation features (application scope of statistical caliber time, adaptation requirements of statistical caliber level), and cross-data source verification features (consistency score of different verification library addresses, matching degree of equity relationship).
[0079] In one embodiment, feature extraction is performed on multidimensional intellectual property data to obtain multidimensional intellectual property features, including:
[0080] Based on basic address data, spatiotemporal and administrative data, and address verification data, address quality features are generated.
[0081] Based on basic address data, spatiotemporal and administrative data, equity relationship data, and supplementary data on main characteristics, statistical and behavioral characteristics are generated through statistical analysis and model calculation.
[0082] Based on basic address data, spatiotemporal and administrative data, text and spatiotemporal features are generated through neural network models and temporal weight calculations.
[0083] Based on equity relationship data, supplementary data on main characteristics, dynamic caliber correlation data, cross-domain verification data, and auxiliary verification data, cross-dimensional correlation features are generated through multi-source correlation calculation.
[0084] In this embodiment of the invention, the address quality features include administrative level contradiction features, geographic logical contradiction features, and format features, which are generated through a combination of model processing and rule calculation.
[0085] When generating administrative level conflict characteristics, the "multi-level address identification results" and "fourth-level indexing results of address change data" are extracted from the basic address data. They are then associated with the "historical versions of administrative divisions" in the spatiotemporal data and the addresses mentioned in third-party information in the address verification data. The multi-level address identification results are compared with the corresponding historical versions of administrative divisions to determine the hierarchical relationship and label them according to the rules. For example: 0 = no conflict, 1 = conflict from level 1 area to level 2 area, 2 = conflict from level 2 area to level 3 area.
[0086] When generating geographic logical contradiction features, the system calls the "original address text" of the basic address data, the "address geographic coordinates (GIS) data" of the spatiotemporal and administrative data, and the "historical version of administrative divisions". When generating enclave detection markers, the system extracts the GIS coordinates (latitude and longitude) of the address, combines them with the regional boundary coordinates in the historical version of the administrative divisions, determines whether the address is located outside its administrative region, and then marks it according to the rules: 0 = non-enclave, 1 = enclave.
[0087] When generating the adaptation degree of special economic zones, the coordinates of the boundary range of special economic zones such as predefined areas and free trade zones are collected, and the percentage of the overlapping area between the address GIS coordinates and the boundary of the special economic zone is calculated. The percentage of the overlapping area is the adaptation degree of the special economic zone (the value is 0-1, and the higher the percentage, the stronger the adaptation degree).
[0088] When generating format features, extract the "original address of the patent applicant" and "multi-level address recognition result" from the basic address data.
[0089] When generating the hierarchical integrity coefficient, the actual number of administrative levels contained in the address text is counted and calculated using the formula: Hierarchical Integrity Coefficient = Actual Number of Levels / Total Number of Levels, with a value range of 0.25-1. When generating the sequence compliance marker, a multi-level address sequence regular expression is constructed to verify whether the order of the multi-level address recognition results conforms to the regular expression rules, and the results are marked according to the rules: 0 = compliant, 1 = non-compliant. When generating the redundant character percentage, redundant characters (irrelevant punctuation, repeated modifiers, meaningless symbols, etc.) are defined and calculated using the formula: Redundant Character Percentage = Number of Redundant Characters / Total Address Length, with a value range of 0-1.
[0090] Statistical and behavioral characteristics include address level (hierarchical combination frequency matrix, rarity index, historical correction preference degree) and applicant level (subject pattern feature vector, cross-cycle stability coefficient), which are generated through statistical analysis and model calculation.
[0091] When generating the hierarchical combination frequency matrix in address-level features, the "multi-level address identification results" of the basic address data and the "time-stamped patent application / change records" (within a time interval limited to the business scope) of the spatiotemporal and administrative data are extracted. Addresses are split according to multi-level address combinations, and the average monthly frequency of each combination over the past 24 months is calculated. A time-series frequency matrix, i.e., the hierarchical combination frequency matrix, is constructed with multi-level address combinations as rows, time slices (months) as columns, and frequency as values. When generating the rarity index, it is calculated according to the formula: R = 1 - (target combination frequency / highest combination frequency in all data), with a value range of 0-1 (the closer to 1, the rarer the combination). When generating historical correction preference, based on historical address correction records (including address features T before correction and correction type C, such as hierarchical missing correction and order reversal correction), the following formula is used to calculate: P(C|T)=[P(T|C)×P(C)] / P(T), where P(C) is the prior probability of a certain type of correction (the proportion of this type in historical corrections), P(T|C) is the probability of the address feature corresponding to this type of correction, and P(T) is the marginal probability of the target address feature; the posterior probability of each type of correction is output, which is the historical correction preference.
[0092] When generating the subject pattern feature vector, equity relationship data (enterprise equity structure, controlling relationship, and school-run enterprise related data) and subject feature supplementary data (applicant industry classification, cross-regional business registration address) are used. Three subject patterns are defined: independent subject, group merger (controlling ratio ≥ 50%), and school-enterprise alliance (including school-run enterprise related data). The three patterns are converted into One-Hot encoded vectors (e.g., independent subject is [1,0,0], group merger is [0,1,0], and school-enterprise alliance is [0,0,1]). The pattern fit score is calculated by combining the equity relationship data (1 point for fit, 0 points for misfit) to form the final subject pattern feature vector.
[0093] When generating the cross-cycle stability coefficient, the "Patent Application Historical Trend Data" (limited to the last 3 years) in the subject feature supplement data is used to statistically analyze the applicant's patent application frequency, address usage frequency and other statistical indicators in the last 3 years. The mean and standard deviation of each indicator are calculated and calculated according to the formula: CV=(standard deviation / mean)×100%, which is the cross-cycle stability coefficient (CV<20% is stable).
[0094] Textual and spatiotemporal features include text similarity features and spatiotemporal dynamic features, which are generated through neural network models and temporal weights.
[0095] When generating text similarity features, historical address error cases (such as missing levels, reversed order, typos, etc.) are collected to construct an annotated error pattern library. Then, semantic similarity scores are generated, including: using the Siamese-CNN network architecture, taking the address text and each pattern in the error pattern library as input pairs, and extracting the feature vectors of the two through a sub-network with shared weights; calculating the cosine similarity between the feature vectors to obtain the semantic similarity score (values range from 0 to 1, ≥0.7 is considered highly similar to the error pattern).
[0096] When generating spatiotemporal dynamic features, the time decay weight is calculated using the "time-stamped patent application / change records" (including the base time t0 = latest application date and target data time t) of spatiotemporal and administrative data and the "multi-level address identification results" of basic address data. This includes substituting preset optimization coefficients (α=1.2, β=0.08) and calculating W_t=α×e (-β×(t0-t)) according to the formula to obtain the decay weight of each time slice (the weight of recent data is higher).
[0097] When generating the hierarchical weighted frequency, a sliding time window (default 24 months) is constructed, and the original frequency of multi-level address combinations in each time slice within the window is statistically analyzed. The hierarchical weighted frequency is calculated using the formula: Hierarchical weighted frequency = Σ(original frequency × W_t of the corresponding time slice), which is the core indicator of spatiotemporal dynamic characteristics.
[0098] Cross-dimensional association features include subject association features, caliber adaptation features, and cross-data source verification features, which are generated through multi-data source association calculations.
[0099] When generating entity association features, the main data used are equity relationship data (corporate equity structure, controlling relationship), auxiliary verification data (patent agency address registration information), and supplementary entity feature data (applicant industry classification, patent application historical trend data).
[0100] When generating the frequency of collaborative applications, the number of times applicants jointly apply for patents within a preset timeframe within the business scope is counted and directly used as the frequency of collaborative applications. When generating the relevance of joint agencies, the number of patent agencies jointly commissioned by the two applicants is counted and divided by the total number of agencies commissioned by both applicants; the result is the relevance of joint agencies (value 0-1, higher values indicate a stronger relevance). When generating the overlap of patent technology fields, the patent technology field classifications of the applicants are extracted (e.g., IPC classification numbers), and the ratio of the number of intersections to the number of unions of the two applicants' technology field classifications is calculated; the ratio is the overlap of patent technology fields (value 0-1, higher values indicate a stronger overlap).
[0101] When generating caliber adaptation features, dynamic caliber-related data (statistical history of caliber changes, comparative data of statistical results of different calibers, and a caliber applicable scenario label library) and spatiotemporal and administrative data (patent application / change records with timestamps) are required.
[0102] When generating the applicable time range of statistical standards, the effective time and expiration time of each standard are extracted from the historical record of statistical standard changes to clarify the time interval; the timestamps of patent application / change records are associated to mark the applicable time range of each data entry.
[0103] When generating hierarchical adaptation requirements, the hierarchical requirements are extracted from statistical documents; combined with historical versions of administrative divisions, the hierarchical requirements are transformed into quantifiable hierarchical identifiers.
[0104] When generating constraints for special scenarios, scenario keywords are extracted from the applicable scenario tag library; the special constraints under each scenario are clarified to form structured constraint conditions.
[0105] When generating cross-data source verification features, it is necessary to use address verification data, cross-domain verification data (addresses mentioned in third-party information), and equity relationship data (corporate equity structure).
[0106] When generating consistency scores for addresses in different verification libraries, the address information of the same applicant in the addresses mentioned in the third-party information is compared, and the number of libraries that completely match is counted; the consistency score is calculated according to the formula: consistency score = number of libraries that completely match / total number of addresses (values range from 0 to 1, with higher values indicating stronger consistency).
[0107] When generating the matching degree of equity relationships across multiple data sources, the enterprise equity structure in the equity relationship data is extracted and compared with the equity records in the third-party database; the number of matching core equity relationships (controlling ratio ≥ 30%) is counted and divided by the total number of core equity relationships to obtain the matching degree (value 0-1).
[0108] In step 103, the extracted multidimensional features of intellectual property are subjected to multi-level detection to obtain the detection results;
[0109] In this embodiment of the invention, the detection result includes the detection process record, the anomaly type, and the anomaly level.
[0110] In one embodiment, the multi-level detection includes a first-level syntax detection; the steps of the first-level syntax detection include:
[0111] Based on the hierarchical integrity coefficient, the system uses hierarchical missing logic rules to determine whether there are hierarchical missing levels in the address data. If so, it marks the anomaly type as hierarchical missing. Combining the multi-level address tree structure, the system uses hierarchical regular expression matching logic rules to identify the missing levels and mark the missing type. The anomaly level is then determined based on the missing type.
[0112] For address data with a hierarchical integrity coefficient of 1, the duplicate hierarchy identification logic rule is used to determine whether there is a duplicate hierarchy. If so, the anomaly type is marked as a duplicate hierarchy, and the anomaly level is determined according to the number of duplicate levels.
[0113] Based on the sequence compliance mark, non-compliant address data is identified through sequence compliance logic rules. For non-compliant address data, the anomaly type is marked as sequence error, and the anomaly level is determined according to the category of sequence error.
[0114] Based on the proportion of redundant characters, the system uses redundancy judgment logic rules to determine whether address data is redundant. If so, the system marks the anomaly type as redundant characters and determines the anomaly level based on the category of redundant characters. If not, the system uses naming convention logic rules to check whether the address data conforms to the preset administrative level naming convention and determines the anomaly level based on the degree of conformity.
[0115] Before the first level of syntax validation, it is necessary to build a multi-level address tree structure, a geographic-specific thesaurus, a hierarchical regular expression library (format matching rules for different administrative levels), and an applicant name standard format library (regular expression templates for enterprise / individual / organization names) in advance, and set judgment threshold parameters: based on business needs, preset syntax error judgment threshold parameters, including hierarchical integrity coefficient threshold parameter (0.5), redundant character proportion threshold parameter (15%), and sequence compliance mark judgment value (1), etc.
[0116] In this embodiment of the invention, during the first-level syntax detection, the hierarchical integrity coefficient of each address data is extracted (from address quality features, with a value of 0.25-1). For example, the hierarchical integrity coefficient = actual number of levels / total number of levels. The hierarchical missing logic rule is that if the hierarchical integrity coefficient < the hierarchical integrity coefficient threshold parameter 0.5 (i.e., actual number of levels < 2 levels), the anomaly type is directly marked as hierarchical missing. When locating missing levels, the missing levels are identified by combining the multi-level address tree structure and hierarchical regular expression matching logic rules, and the missing type (partial missing / complete missing) is recorded. The anomaly level is determined according to the missing type; for example, complete missing is a high level, and partial missing is a low level. For addresses with a hierarchical integrity coefficient of 1, the existence of duplicate levels is further analyzed by the duplicate level identification logic rules. The more duplicate levels there are, the higher the anomaly level.
[0117] During hierarchical order detection, a compliance flag for each address data point (derived from address quality features, 0 = compliant / 1 = non-compliant) is required. This flag is determined by the ERNIE-DPCNN model during the feature extraction stage based on address semantic association. The logical rule for order compliance is that if the compliance flag is 1, the address data is non-compliant, and the anomaly type is marked as an order error. Specifically, it is divided into different categories: hierarchical inversion and hierarchical intersection. The anomaly level for hierarchical inversion can be high, and the anomaly level for hierarchical intersection can be medium. The error location is marked and recorded accordingly.
[0118] Extract the percentage of redundant characters from address data (from address quality features, = number of redundant characters / total address length). The redundancy determination logic rule is: if the percentage of redundant characters > 15% of the redundancy percentage threshold parameter, mark the anomaly type as redundant characters, and determine the anomaly level based on the category of redundant characters (e.g., irrelevant punctuation, repeated modifiers, incorrect symbols, etc.). For address data with a percentage of redundant characters ≤ the redundancy percentage threshold parameter, the naming convention logic rule is to verify whether the address data conforms to the geographic proprietary thesaurus.
[0119] In one embodiment, the first-level syntax detection step further includes:
[0120] The system uses logical rules to determine whether there are integrity errors in the applicant's data. If so, it marks the exception type as an integrity error and determines the exception level based on the degree of the integrity error.
[0121] Based on the subject pattern feature vector, the pattern consistency of the applicant's data is checked through pattern judgment logic rules. If the check fails, the anomaly type is marked as subject pattern format mismatch error, and the anomaly level is determined according to the degree of subject pattern format mismatch.
[0122] In this embodiment of the invention, the format integrity judgment logic rule is as follows: For company applicants, the system checks the enterprise templates in the applicant name standard format library (including regular expressions for suffixes such as "Limited Company", "Joint-Stock Company", "Group") to see if any core identifiers are missing (e.g., "Technology" without the "Limited Company" suffix and without a matching record in the business registration database). For individual applicants, the system verifies whether the name contains complete information and whether there are any meaningless characters or format confusion. When checking institutional applicants, the system can target universities, research institutions, etc., to verify whether the format conforms to "Full Name of Institution + Department (Optional)".
[0123] The pattern determination logic rule for pattern consistency verification of the applicant's main pattern feature vector (derived from statistical and behavioral features, One-Hot encoded vector) is as follows: verify whether the main pattern (independent entity / group merger / university-enterprise joint venture) corresponding to the main pattern feature vector is consistent with the applicant's name format (e.g., if an applicant marked as "university-enterprise joint venture" does not contain related identifiers such as "university" or "university-run enterprise" in its name, and there is no equity relationship data to support it, the anomaly type is marked as "main pattern format mismatch error").
[0124] In one embodiment, the multi-level detection includes a second-level semantic matching detection; the steps of the second-level semantic matching detection include:
[0125] Based on the characteristics of administrative hierarchy contradictions, the hierarchical subordinate relationship verification logic rules are used to verify the hierarchical subordinate relationship of address data. When the verification fails, the anomaly type is marked as a hierarchical contradiction error. The anomaly type is determined according to the degree of semantic error of the hierarchical contradiction.
[0126] Based on the characteristics of geographic logical contradictions, geospatial consistency is checked through geospatial verification logic rules. When the verification fails, the anomaly type is marked as a geospatial verification error, and the anomaly level is determined according to the category of geospatial verification error.
[0127] Based on text similarity features and cross-data source verification features, the address data is verified for authenticity through authenticity verification logic rules, and the abnormal type is marked as authenticity verification error. The abnormality level is determined according to the category of authenticity verification error.
[0128] In this embodiment of the invention, classification features of hierarchical contradictions are extracted, and hierarchical subordinate relationship verification logic rules are established: for address data with administrative hierarchical contradiction features ≥1, the hierarchical subordinate relationship is verified by combining the historical changes of administrative divisions in spatiotemporal and administrative data. If the verification fails, the abnormal type is marked as a semantic error of hierarchical contradiction, and the abnormal type is determined according to the degree of semantic error of hierarchical contradiction.
[0129] In this embodiment of the invention, the geospatial verification logic rules are as follows: Enclave detection markers and economic zone fit are extracted from the geospatial logical contradiction features; based on the enclave detection markers, it is determined whether a geospatial contradiction error exists; for address data marked as 1 for geospatial contradiction errors, the distance between the geographic center point is calculated using GIS data. If the distance between the city-level and district-level center points is >50 kilometers (not within the normal range of an enclave), the verification is deemed unsuccessful, and the anomaly type is marked as a geospatial verification error, with the category being a geospatial contradiction error; if the economic zone fit is <0.6 (low boundary overlap rate), and the statistical caliber requires independent statistics for economic zones (cross-dimensional correlation features → caliber fit features → hierarchical fit requirements), the anomaly type is marked as a geospatial verification error, with the category being an economic zone and administrative zone fit contradiction error.
[0130] In this embodiment of the invention, authenticity verification is performed based on text similarity features and cross-data source verification features. The authenticity verification includes text similarity verification, multi-source verification, and association verification.
[0131] The authenticity verification logic rules include:
[0132] During text similarity verification, text similarity features between the text and spatiotemporal features are extracted. If the score is greater than or equal to the threshold parameter 0.7 (high similarity to the fictitious address pattern), the authenticity verification error category is "authenticity questionable".
[0133] During multi-source verification, cross-data source verification features of cross-dimensional correlation features are extracted and compared with address verification data and cross-domain verification data. If the consistency score is less than the threshold parameter 0.5 (≥3 databases have no matching records), the category of authenticity verification error is address authenticity semantic error.
[0134] During the correlation verification, the patent agency address filing information in the auxiliary verification data is used as a reference. If the applicant's address and the agency's address are not geographically related, the category of authenticity verification error is address correlation semantic error.
[0135] In one embodiment, the second-level semantic matching detection step further includes:
[0136] Based on the subject association features and subject pattern feature vectors, subject relationship semantic verification is performed through subject relationship semantic verification logic rules, and the abnormal type is marked as subject relationship semantic contradiction error. Based on the category of subject relationship semantic contradiction error, the marking judgment logic is determined to be association matching logic.
[0137] In this embodiment of the invention, the semantic verification logic rule for subject relationship is as follows: extract the frequency of applicant cooperation applications and the degree of association of joint agency from the subject association features, and then perform association determination. If the frequency of cooperation between the two applicants is ≥ 5 times and the degree of association of joint agency is ≥ 0.8 (full score 1), but the subject pattern feature vector is not marked as "group merger / university-enterprise joint model", the abnormal type is marked as subject relationship semantic contradiction error; in addition, equity relationship data can be analyzed. If the applicant is a parent company or subsidiary (holding ratio ≥ 50%) but is not associated with the index, the same type of error is supplemented and marked.
[0138] In one embodiment, the multi-level detection includes a third-level spatiotemporal statistical anomaly detection; the steps of the third-level spatiotemporal statistical anomaly detection include:
[0139] Based on the hierarchical combination frequency matrix and spatiotemporal dynamic weight characteristics, the address data is judged to be dynamic probability distribution anomaly by the dynamic probability distribution logic rules. If so, the anomaly type is marked as dynamic probability distribution anomaly, and the anomaly level is determined according to the category of dynamic probability distribution anomaly.
[0140] Based on the rarity index and weighted frequency, the anomaly score of the address data is calculated. The existence of errors in the address data is determined by the spatiotemporal anomaly logic rules. If so, the anomaly type is marked as a spatiotemporal anomaly error, and the anomaly level is determined according to the category of the spatiotemporal anomaly error.
[0141] Based on the cross-cycle stability coefficient of the applicant's data, the system uses stability logic rules to determine whether there are errors in the applicant's data. If so, the anomaly type is marked as an applicant statistical pattern anomaly error, and the anomaly level is determined according to the category of the applicant statistical pattern anomaly error.
[0142] In this embodiment of the invention, a hierarchical combination frequency matrix (statistical and behavioral features → address level, for example, the monthly average frequency of multi-level address combinations in the past 24 months) can be extracted. The dynamic probability distribution logic rule is: based on the formula W_t = α×e (-β×(t0-t)) in the spatiotemporal dynamic weight feature (t0 = latest application date, t = target data time), each time slice is assigned a weight (recent data has a higher weight). For the target address data, the weighted frequency F_w = Σ(F_t×W_t) in the sliding window is calculated by combining the hierarchical combination frequency matrix and the spatiotemporal dynamic weight feature (F_t is the spatiotemporal dynamic weight feature of the t-th time slice).
[0143] When assessing the degree of anomaly based on the rarity index of statistical and behavioral characteristics, the rarity index is extracted (statistical and behavioral characteristics → address level, R = 1 - (target combination frequency / highest combination frequency of all data), 0-1), and the anomaly score is calculated using the formula B = -log(P)×R (P = F_w / average weighted frequency of all data); the spatiotemporal anomaly logic rule is: if B ≥ score threshold parameter 2.5, the anomaly type is marked as a spatiotemporal anomaly error.
[0144] In this embodiment of the invention, the stability logic rule is as follows: if the cross-cycle stability coefficient CV ≥ stability threshold parameter 20% (unstable), it is determined that there is an error in the applicant's data. The reasons for the fluctuation can be analyzed by combining the patent application records with timestamps (such as the applicant suddenly changing the statistical mode in the past year without reasonable business basis), and the abnormal type is marked as an abnormal error in the applicant's statistical mode.
[0145] In one embodiment, the multi-level detection includes a fourth-level cross-dimensional correlation detection; the steps of the fourth-level cross-dimensional correlation detection include:
[0146] Based on the overlap between the applicant's industry classification and the patent technology field in the supplementary data of the main features, the first association logic rule is used to determine whether there is an association anomaly. If so, the anomaly type is marked as address and industry association error, and the anomaly level of address and industry association error is determined according to the cross-regional business filing address in the supplementary data of the main features.
[0147] Based on the correlation between the joint agency and the matching degree of equity relationship, the second association logic rule is used to determine whether there is an association anomaly. If so, the anomaly type is marked as subject relationship and address association error, and the anomaly level of subject relationship and address association error is determined.
[0148] Based on the applicable scope of statistical time and patent change records with timestamps, the time adaptation of statistical time is verified through the third association logic rule. When the verification fails, the anomaly type is marked as data and time adaptation error, and the anomaly level of data and time adaptation error is determined.
[0149] Based on the statistical caliber level adaptation requirements and economic zone adaptation, the caliber level adaptation is verified through the third association logic rule. When the verification fails, the anomaly type is marked as data and caliber level adaptation error, and the anomaly level of data and caliber level adaptation error is determined.
[0150] Based on the address consistency scores of different verification libraries, the address verification data and cross-domain verification data are compared through the fourth association logic rule to perform multi-source address consistency verification. When the verification fails, the anomaly type is marked as cross-data source address contradiction error, and the anomaly level of cross-data source address contradiction error is determined.
[0151] Based on the matching degree of equity relationships in multiple data sources, the consistency of equity relationships across multiple sources is verified through the fifth association logic rule. When the verification fails, the anomaly type is marked as cross-data source equity relationship contradiction error, and the anomaly level of cross-data source equity relationship contradiction error is determined.
[0152] In this embodiment of the invention, when determining whether there is an abnormal correlation between the applicant's industry classification in the multidimensional intellectual property data and the patent technology field in the subject association features, the first association logic rule is as follows: If the applicant is in the agricultural industry (such as planting), but the address is marked as a core business district (combined with GIS data), and there is no overlap in the patent technology field with agricultural technology (overlap < 0.3), the abnormality type is marked as address and industry association error; combined with the cross-regional business registration address in the subject feature supplementary data, if the applicant has no registration in that region and no reasonable change record, the abnormality level of address and industry association error is high level, otherwise it is low level.
[0153] The second association logic rule is as follows: If the association degree of the two applicants' joint agency is ≥ threshold parameter 0.9 (which may be determined according to the actual situation) and the equity relationship matching degree is ≥ threshold parameter 0.8 (which may be determined according to the actual situation) (controlling / participating relationship), but the addresses belong to cross-provincial regions with no business relationship and there is no cross-regional business filing, the abnormal type is marked as incorrect association between the subject relationship and the address.
[0154] During the time-based calibration verification, the applicable scope of the statistical caliber (belonging to the caliber calibration feature in cross-dimensional association features) and the patent change records with timestamps (spatiotemporal and administrative data) are extracted. The third association logic rule is: if the statistical caliber requires statistics to be based on the merged regional division after the first preset year, but the address indexing is incorrect and the timestamp is ≥ the first preset year, the anomaly type is marked as a data and caliber time-based calibration error. Combining the statistical caliber change history in the dynamic caliber association data, if the address corresponds to the historical caliber (before the first preset year) but is not indexed according to the historical regional division, the same type of error is supplemented and marked.
[0155] During the statistical caliber level adaptation verification, the statistical caliber level adaptation requirements (belonging to the caliber adaptation features in cross-dimensional association features, such as independent statistics of predefined areas) and the economic zone adaptation degree (belonging to the geographic logic contradiction features in address quality features) are extracted. The third association logic rule is: if the statistical caliber requirement is "independent statistics of predefined areas", but the address indexing is not indexed according to the independent level, and the economic zone adaptation degree is ≥ threshold parameter 0.8 (which can be determined according to the actual situation, and it is indeed a predefined area), the anomaly type is marked as data and caliber level adaptation error; result comparison: combine the statistical results of different calibers in the dynamic caliber association data to compare the data. If the result of the current indexing statistics deviates from the caliber requirement by ≥ threshold parameter 20% (which can be determined according to the actual situation), the error judgment is strengthened.
[0156] When verifying address consistency across multiple sources, extract the address consistency scores from different verification databases (belonging to the cross-data source verification feature in cross-dimensional association features), and compare them with address verification data and cross-domain verification data; the fourth association logic rule is: if there are ≥3 (which can be determined according to the actual situation) differences in address records among verification databases (such as different house numbers or street names), and there are no address change records (address change data in the basic address data), the verification fails, and the anomaly type is marked as a cross-data source address contradiction error; the patent agency address filing information in the auxiliary verification data can be combined, and if the applicant's address differs greatly from the client address filed by the agency, the same type of error can be additionally marked.
[0157] When verifying the consistency of equity relationships across multiple sources, the matching degree of equity relationships in multiple data sources is extracted (belonging to the cross-data source verification feature in the cross-dimensional association features). The fifth association logic rule is: if the matching degree < matching degree threshold parameter 0.6 (which can be determined according to the actual situation, such as inconsistent core equity relationships, such as a difference in the controlling ratio ≥ 30%), and the main body pattern feature vector is marked as a group merger pattern, the anomaly type is marked as a cross-data source equity relationship contradiction error.
[0158] In step 104, a training dataset is constructed based on the multidimensional features of intellectual property and the detection results. Based on the training dataset, a rule learning model is generated. The rule learning model includes a rule engine, optimized threshold parameters, and a detection strategy reinforcement learning model.
[0159] Figure 2 This is a flowchart illustrating the generation of a rule learning model in one embodiment of the present invention. In one embodiment, a training dataset is constructed based on multidimensional features of intellectual property and detection results. Based on the training dataset, a rule learning model is generated, including:
[0160] Step 201: Combine the multidimensional feature vectors of intellectual property rights, detection results, and manual review results to form a training dataset;
[0161] Step 202: Extract feature condition terms from the training dataset, treating each sample in the training dataset as a transaction, with each sample including multiple feature condition terms;
[0162] Step 203: Mine the association logic rules based on the feature condition items, and use the association logic rules and the current logic rules as the logic rules in the constructed rule engine;
[0163] Step 204: Based on the Bayesian optimization algorithm, optimize the threshold parameters for each detection level to obtain the optimized threshold parameters;
[0164] Step 205: Construct an extended feature set based on the multidimensional feature vector of intellectual property rights;
[0165] Step 206: Model the multi-level detection process as a Markov decision process, perform reinforcement learning training, and obtain a trained detection strategy reinforcement learning model. The state space consists of the intellectual property multi-dimensional feature vector, the extended feature set, and the historical detection path encoding. The execution order of the detection levels, the logical rules used for each level of detection and the corresponding rule execution order, and the threshold parameters are also included.
[0166] Step 207: Combine the rule engine, threshold parameters, and trained detection strategy reinforcement learning model to form a rule learning model.
[0167] In this embodiment of the invention, the multidimensional feature vector of intellectual property is represented as follows: (where m is the total number of features), the detection process record can be represented as follows: ( (This is a record of the detection process for the i-th data at level j). The anomaly type can be represented as: (k is the total number of exception types, This indicates the existence of this type of exception. (Indicates none); This is the result of manual review;
[0168] The extended feature set includes interactive features, temporally derived features, and detection process features. Interactive features are obtained by calculating the interaction relationship between features, such as calculating the hierarchical integrity coefficient × economic zone fit to obtain an interactive feature. Temporally derived features are obtained by calculating the trend of feature changes over time based on the detection process record. Detection process features are the intermediate results of the first-level syntax detection, the second-level semantic matching detection, the third-level spatiotemporal statistical anomaly detection, and the fourth-level cross-dimensional association detection encoded as features, such as the number of syntax detection anomalies and the proportion of semantic detection time.
[0169] In one embodiment, mining association logic rules based on feature condition terms includes:
[0170] Traverse all transactions, count the frequency of each feature condition item, calculate the support, filter feature condition items with support greater than the minimum support, form frequent itemsets, and then use the frequent itemsets as a basis.
[0171] Based on frequent itemsets, generate candidate association logic rules;
[0172] The candidate association logic rules are subjected to triple statistical verification to obtain the association logic rules;
[0173] Conflict resolution is performed on the associated logic rules to obtain all associated logic rules output after conflict resolution.
[0174] In this embodiment of the invention, the feature condition term is to transform the continuous / discrete features in the training dataset into discrete terms that can be used for association logic rules. For example, the enclave detection label (0 / 1) directly uses discrete values as condition terms. Transaction examples are shown in Table 1.
[0175] Table 1
[0176]
[0177] The minimum support (min_support=0.01) can be preset: the frequency of an itemset appearing in all transactions must be ≥1% (i.e., if the total number of transactions is 10,000, the itemset must appear at least 100 times to be considered "frequent"), and the minimum confidence (min_confidence=0.7) can be preset: the probability of the "predecessor → consequent" relationship of subsequent association logic rules must be ≥70% to ensure the validity of the rules.
[0178] In this embodiment of the invention, the frequent itemsets obtained after the first screening of feature condition items with support greater than the minimum support can be called frequent 1-itemsets. Based on the frequent 1-itemsets, candidate 2-itemsets (such as {enclave marker = 1, cross-cycle stability coefficient < 0.15}, {enclave marker = 1, equity relationship matching degree < 0.5}) are generated through a "join operation". The support of the candidate 2-itemsets is calculated, and frequent 2-itemsets with support ≥ min_support are screened. The above process is repeated to gradually generate candidate 3-itemsets, candidate 4-itemsets, until no new frequent k-itemsets can be generated.
[0179] Examples of frequent itemsets obtained: Frequent 2-itemsets: {Enclave tag = 1, cross-cycle stability coefficient < 0.15} (support 0.011), {Enclave tag = 1, equity relationship matching degree < 0.5} (support 0.0105); Frequent 3-itemsets: {Enclave tag = 1, economic zone fit degree > 0.8 (corrected support 0.0102), cross-cycle stability coefficient < 0.15} (support 0.010).
[0180] Candidate association logic rules are generated using the structure "frequent itemsets → anomaly labels", and the confidence of the rules is calculated.
[0181] Example of filtering rules with confidence levels greater than or equal to the minimum confidence level (min_confidence):
[0182] Frequent 3-itemset {enclave tag = 1, economic zone fit > 0.8, cross-cycle stability coefficient < 0.15} → anomaly tag {cross-data source address conflict = 1}, confidence level = 0.78 (satisfies ≥ 0.7);
[0183] Frequent 2-itemsets {enclave tag = 1, equity relationship matching degree < 0.5} → anomaly tag {cross-data source address contradiction = 1}, confidence level = 0.72 (satisfies ≥ 0.7);
[0184] Output the candidate association logic rules that meet the conditions, and then perform triple statistical verification, specifically including:
[0185] Coverage test: (Scope of application of the measurement rules);
[0186] Accuracy test (Measure the accuracy of rule-based predictions);
[0187] Lift test: (P(anomaly) is the prior probability of an anomaly in the dataset, which measures the improvement effect of the rule compared to random guessing).
[0188] Retention conditions:
[0189] .
[0190] Establish a rule conflict detection mechanism to handle two types of conflicts separately:
[0191] Direct conflict: Two rules have highly similar antecedents (similarity ≥ 0.8) but opposite conclusions;
[0192] Implicit conflict: The application of multiple rules in combination leads to contradictory conclusions;
[0193] Conflict resolution strategy: A weighted priority ranking based on confidence and coverage is adopted, using the following formula:
[0194] Final adopted rule = The rule with the highest weighted score is selected for execution.
[0195] In one embodiment, the threshold parameters for each detection level are optimized based on a Bayesian optimization algorithm, including:
[0196] An interval is set for each threshold parameter, and a Gaussian process proxy model is established based on the optimization objective function;
[0197] The threshold parameter at the termination of the iteration is obtained by iteratively solving the Gaussian process surrogate model using the expected improvement criterion.
[0198] In this embodiment of the invention, an optimization objective function can be predefined:
[0199]
[0200] in, For p threshold parameters, For multi-class anomaly detection, the macro F1 score is... (Measure processing efficiency). A value of 0.7 can be chosen (precision weight is higher than efficiency weight).
[0201] For each threshold parameter Set interval The Gaussian process surrogate model can be represented as: , It is a mean function. For the kernel function, the expected improvement criterion is the expected criterion. , The current optimal objective function value is set, and the iteration can be terminated when a preset number of iterations (e.g., 50 times) is reached or the objective function converges (the change is ≤0.001).
[0202] When training a reinforcement learning model for a detection strategy, the multi-level detection process is modeled as a Markov decision process (MDP): after defining the state space S and action space A, the state transition function is determined, and a multi-objective reward function is designed to balance accuracy, efficiency, and consistency.
[0203]
[0204] in: Alternatively, other values can be set to ensure that the sum of the weights is 1.
[0205] Accuracy bonus (+1 for correct anomaly detection, -0.5 for false positives, -1 for false negatives);
[0206] Efficiency rewards ( (The fewer steps, the higher the reward)
[0207] Consistency bonus (+0.2 when the detection result is consistent with the historical data pattern).
[0208] The policy network training uses the Deep Deterministic Policy Gradient (DDPG) algorithm to train the detection policy reinforcement learning model, which is an agent, until the preset training termination condition is met.
[0209] In step 105, after constructing an initial state space based on the multidimensional features of intellectual property, it is input into the detection strategy reinforcement learning model to obtain the optimal detection strategy. According to the optimal detection strategy, the rule engine and optimized threshold parameters are called to perform multi-level detection to obtain optimized detection results. The optimal detection strategy includes the execution order of detection levels, the logical rules used for each level of detection and the corresponding rule execution order, and the optimized threshold parameters used.
[0210] In this embodiment of the invention, the historical detection path encoding is empty in the initial state of the initial state space. The optimal detection strategy includes the execution order of detection levels (levels that do not need to be executed can be skipped), the logical rules used for each level of detection and the corresponding rule execution order (some rules can be skipped), and threshold parameters.
[0211] In step 106, the multidimensional data of intellectual property rights are corrected based on the detection results and the optimized detection results;
[0212] Figure 3 This is a flowchart illustrating the correction of multidimensional intellectual property data in an embodiment of the present invention. In one embodiment, the correction of multidimensional intellectual property data based on detection results and optimized detection results includes:
[0213] Step 301: For each intellectual property multidimensional data, generate structured error labels according to the intellectual property multidimensional features, anomaly types, and anomaly levels associated with the intellectual property multidimensional data;
[0214] Step 302: Based on the correction rules corresponding to the structured error labels, perform correction on each intellectual property multidimensional data.
[0215] In this embodiment of the invention, the core objective of step 106 is to achieve automated and accurate correction of multidimensional intellectual property data based on the anomaly type, anomaly level, and structured error labels output by multi-level detection, through preset correction rules.
[0216] Structured error labels can be represented as:
[0217] Anomaly Dimensions: Address Data (AD) / Applicant Data (SB)
[0218] Anomaly Level: High (H) / Medium (M) / Low (L)
[0219] Below is an example of an address where the hierarchy is missing:
[0220] Detection results: Anomaly type = missing hierarchy, anomaly level = high, associated intellectual property multidimensional features = hierarchy integrity coefficient (0.25).
[0221] Structured error labels: AD_address format_missing level_H_hierarchical integrity coefficient
[0222] Below is an example of a correction rule base, see Table 2.
[0223] Table 2
[0224]
[0225] Based on the matching correction rules, the dependent data source is invoked, and the correction operation is performed in the order of "format first, then semantics; basic first, then correlation" to generate the corrected data. The correction process records "original value - corrected value - correction basis - rule ID" to ensure traceability.
[0226] This invention also proposes an intellectual property data statistics device, the principle of which is similar to the intellectual property data statistics method, and will not be described in detail here.
[0227] Figure 4 This is a schematic diagram of the structure of the intellectual property data statistics device in an embodiment of the present invention, including:
[0228] The data acquisition module 401 is used to collect statistical documents, determine statistical objectives and business scope, and collect multidimensional intellectual property data based on statistical objectives and business scope.
[0229] The feature extraction module 402 is used to extract features from multidimensional intellectual property data to obtain multidimensional intellectual property features;
[0230] The multi-level detection module 403 is used to perform multi-level detection on the extracted multi-dimensional features of intellectual property rights to obtain detection results;
[0231] The rule learning module 404 is used to construct a training dataset based on the multidimensional features of intellectual property rights and the detection results, and to generate a rule learning model based on the training dataset. The rule learning model includes a rule engine, optimized threshold parameters, and a detection strategy reinforcement learning model.
[0232] The optimization detection module 405 is used to construct an initial state space based on the multidimensional features of intellectual property rights and then input it into the detection strategy reinforcement learning model to obtain the optimal detection strategy. According to the optimal detection strategy, the rule engine and the optimized threshold parameters are called to perform multi-level detection to obtain the optimized detection result. The optimal detection strategy includes the execution order of detection levels, the logical rules used for each level of detection and the corresponding rule execution order, and the optimized threshold parameters used.
[0233] The calibration module 406 is used to calibrate multidimensional intellectual property data based on the detection results and optimized detection results.
[0234] In one embodiment, the multidimensional features of intellectual property rights include at least one of the following: address quality features, statistical and behavioral features, textual and spatiotemporal features, and cross-dimensional correlation features;
[0235] The address quality features include attribution conflict classification features, geographic logical conflict features, and format features; geographic logical conflict features include enclave detection markers and economic zone adaptability; format features include hierarchical integrity coefficients, sequence compliance markers, and redundant character ratios.
[0236] The statistical and behavioral features include a hierarchical combination frequency matrix, a rarity index, a main pattern feature vector, and a cross-cycle stability coefficient.
[0237] The text and spatiotemporal features include text similarity features and spatiotemporal dynamic weight features;
[0238] The cross-dimensional correlation features include subject correlation features, caliber adaptation features, and cross-data source verification features; the subject correlation features include the correlation degree of joint agency and the overlap of patent technology fields; the caliber adaptation features include the applicable time range of statistical caliber and the adaptation requirements of statistical caliber level; the cross-data source verification features include the consistency score of different verification library addresses and the equity relationship matching degree.
[0239] In one embodiment, the feature extraction module is used for:
[0240] Based on basic address data, spatiotemporal and administrative data, and address verification data, address quality features are generated.
[0241] Based on basic address data, spatiotemporal and administrative data, equity relationship data, and supplementary data on main characteristics, statistical and behavioral characteristics are generated through statistical analysis and model calculation.
[0242] Based on basic address data, spatiotemporal and administrative data, text and spatiotemporal features are generated through neural network models and temporal weight calculations.
[0243] Based on equity relationship data, supplementary data on main characteristics, dynamic caliber correlation data, cross-domain verification data, and auxiliary verification data, cross-dimensional correlation features are generated through multi-source correlation calculation.
[0244] In one embodiment, the multi-level detection includes a first-level syntax detection; the steps of the first-level syntax detection include:
[0245] Based on the hierarchical integrity coefficient, the system uses hierarchical missing logic rules to determine whether there are hierarchical missing levels in the address data. If so, it marks the anomaly type as hierarchical missing. Combining the multi-level address tree structure, the system uses hierarchical regular expression matching logic rules to identify the missing levels and mark the missing type. The anomaly level is then determined based on the missing type.
[0246] For address data with a hierarchical integrity coefficient of 1, the duplicate hierarchy identification logic rule is used to determine whether there is a duplicate hierarchy. If so, the anomaly type is marked as a duplicate hierarchy, and the anomaly level is determined according to the number of duplicate levels.
[0247] Based on the sequence compliance mark, non-compliant address data is identified through sequence compliance logic rules. For non-compliant address data, the anomaly type is marked as sequence error, and the anomaly level is determined according to the category of sequence error.
[0248] Based on the proportion of redundant characters, the system uses redundancy judgment logic rules to determine whether address data is redundant. If so, the system marks the anomaly type as redundant characters and determines the anomaly level based on the category of redundant characters. If not, the system uses naming convention logic rules to check whether the address data conforms to the preset administrative level naming convention and determines the anomaly level based on the degree of conformity.
[0249] In one embodiment, the first-level syntax detection step further includes:
[0250] The system uses logical rules to determine whether there are integrity errors in the applicant's data. If so, it marks the exception type as an integrity error and determines the exception level based on the degree of the integrity error.
[0251] Based on the subject pattern feature vector, the pattern consistency of the applicant's data is checked through pattern judgment logic rules. If the check fails, the anomaly type is marked as subject pattern format mismatch error, and the anomaly level is determined according to the degree of subject pattern format mismatch.
[0252] In one embodiment, the multi-level detection includes a second-level semantic matching detection; the steps of the second-level semantic matching detection include:
[0253] Based on the characteristics of administrative hierarchy contradictions, the hierarchical subordinate relationship verification logic rules are used to verify the hierarchical subordinate relationship of address data. When the verification fails, the anomaly type is marked as a hierarchical contradiction error. The anomaly type is determined according to the degree of semantic error of the hierarchical contradiction.
[0254] Based on the characteristics of geographic logical contradictions, geospatial consistency is checked through geospatial verification logic rules. When the verification fails, the anomaly type is marked as a geospatial verification error, and the anomaly level is determined according to the category of geospatial verification error.
[0255] Based on text similarity features and cross-data source verification features, the address data is verified for authenticity through authenticity verification logic rules, and the abnormal type is marked as authenticity verification error. The abnormality level is determined according to the category of authenticity verification error.
[0256] In one embodiment, the second-level semantic matching detection step further includes:
[0257] Based on the subject association features and subject pattern feature vectors, subject relationship semantic verification is performed through subject relationship semantic verification logic rules, and the abnormal type is marked as subject relationship semantic contradiction error. Based on the category of subject relationship semantic contradiction error, the marking judgment logic is determined to be association matching logic.
[0258] In one embodiment, the multi-level detection includes a third-level spatiotemporal statistical anomaly detection; the steps of the third-level spatiotemporal statistical anomaly detection include:
[0259] Based on the hierarchical combination frequency matrix and spatiotemporal dynamic weight characteristics, the address data is judged to be dynamic probability distribution anomaly by the dynamic probability distribution logic rules. If so, the anomaly type is marked as dynamic probability distribution anomaly, and the anomaly level is determined according to the category of dynamic probability distribution anomaly.
[0260] Based on the rarity index and weighted frequency, the anomaly score of the address data is calculated. The existence of errors in the address data is determined by the spatiotemporal anomaly logic rules. If so, the anomaly type is marked as a spatiotemporal anomaly error, and the anomaly level is determined according to the category of the spatiotemporal anomaly error.
[0261] Based on the cross-cycle stability coefficient of the applicant's data, the system uses stability logic rules to determine whether there are errors in the applicant's data. If so, the anomaly type is marked as an applicant statistical pattern anomaly error, and the anomaly level is determined according to the category of the applicant statistical pattern anomaly error.
[0262] In one embodiment, the multi-level detection includes a fourth-level cross-dimensional correlation detection; the steps of the fourth-level cross-dimensional correlation detection include:
[0263] Based on the overlap between the applicant's industry classification and the patent technology field in the supplementary data of the main features, the first association logic rule is used to determine whether there is an association anomaly. If so, the anomaly type is marked as address and industry association error, and the anomaly level of address and industry association error is determined according to the cross-regional business filing address in the supplementary data of the main features.
[0264] Based on the correlation between the joint agency and the matching degree of equity relationship, the second association logic rule is used to determine whether there is an association anomaly. If so, the anomaly type is marked as subject relationship and address association error, and the anomaly level of subject relationship and address association error is determined.
[0265] Based on the applicable scope of statistical time and patent change records with timestamps, the time adaptation of statistical time is verified through the third association logic rule. When the verification fails, the anomaly type is marked as data and time adaptation error, and the anomaly level of data and time adaptation error is determined.
[0266] Based on the statistical caliber level adaptation requirements and economic zone adaptation, the caliber level adaptation is verified through the third association logic rule. When the verification fails, the anomaly type is marked as data and caliber level adaptation error, and the anomaly level of data and caliber level adaptation error is determined.
[0267] Based on the address consistency scores of different verification libraries, the address verification data and cross-domain verification data are compared through the fourth association logic rule to perform multi-source address consistency verification. When the verification fails, the anomaly type is marked as cross-data source address contradiction error, and the anomaly level of cross-data source address contradiction error is determined.
[0268] Based on the matching degree of equity relationships in multiple data sources, the consistency of equity relationships across multiple sources is verified through the fifth association logic rule. When the verification fails, the anomaly type is marked as cross-data source equity relationship contradiction error, and the anomaly level of cross-data source equity relationship contradiction error is determined.
[0269] In one embodiment, the rule learning module is used for:
[0270] The training dataset is formed by combining the multidimensional feature vectors of intellectual property rights, detection results, and manual review results.
[0271] Extract feature condition terms from the training dataset, treating each sample in the training dataset as a transaction, with each sample including multiple feature condition terms;
[0272] Based on the feature condition terms, we mine the associated logical rules and use the associated logical rules and the current logical rules as the logical rules in the constructed rule engine;
[0273] Based on the Bayesian optimization algorithm, the threshold parameters for each detection level are optimized to obtain the optimized threshold parameters.
[0274] Based on the multidimensional feature vectors of intellectual property rights, an extended feature set is constructed;
[0275] The multi-level detection process is modeled as a Markov decision process and trained by reinforcement learning to obtain a trained detection strategy reinforcement learning model. The state space consists of the intellectual property multi-dimensional feature vector, the extended feature set and the historical detection path encoding. The execution order of the detection levels, the logical rules used for each level of detection and the corresponding rule execution order, and the threshold parameters are also included.
[0276] The rule engine, threshold parameters, and trained detection strategy reinforcement learning model are combined to form a rule learning model.
[0277] In one embodiment, the rule learning module is used for:
[0278] Traverse all transactions, count the frequency of each feature condition item, calculate the support, filter feature condition items with support greater than the minimum support, form frequent itemsets, and then use the frequent itemsets as a basis.
[0279] Based on frequent itemsets, generate candidate association logic rules;
[0280] The candidate association logic rules are subjected to triple statistical verification to obtain the association logic rules;
[0281] Conflict resolution is performed on the associated logic rules to obtain all associated logic rules output after conflict resolution.
[0282] In one embodiment, the rule learning module is used for:
[0283] An interval is set for each threshold parameter, and a Gaussian process proxy model is established based on the optimization objective function;
[0284] The threshold parameter at the termination of the iteration is obtained by iteratively solving the Gaussian process surrogate model using the expected improvement criterion.
[0285] In one embodiment, the calibration module is used to:
[0286] For each intellectual property multidimensional data, a structured error label is generated according to the intellectual property multidimensional features, anomaly type, and anomaly level associated with the intellectual property multidimensional data;
[0287] Based on the correction rules corresponding to the structured error labels, corrections are performed on each intellectual property multidimensional data.
[0288] The proposed solution in this invention first clarifies the statistical objectives and business scope, then collects multidimensional intellectual property data in a targeted manner. This ensures the accuracy and relevance of data collection, avoids interference from irrelevant data, and lays a high-quality data foundation for subsequent statistical analysis, adapting to the needs of different statistical scenarios. Feature extraction is performed on the multidimensional intellectual property data, comprehensively mining features such as address quality, statistics and behavior, text and spatiotemporal relationships, and cross-dimensional correlations. This provides rich and effective judgment criteria for subsequent anomaly detection, improving the comprehensiveness of anomaly identification. A multi-level detection mechanism is employed to detect the extracted features, enabling progressive identification of explicit and implicit anomalies in the data, improving the accuracy of anomaly detection, and reducing anomaly omissions or misjudgments. A training dataset is constructed based on multidimensional features and detection results, generating a rule learning model that includes a rule engine, optimized threshold parameters, and a detection strategy reinforcement learning model. This enables dynamic optimization of detection rules and threshold parameters, enhancing the model's adaptability to complex data scenarios. The detection strategy reinforcement learning model outputs the optimal detection strategy, and the rule engine and optimized thresholds are invoked according to the optimal strategy to execute multi-level detection, optimizing the detection process, improving detection efficiency, and ensuring the consistency and reliability of detection results. By combining the test results with optimized test results to correct the data, the system can accurately correct data errors and solve problems such as address indexing errors, statistical mismatches, and insufficient identification of subject relationships. It balances statistical accuracy and processing efficiency, providing accurate and reliable data support for intellectual property data statistics and fully supporting the accurate statistical needs of actual business operations.
[0289] This invention also provides a computer device. Figure 5 This is a schematic diagram of a computer device in an embodiment of the present invention. The computer device 500 includes a memory 510, a processor 520, and a computer program 530 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 530, it implements the above-described method.
[0290] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0291] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method.
[0292] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0293] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0294] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0295] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0296] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for statistical analysis of intellectual property data, characterized in that, include: Collect statistical documents, determine statistical objectives and business scope, and collect multidimensional intellectual property data based on statistical objectives and business scope; Feature extraction is performed on multidimensional intellectual property data to obtain multidimensional intellectual property features; Multi-level detection is performed on the extracted multi-dimensional features of intellectual property to obtain detection results; A training dataset is constructed based on the multidimensional features of intellectual property rights and the detection results. Based on the training dataset, a rule learning model is generated. The rule learning model includes a rule engine, optimized threshold parameters, and a detection strategy reinforcement learning model. After constructing an initial state space based on the multidimensional features of intellectual property, it is input into the detection strategy reinforcement learning model to obtain the optimal detection strategy. According to the optimal detection strategy, the rule engine and optimized threshold parameters are called to perform multi-level detection to obtain optimized detection results. The optimal detection strategy includes the execution order of detection levels, the logical rules used for each level of detection and the corresponding rule execution order, and the optimized threshold parameters used. Based on the detection results and optimized detection results, the multidimensional data of intellectual property rights are corrected.
2. The method according to claim 1, characterized in that, The intellectual property multidimensional features include at least one of the following: address quality features, statistical and behavioral features, textual and spatiotemporal features, and cross-dimensional correlation features; The address quality characteristics include attribution conflict classification characteristics, geographic logical conflict characteristics, and format characteristics; Geographical logical contradictions include enclave detection markers and economic zone fit. Formatting features include hierarchical integrity coefficient, sequence compliance markers, and the percentage of redundant characters; The statistical and behavioral features include a hierarchical combination frequency matrix, a rarity index, a main pattern feature vector, and a cross-cycle stability coefficient. The text and spatiotemporal features include text similarity features and spatiotemporal dynamic weight features; The cross-dimensional association features include subject association features, caliber adaptation features, and cross-data source verification features; The main entity association features include the degree of association of joint agencies and the degree of overlap in patent technology fields; the caliber adaptation features include the applicable time range of statistical caliber and the hierarchical adaptation requirements of statistical caliber; the cross-data source verification features include the consistency score of different verification library addresses and the degree of matching of equity relationships.
3. The method according to claim 2, characterized in that, Feature extraction is performed on multidimensional intellectual property data to obtain multidimensional intellectual property features, including: Based on basic address data, spatiotemporal and administrative data, and address verification data, address quality features are generated. Based on basic address data, spatiotemporal and administrative data, equity relationship data, and supplementary data on subject characteristics, statistical and behavioral characteristics are generated through statistical analysis and model calculation. Based on basic address data, spatiotemporal and administrative data, text and spatiotemporal features are generated through neural network models and temporal weight calculations. Based on equity relationship data, supplementary data on main characteristics, dynamic caliber correlation data, cross-domain verification data, and auxiliary verification data, cross-dimensional correlation features are generated through multi-source correlation calculation.
4. The method according to claim 2, characterized in that, The multi-level detection includes a first-level syntax detection; the steps of the first-level syntax detection include: Based on the hierarchical integrity coefficient, the system uses hierarchical missing logic rules to determine whether there are hierarchical missing levels in the address data. If so, it marks the anomaly type as hierarchical missing. Combining the multi-level address tree structure, the system uses hierarchical regular expression matching logic rules to identify the missing levels and mark the missing type. The anomaly level is then determined based on the missing type. For address data with a hierarchy integrity coefficient of 1, the duplicate hierarchy identification logic rule is used to determine whether there is a duplicate hierarchy. If so, the anomaly type is marked as a duplicate hierarchy, and the anomaly level is determined according to the number of duplicate levels. Based on the sequence compliance mark, non-compliant address data is identified through sequence compliance logic rules. For non-compliant address data, the anomaly type is marked as sequence error, and the anomaly level is determined according to the category of sequence error. Based on the proportion of redundant characters, the system uses redundancy judgment logic rules to determine whether address data is redundant. If so, the system marks the anomaly type as redundant characters and determines the anomaly level based on the category of redundant characters. If not, the system uses naming convention logic rules to check whether the address data conforms to the preset administrative level naming convention and determines the anomaly level based on the degree of conformity.
5. The method according to claim 4, characterized in that, The steps of the first-level syntax check also include: The system uses logical rules to determine whether there are integrity errors in the applicant's data. If so, it marks the exception type as an integrity error and determines the exception level based on the degree of the integrity error. Based on the subject pattern feature vector, the pattern consistency of the applicant's data is checked through pattern judgment logic rules. If the check fails, the anomaly type is marked as subject pattern format mismatch error, and the anomaly level is determined according to the degree of subject pattern format mismatch.
6. The method according to claim 2, characterized in that, The multi-level detection includes a second-level semantic matching detection; The steps of the second-level semantic matching detection include: Based on the characteristics of administrative hierarchy contradictions, the hierarchical subordinate relationship verification logic rules are used to verify the hierarchical subordinate relationship of address data. When the verification fails, the anomaly type is marked as a hierarchical contradiction error. The anomaly type is determined according to the degree of semantic error of the hierarchical contradiction. Based on the characteristics of geographic logical contradictions, geospatial consistency is checked through geospatial verification logic rules. When the verification fails, the anomaly type is marked as a geospatial verification error, and the anomaly level is determined according to the category of geospatial verification error. Based on text similarity features and cross-data source verification features, the address data is verified for authenticity through authenticity verification logic rules, and the abnormal type is marked as authenticity verification error. The abnormality level is determined according to the category of authenticity verification error.
7. The method according to claim 6, characterized in that, The steps of the second-level semantic matching detection also include: Based on the subject association features and subject pattern feature vectors, subject relationship semantic verification is performed through subject relationship semantic verification logic rules, and the abnormal type is marked as subject relationship semantic contradiction error. Based on the category of subject relationship semantic contradiction error, the marking judgment logic is determined to be association matching logic.
8. The method according to claim 2, characterized in that, The multi-level detection includes a third-level spatiotemporal statistical anomaly detection; the steps of the third-level spatiotemporal statistical anomaly detection include: Based on the hierarchical combination frequency matrix and spatiotemporal dynamic weight characteristics, the address data is judged to be dynamic probability distribution anomaly by the dynamic probability distribution logic rules. If so, the anomaly type is marked as dynamic probability distribution anomaly, and the anomaly level is determined according to the category of dynamic probability distribution anomaly. Based on the rarity index and weighted frequency, the anomaly score of the address data is calculated. The existence of errors in the address data is determined by the spatiotemporal anomaly logic rules. If so, the anomaly type is marked as a spatiotemporal anomaly error, and the anomaly level is determined according to the category of the spatiotemporal anomaly error. Based on the cross-cycle stability coefficient of the applicant's data, the system uses stability logic rules to determine whether there are errors in the applicant's data. If so, the anomaly type is marked as an applicant statistical pattern anomaly error, and the anomaly level is determined according to the category of the applicant statistical pattern anomaly error.
9. The method according to claim 2, characterized in that, The multi-level detection includes a fourth-level cross-dimensional correlation detection; the steps of the fourth-level cross-dimensional correlation detection include: Based on the overlap between the applicant's industry classification and the patent technology field in the supplementary data of the main features, the first association logic rule is used to determine whether there is an association anomaly. If so, the anomaly type is marked as address and industry association error, and the anomaly level of address and industry association error is determined according to the cross-regional business filing address in the supplementary data of the main features. Based on the correlation between the joint agency and the matching degree of equity relationship, the second association logic rule is used to determine whether there is an association anomaly. If so, the anomaly type is marked as subject relationship and address association error, and the anomaly level of subject relationship and address association error is determined. Based on the applicable scope of statistical time and patent change records with timestamps, the time adaptation of statistical time is verified through the third association logic rule. When the verification fails, the anomaly type is marked as data and time adaptation error, and the anomaly level of data and time adaptation error is determined. Based on the statistical caliber level adaptation requirements and economic zone adaptation, the caliber level adaptation is verified through the third association logic rule. When the verification fails, the anomaly type is marked as data and caliber level adaptation error, and the anomaly level of data and caliber level adaptation error is determined. Based on the address consistency scores of different verification libraries, the address verification data and cross-domain verification data are compared through the fourth association logic rule to perform multi-source address consistency verification. When the verification fails, the anomaly type is marked as cross-data source address contradiction error, and the anomaly level of cross-data source address contradiction error is determined. Based on the matching degree of equity relationships in multiple data sources, the consistency of equity relationships across multiple sources is verified through the fifth association logic rule. When the verification fails, the anomaly type is marked as cross-data source equity relationship contradiction error, and the anomaly level of cross-data source equity relationship contradiction error is determined.
10. The method according to claim 1, characterized in that, A training dataset is constructed based on the multidimensional features of intellectual property rights and the detection results. Based on this training dataset, a rule learning model is generated, including: The training dataset is formed by combining the multidimensional feature vectors of intellectual property rights, detection results, and manual review results. Extract feature condition terms from the training dataset, treating each sample in the training dataset as a transaction, with each sample including multiple feature condition terms; Based on the feature condition terms, we mine the associated logical rules and use the associated logical rules and the current logical rules as the logical rules in the constructed rule engine; Based on the Bayesian optimization algorithm, the threshold parameters for each detection level are optimized to obtain the optimized threshold parameters. Based on the multidimensional feature vectors of intellectual property rights, an extended feature set is constructed; The multi-level detection process is modeled as a Markov decision process and trained by reinforcement learning to obtain a trained detection strategy reinforcement learning model. The state space consists of the intellectual property multi-dimensional feature vector, the extended feature set and the historical detection path encoding. The execution order of the detection levels, the logical rules used for each level of detection and the corresponding rule execution order, and the threshold parameters are also included. The rule engine, threshold parameters, and trained detection strategy reinforcement learning model are combined to form a rule learning model.
11. The method according to claim 10, characterized in that, Based on feature condition terms, we can mine association logic rules, including: Traverse all transactions, count the frequency of each feature condition item, calculate the support, filter feature condition items with support greater than the minimum support, form frequent itemsets, and then use the frequent itemsets as a basis. Based on frequent itemsets, generate candidate association logic rules; The candidate association logic rules are subjected to triple statistical verification to obtain the association logic rules; Conflict resolution is performed on the associated logic rules to obtain all associated logic rules output after conflict resolution.
12. The method according to claim 10, characterized in that, Based on the Bayesian optimization algorithm, the threshold parameters for each detection level are optimized, including: An interval is set for each threshold parameter, and a Gaussian process proxy model is established based on the optimization objective function; The threshold parameter at the termination of the iteration is obtained by iteratively solving the Gaussian process surrogate model using the expected improvement criterion.
13. The method according to claim 1, characterized in that, Based on the detection results and optimized detection results, the multidimensional data of intellectual property rights are corrected, including: For each intellectual property multidimensional data, a structured error label is generated according to the intellectual property multidimensional features, anomaly type, and anomaly level associated with the intellectual property multidimensional data; Based on the correction rules corresponding to the structured error labels, corrections are performed on each intellectual property multidimensional data.
14. An intellectual property data statistics device, characterized in that, include: The data acquisition module is used to collect statistical documents, determine statistical objectives and business scope, and collect multidimensional intellectual property data based on statistical objectives and business scope. The feature extraction module is used to extract features from multidimensional intellectual property data to obtain multidimensional intellectual property features; The multi-level detection module is used to perform multi-level detection on the extracted multi-dimensional features of intellectual property rights to obtain detection results; The rule learning module is used to construct a training dataset based on the multidimensional features of intellectual property and the detection results, and to generate a rule learning model based on the training dataset. The rule learning model includes a rule engine, optimized threshold parameters, and a detection strategy reinforcement learning model. The optimization detection module is used to construct an initial state space based on the multidimensional features of intellectual property rights and then input it into the detection strategy reinforcement learning model to obtain the optimal detection strategy. According to the optimal detection strategy, the rule engine and optimized threshold parameters are called to perform multi-level detection to obtain the optimized detection result. The optimal detection strategy includes the execution order of detection levels, the logical rules used for each level of detection and the corresponding rule execution order, and the optimized threshold parameters used. The calibration module is used to calibrate multidimensional intellectual property data based on the detection results and optimized detection results.
15. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 13.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 13.
17. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 13.