Multi-dimensional credibility dynamic evaluation method based on automatic information processing system

CN122838696APending Publication Date: 2026-09-29SHANDONG XUNLAI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610980898.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

1、现有方案中,部分仅依赖来源域名信誉、发布时间等单一维度进行可信度判断,未综合考量作者权威性、内容语义质量、传播链路可信度等多维度要素,在权威站点转载低质内容、新发布权威信息暂未积累传播数据等场景下极易出现评估误判,准确性难以保障,另一部分虽采用多维度评估思路,但维度设计缺乏系统性,且各维度的特征提取、量化评分采用串行处理方式,单条信息需依次完成所有维度计算后才输出结果,评估效率难以支撑大规模信息批量处理场景,无法兼顾评估准确性与处理效率

Benefits of technology

1.本申请通过从来源可信度、信息新鲜度、作者权威性、内容质量及URL规范性五个维度对网络信息进行可信度评估,并采用并行化处理架构同时对五个维度的原始特征进行提取和评分,有效解决现有多维度评估方案中因串行处理导致评估效率低下以及单一维度评估易产生误判的技术问题,在保证评估准确性的同时提升大规模信息处理场景下的评估效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838696A_ABST
    Figure CN122838696A_ABST
Patent Text Reader

Abstract

The application discloses a multi-dimensional credibility dynamic evaluation method based on an automatic information processing system, belongs to the technical field of computers, acquires metadata of an information item to be evaluated, synchronously extracts original features of five dimensions of source credibility, information freshness, author authority, content quality and URL standardization from the metadata, generates a structured data set through data standardization, configures independent scoring engines for the five dimensions, each scoring engine acquires standardized features of the corresponding dimension in parallel and outputs a score value of each dimension, acquires the score values of the dimensions, configures weight coefficients for the dimensions, reads data quality identifiers of the original features of the dimensions and calculates confidence factors of the dimensions according to the data quality identifiers, adopts weighted summation with confidence correction to calculate a final comprehensive credibility total score, compares the final comprehensive credibility total score with preset double thresholds, outputs learning instructions, review instructions or rejection instructions according to a comparison result and sends the instructions to a downstream processing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, specifically relating to a multi-dimensional dynamic credibility evaluation method based on an automated information processing system. Background Technology

[0002] With the rapid development of internet technology, online information resources are growing explosively. Various artificial intelligence knowledge bases, content aggregation platforms, and public opinion monitoring systems need to continuously capture information from public network channels to complete data iteration and knowledge updates. Due to the uneven quality of information on the open network, false information, outdated content, low-quality marketing copy, and falsely reprinted content are mixed in. If these are directly included in the knowledge base or business system without verification, it is easy to cause knowledge base pollution, leading to a decline in the performance of downstream AI models and distortion of decision-making basis. Therefore, efficient and accurate automated information credibility assessment technology has become a core requirement for information quality control.

[0003] Currently, existing network information credibility assessment schemes have the following shortcomings in their core assessment mechanisms: 1. In existing solutions, some rely solely on a single dimension such as the reputation of the source domain and the publication time for credibility judgment, without comprehensively considering multiple dimensions such as the author's authority, the semantic quality of the content, and the credibility of the dissemination chain. In scenarios such as authoritative sites reprinting low-quality content or newly released authoritative information that has not yet accumulated dissemination data, it is easy to make misjudgments, and the accuracy is difficult to guarantee. Other solutions, although they adopt a multi-dimensional evaluation approach, lack a systematic dimension design, and the feature extraction and quantitative scoring of each dimension are processed serially. A single piece of information must complete the calculation of all dimensions in sequence before the result is output. The evaluation efficiency is difficult to support large-scale information batch processing scenarios, and it is impossible to balance evaluation accuracy and processing efficiency.

[0004] 2. Existing multi-dimensional assessment schemes generally use a weighted summation method with fixed weights to calculate the overall credibility score. Once the weights are set, they take effect statically throughout the process, without taking into account the quality differences of the feature data of each dimension under different data sources. In actual multi-source information collection scenarios, the completeness of fields and the reliability of data returned by different channels vary greatly. When a certain dimension's data is missing, the fields are incomplete, or the collection channel has low reliability, the score result of that dimension itself has limited reference value, but it will still participate in the comprehensive calculation according to the preset fixed weights. Unreliable feature data will cause excessive interference to the final score, directly leading to the distortion of the assessment results. Summary of the Invention

[0005] The purpose of this invention is to provide a multi-dimensional dynamic credibility evaluation method based on an automated information processing system, which can effectively solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A multi-dimensional dynamic credibility assessment method based on automated information processing systems includes: Obtain the metadata of the information item to be evaluated, and simultaneously extract the original features of five dimensions from the metadata: source credibility features, information freshness features, author authority features, content quality features, and URL standardization features. Perform data standardization processing on the original features of each dimension to generate a structured dataset containing the standardized features of the five dimensions. Each of the five dimensions is configured with an independent scoring engine. Each scoring engine acquires the standardized features of the corresponding dimension in the structured dataset in parallel and outputs the score value of each dimension. Obtain the score values ​​of the five dimensions, configure the weight coefficients for each dimension, read the data quality identifiers of the original features of each dimension in the structured dataset and calculate the confidence factor of each dimension accordingly, and calculate the final comprehensive credibility score using a weighted summation method with confidence correction. The final overall credibility score is compared with preset dual thresholds, which include a preset high threshold and a preset low threshold. When the final overall credibility score is greater than or equal to the preset high threshold, a learning instruction is output. When the final overall credibility score is greater than or equal to the preset low threshold and less than the preset high threshold, a review instruction is output. When the final overall credibility score is less than the preset low threshold, a rejection instruction is output, and the instruction is sent to the downstream processing system.

[0007] Preferably, the source credibility feature includes the source domain name and the associated records of the source domain name in the public reputation query system; The information freshness feature includes a publication timestamp and the time zone attribute and synchronization calibration flag of the publication timestamp; The author authority feature includes the author name field and the author's cumulative credibility score in the internal historical evaluation records corresponding to the author name field; The content quality features include the main text and the effective total number of words and paragraph structure distribution data of the main text; The URL canonical characteristics include the complete path structure of the Uniform Resource Locator and the list of query parameters carried by the complete path structure.

[0008] Preferably, the information freshness scoring engine obtains the publication timestamp and the current evaluation time, calculates the day difference t between the publication timestamp and the current evaluation time, and calls the segmented decay function to calculate the score; The calculation rules for the piecewise decay function are as follows: when t≤1, output 1.0; when 1<t≤7, output 0.8; when 7<t≤30, output 0.6; when 30<t≤365, output 0.4; and when t>365, output 0.2.

[0009] Preferably, when extracting the original features of the five dimensions simultaneously, a parallel processing architecture is adopted, wherein the parallel processing architecture simultaneously starts five independent data parsing threads; After each of the five independent data parsing threads extracts the original feature data of its corresponding dimension, the extracted original feature data is uniformly sent to the data standardization processor. The data standardization processor converts heterogeneous data formats from different data sources into a unified internally predefined standard data structure according to preset standardization rules. The standard data structure includes five required fields: source domain name, publication timestamp, author name, text content, and Uniform Resource Locator.

[0010] Preferably, the content quality scoring engine obtains the main text content field, calls the word segmentation processor to count the total number of valid words w in the main text content field, and calls the paragraph separator recognizer to count the number of paragraphs p in the main text. The length score L is calculated using a piecewise linear function when w < 500. When 500 ≤ w ≤ 2000, L = 1.0; when w > 2000, L = 1.0. ; Calculate paragraph score P. ; The average of the length score and the paragraph score is used as the content quality score.

[0011] Preferably, the URL canonical scoring engine performs pattern matching scanning on the Uniform Resource Locator (URL), and uses a regular expression matching algorithm to traverse a preset list of suspicious keywords to detect whether the URL contains a preset path identifier. The preset path identifier includes at least / sponsored, / ad, / clickbait, and / promo. If the Uniform Resource Locator contains any of the above identifiers, a preset low score is output; if the Uniform Resource Locator does not contain any suspicious patterns, a preset high score is output.

[0012] Preferably, the data quality identifier includes an integrity status and a credibility level. The integrity status is divided into a complete level, a partially missing level, and a completely missing level. The credibility level is divided into a high level, a medium level, and a low level. The confidence factor manager uses the combination of the integrity status and the confidence level as an index to query a pre-set confidence mapping table; In the confidence mapping table, the confidence factor corresponding to the combination of the partially missing level and the medium level is 0.7, the confidence factor corresponding to the combination of the partially missing level and the low level is 0.5, the confidence factor corresponding to the combination of the completely missing level and the high level is 0.4, the confidence factor corresponding to the combination of the completely missing level and the medium level is 0.2, and the confidence factor corresponding to the combination of the completely missing level and the low level is 0.1.

[0013] Preferably, the review instruction includes a reason tag for triggering the review instruction. The reason tag is automatically generated based on the difference between the scores of each dimension and the threshold: when the source credibility score is lower than the preset high threshold but higher than the preset low threshold, a source credibility score in the middle range tag is generated; when the content quality score is lower than 0.5, a content quality score low tag is generated; when the information freshness score is lower than 0.4, an information timeliness insufficient tag is generated; when multiple dimensions are triggered simultaneously, a combined tag is generated. The rejection instruction includes a rejection reason label, which is automatically generated based on the dimensions whose scores are below a preset single-item rejection threshold. The preset single-item rejection threshold for each dimension is 0.3.

[0014] Preferably, it also includes dynamic feedback optimization: the weight adjustment engine listens to the review result data of the manual review queue, calculates the degree of deviation between the manual review conclusion and the evaluation conclusion, calculates the adjustment range of the weight coefficients of each dimension according to the degree of deviation, and writes the updated weight coefficients into the weight configuration manager. The threshold optimization engine calculates the proportion of learning instructions, review instructions, and rejection instructions within a statistical period of 1000 evaluation records. The preset expected distribution range is 40% to 60% for learning instructions, 20% to 40% for review instructions, and 10% to 30% for rejection instructions. When the proportion of a certain type of instruction deviates from the expected distribution range, the preset high threshold and the preset low threshold are finely adjusted in the same or opposite direction according to a preset step size of 0.02.

[0015] Preferably, it also includes scene switching: the scene adaptation manager maintains a scene parameter configuration table, which stores the recommended value combination of the preset high threshold and the preset low threshold under the academic knowledge base construction scene, the news aggregation scene and the social media monitoring scene, as well as the recommended weight coefficient combination of each of the five dimensions; The preset high threshold for the academic knowledge base construction scenario is 0.85 and the preset low threshold is 0.65; the preset high threshold for the news aggregation scenario is 0.75 and the preset low threshold is 0.55; and the preset high threshold for the social media monitoring scenario is 0.60 and the preset low threshold is 0.40. The scene adaptation manager queries the scene parameter configuration table based on the target scene identifier carried in the received scene switching instruction, writes the read threshold into the dual threshold configuration parameter, writes the read weight coefficient into the current active configuration of the weight configuration manager, and records the scene switching operation log.

[0016] In summary, this application includes at least one of the following beneficial technical effects: 1. This application evaluates the credibility of online information from five dimensions: source credibility, information freshness, author authority, content quality, and URL standardization. It also adopts a parallel processing architecture to extract and score the original features of the five dimensions simultaneously. This effectively solves the technical problems of low evaluation efficiency due to serial processing and easy misjudgment in single-dimensional evaluation in existing multi-dimensional evaluation schemes. It improves the evaluation efficiency in large-scale information processing scenarios while ensuring the accuracy of the evaluation.

[0017] 2. This application addresses the technical problem in existing fixed-weight configuration schemes where unreliable features still participate in the calculation with fixed weights when data for a certain dimension is missing or has low reliability. This is achieved by attaching data quality identifiers containing integrity status and confidence level to the original features of each dimension. The confidence factor manager uses the combination of integrity status and confidence level as an index to query the confidence mapping table to obtain the confidence factors for each dimension, and uses a weighted summation method with confidence correction to calculate the final comprehensive confidence score.

[0018] 3. This application compares the final overall credibility score with preset dual thresholds, and automatically outputs learning instructions, review instructions, or rejection instructions based on the comparison results and sends them to the downstream processing system. The learning instructions instruct the downstream processing system to directly include the information item to be evaluated into the knowledge base, the review instructions instruct the downstream processing system to transfer the information item to be evaluated to the manual review queue, and the rejection instructions instruct the downstream processing system to directly discard the information item to be evaluated. This achieves seamless connection between the evaluation results and the downstream processing flow, and the entire process of processing learning instructions and rejection instructions does not require manual intervention. Attached Figure Description

[0019] Figure 1 This is the overall flowchart of the multi-dimensional credibility dynamic evaluation method of this application.

[0020] Figure 2 This is a flowchart of the feature extraction process in this application.

[0021] Figure 3 This is a flowchart of the single-dimensional scoring calculation process for this application.

[0022] Figure 4 This is a flowchart of the dynamic weighted fusion process in this application.

[0023] Figure 5 This is a flowchart of the automated decision-making output process in this application. Detailed Implementation

[0024] To further illustrate the technical means and effects adopted by the present invention to achieve the intended purpose, the following detailed description of specific embodiments based on the present invention is provided in conjunction with the accompanying drawings and preferred embodiments.

[0025] like Figure 1 As shown, this invention provides a multi-dimensional dynamic credibility assessment method based on an automated information processing system, including feature extraction, single-dimensional score calculation, dynamic weighted fusion, and automated decision output. The method achieves a comprehensive quantitative assessment of the credibility of information such as online articles by processing the five core dimensions of the information items to be evaluated in parallel, and directly drives the decision execution of downstream automated processing procedures.

[0026] Step S1, feature extraction, establishes a communication connection with an external system through a preset data acquisition interface to obtain metadata of the information items to be evaluated. Simultaneously, it extracts raw features from five dimensions: source credibility, information freshness, author authority, content quality, and URL standardization. The extracted results are then standardized to generate a structured dataset for subsequent scoring calculations. For example... Figure 2 As shown, the specific steps include the following: S101, the feature extraction module establishes a communication connection with an external web crawler system or application interface through a preset data acquisition interface. This interface is based on the standard HTTP or HTTPS protocol to receive and obtain the complete metadata of the information item to be evaluated in real time. The information item here specifically refers to a web article, whose metadata includes at least the source domain name, publication timestamp, author name, full text, and Uniform Resource Locator. The metadata is encapsulated by the external system according to a pre-agreed communication format and pushed to the feature extraction module.

[0027] S102, the feature extraction module simultaneously extracts the five core dimensions of the original features of the information item. The source credibility feature includes the source domain name and the associated records of the domain name in the public reputation query system. The information freshness feature is based on the publication timestamp and includes the time zone attribute and synchronization calibration mark of the timestamp. The author authority feature includes the author name field and the author's cumulative credibility score in the historical evaluation records within the feature extraction module. The content quality feature covers the main text itself as well as the word count and paragraph structure distribution data of the main text. The URL normalization feature includes the complete path structure of the Uniform Resource Locator and the list of query parameters it carries.

[0028] The original features of each of the above dimensions were directly extracted from the metadata obtained in S101.

[0029] S103, the feature extraction module adopts a parallel processing architecture. This architecture simultaneously starts five independent data parsing threads, each corresponding to a specific dimension. Each thread is equipped with a dedicated feature parser and data validator.

[0030] The feature parser performs field identification and extraction operations based on the format specifications of the data items corresponding to the dimension, while the data validator performs validity verification on the extracted data format to exclude abnormal data that does not meet the format requirements, thereby ensuring that the original features entering the subsequent process have basic integrity.

[0031] S104, the raw feature data extracted by each thread is uniformly sent to the built-in data standardization processor. The processor converts the heterogeneous data formats from different data sources into a predefined standard data structure according to the standardization rules preset by the feature extraction module. The standard data structure includes five required fields: source domain name, publication timestamp, author name, text content, and Uniform Resource Locator.

[0032] The data standardization processor performs corresponding format merging operations on different fields. For example, it performs case normalization on domain name fields, performs a unified time reference system conversion on timestamp fields, and performs whitespace compression on text fields, thereby ensuring the consistency of the output structure.

[0033] The S105 data standardization processor uses an extensible configuration mechanism to adapt to heterogeneous data sources. This mechanism relies on external XML or JSON configuration files to define the mapping relationship between external fields and standard fields, data type conversion rules, and default value filling strategies when data is missing.

[0034] When a field specified in the configuration file is missing or has an empty value in the original metadata, the data standardization processor fills it in according to the configured default value and adds a missing status flag to the field. This configuration file allows administrators to dynamically update it based on changes in the format of external data sources without modifying the system code.

[0035] S106 After the data standardization processor completes the transformation, it generates a structured dataset containing five original features. This structured dataset not only contains the final values ​​of the five standard fields mentioned above, but also adds a data quality identifier to each field. The data quality identifier is used to indicate the integrity status and credibility level of the feature data.

[0036] The integrity status is determined based on whether the field was successfully obtained and whether the verification result is qualified, and is specifically divided into three levels: complete, partially missing, or completely missing. The credibility level is determined based on the preset reliability of the original source channel corresponding to the feature data and is specifically divided into three levels: high, medium, or low.

[0037] Data quality identifiers are stored in the dataset as enumerated values ​​for use in the subsequent dynamic calculation of confidence factors.

[0038] After steps S101 to S106 are completed, the feature extraction stage completes all preprocessing work for a single piece of information to be evaluated. The standardized structured dataset output by this stage will serve as the input for the subsequent single-dimensional scoring calculation stage and provide data support with a unified format for the parallel operation of the five independent scoring engines in step S2.

[0039] Step S2, the single-dimensional score calculation stage, configures an independent scoring engine for each of the five core dimensions. Each engine receives the standardized feature data output from the feature extraction stage and maps the original features to a unified scoring range according to its own preset scoring algorithm, ultimately outputting the score results for each dimension. For example... Figure 3 As shown, the specific steps include the following: S201, The source credibility scoring engine obtains the source domain name field from the standardized dataset, and queries the trusted domain name score mapping table preset in the configurable policy storage module with the source domain name as the key. The mapping table adopts a key-value pair storage structure, where the key is the domain name and the value is the credibility score corresponding to the domain name.

[0040] The trusted domain score mapping table is pre-configured with at least one mapping entry for arxiv.org with a score of 0.95, nature.com with a score of 0.9, and techcrunch.com with a score of 0.8. It also stores the mapping relationship between various platform domains and their corresponding preset scores. Different platform types are configured with different scores according to their content quality trustworthiness level. If the queried domain is not in the mapping table, the preset default value of 0.5 is returned. The mapping table allows administrators to dynamically add, modify, or delete mapping entries at runtime through the management interface to achieve real-time adjustment of the scoring strategy.

[0041] S202, the information freshness scoring engine obtains the publication timestamp field and the current evaluation time from the standardized dataset, calculates the difference in days between the two and records the difference in days as t, and then calls the segmented decay function to calculate the score.

[0042] The calculation rules for the piecewise decay function are as follows: when t≤1, output the first preset score of 1.0; when 1<t≤7, output the second preset score of 0.8; when 7<t≤30, output the third preset score of 0.6; when 30<t≤365, output the fourth preset score of 0.4; and when t>365, output the fifth preset score of 0.2.

[0043] The segmented decay function employs a boundary value optimization algorithm to perform unique branch matching for cases where the difference in days is exactly equal to the threshold boundary, preventing the same value from triggering multiple scoring branches. The information freshness scoring engine maintains an internal time synchronization service that periodically calibrates with an external network time protocol server to ensure the accuracy of the scoring time benchmark.

[0044] S203, the author authority rating engine obtains the author name field from the standardized dataset. First, it performs an exact match query between the author name and the preset expert whitelist. The whitelist adopts a fast lookup structure based on a hash table, and the whitelist entries are stored with encrypted salt values ​​to prevent the leakage of sensitive expert information.

[0045] If the author's name exists in the whitelist, the author's authority score will be the first preset score of 0.9. If the author's name is not in the whitelist, the process of judging the integrity of the name format will begin.

[0046] The author authority scoring engine calls a name format parser to analyze the composition of the author's name to determine whether the author's name contains both a first name and a last name. The name format parser supports the recognition of at least Western name formats, Eastern name formats, and compound name formats. If the author's name is complete and contains both a first name and a last name, the author authority score is the second preset score of 0.7. If the author's name is incomplete or the author's name information cannot be obtained, the author authority score is the third preset score of 0.5.

[0047] S204, the content quality scoring engine obtains the main text content field from the standardized dataset, calls the built-in text preprocessing module to process the main text content to filter out non-main text content, after preprocessing, calls the word segmentation processor to count the total number of effective words in the main text and records the total number of effective words as w, and calls the paragraph separator recognizer to count the number of paragraphs in the main text and records the number of paragraphs as p, and then uses a piecewise linear function to calculate the length score L and the paragraph score P respectively.

[0048] The calculation rules for the length score L are as follows: When w < 500 ; When 500≤w≤2000, L=1.0; When w > 2000, ; Where w represents the total number of valid words in the main text, in units of words, and max represents taking the maximum of the two values ​​in parentheses.

[0049] The calculation rules for paragraph score P are as follows: ; Where p represents the number of paragraphs in the main text, in units of , and min represents taking the minimum value between the two values ​​in parentheses.

[0050] Final content quality score The calculation formula is: ; The score ranges from 0 to 1.

[0051] S205, the URL canonical scoring engine obtains the Uniform Resource Locator (URL) field from the standardized dataset, performs pattern matching scanning on the URL, and uses a regular expression matching algorithm to traverse a preset list of suspicious keywords to detect whether the URL contains preset path identifiers of commercial promotion or click inducement.

[0052] The default path identifiers include at least / sponsored, / ad, / clickbait, and / promo. The built-in pattern matching optimizer compiles and caches commonly used regular expression patterns to improve processing efficiency during large-scale detection. If the Uniform Resource Locator contains any of the above identifiers, the URL normalization score is set to the default low score of 0.2. If the Uniform Resource Locator does not contain any suspicious patterns, the URL normalization score is set to the default high score of 0.8. The suspicious keyword pattern library can be extended and added by the administrator through the configuration file.

[0053] After step S2 is completed, the five independent scoring engines will output scores for five dimensions: source credibility score, information freshness score, author authority score, content quality score, and URL standardization score. These scores will be used as inputs for the subsequent dynamic weighted fusion process and will be called for the weighted summation calculation in step S3.

[0054] Step S3, the dynamic weighted fusion stage, obtains the five-dimensional score values ​​output from step S2, assigns weight coefficients to each dimension and introduces confidence factors for correction, and calculates the overall credibility score through weighted summation. For example... Figure 4 As shown, the specific steps include the following: S301, the dynamic weighted fusion module calls the built-in weight configuration manager. The weight configuration manager stores weight configuration templates for multiple predefined application scenarios. Each weight configuration template specifies corresponding weight coefficients for five dimensions: source credibility, information freshness, author authority, content quality, and URL standardization. The initial values ​​of the weight coefficients for each dimension are read from the corresponding template by the weight configuration manager according to the currently selected application scenario identifier, and the sum of the weight coefficients for all dimensions is always equal to 1.

[0055] The weight configuration manager adopts a versioned storage mechanism. Each time a weight is adjusted, the version number and timestamp of the current configuration are recorded to support subsequent configuration backtracking and auditing. The application scenario identifier is passed in by the upper-layer caller when starting the evaluation process, or it is automatically matched and determined by the weight configuration manager based on the source channel type of the information item to be evaluated.

[0056] In step S302, the basic weighted fusion calculation obtains the weight coefficients of each dimension in the current scenario from the weight configuration manager, and also obtains the score values ​​of each dimension output in step S2. The basic comprehensive credibility score is calculated using a weighted summation method. The calculation formula is as follows: ; in This represents the total score of basic overall credibility. An integer representing the dimension index, with a value ranging from 1 to 5; Indicates the first Weight coefficients for each dimension; Indicates the first The scores for each dimension are calculated and distributed in the range of 0 to 1.

[0057] In step S303, the confidence factor calculation step involves the confidence factor manager reading the data quality identifiers of each field in the structured dataset output from the feature extraction stage. These data quality identifiers record the integrity status and confidence level of the original features for each dimension. The confidence factor manager then calculates the confidence factor for each dimension separately. .

[0058] Confidence factor The calculation is implemented using a lookup table method. The confidence factor manager has a pre-set confidence mapping table. The confidence mapping table uses the combination of integrity status and confidence level as an index to directly give the corresponding confidence factor value. The integrity status is divided into three levels: complete, partially missing, and completely missing. The confidence level is divided into three levels: high, medium, and low. There are a total of 9 index items in the two combinations. Each index item corresponds to a confidence factor value in the range of 0 to 1 in the confidence mapping table.

[0059] When the data quality identifier indicates an integrity status of completeness and a confidence level of high, the confidence factor... The preset baseline value is 1.0. When the integrity status is partially missing or the confidence level is medium or low, the confidence factor manager queries the confidence mapping table based on the combined index of integrity status and confidence level and directly reads the corresponding confidence factor value.

[0060] In the confidence mapping table, the confidence factor for the combination of partially missing grades and medium grades is 0.7, the confidence factor for the combination of partially missing grades and low grades is 0.5, the confidence factor for the combination of completely missing grades and high grades is 0.4, the confidence factor for the combination of completely missing grades and medium grades is 0.2, and the confidence factor for the combination of completely missing grades and low grades is 0.1.

[0061] The confidence factor manager calls a built-in multi-validation algorithm to validate the input feature data during calculation. The multi-validation algorithm consists of three validation levels executed sequentially. The first level is integrity checking, which checks whether each standard field is empty or consists of entirely blank characters. The second level is format validation, which checks whether the timestamp field is a valid date format, whether the domain name field conforms to the standard domain name format specification, and whether the Uniform Resource Locator field conforms to the URL format specified in RFC 3986. The third level is logical consistency checking, which checks whether the publication timestamp is later than the current evaluation time and whether the text content matches the word count.

[0062] If any of the above verification levels detects an anomaly, the confidence factor manager lowers the confidence level of that dimension by one level and then re-queries the confidence mapping table to determine the final confidence factor value. The value is limited to a continuous range of 0 to 1, and the confidence mapping table allows administrators to customize it through an external configuration file.

[0063] S304, In the corrected weighted fusion calculation step, the confidence factors of each dimension calculated in step S303 are... The weighting coefficients used in step S302 and ratings for each dimension The final overall credibility score is calculated by merging the results and using a weighted summation method with confidence level correction. The calculation formula is as follows: ; in, This represents the final overall credibility score; An integer representing the dimension index, with a value ranging from 1 to 5; Indicates the first Weight coefficients for each dimension; Indicates the first Confidence factors in each dimension; Indicates the first The score values ​​for each dimension are used to ensure that the confidence factor is adjusted when feature data in a certain dimension is missing or has low reliability. The corresponding reduction effectively suppresses the actual contribution of this dimension's score to the overall score, preventing unreliable features from having an excessive impact on the overall score results.

[0064] After step S3 is completed, the dynamic weighted fusion stage outputs the final overall credibility score. , This will serve as input for the subsequent automated decision-making output stage, and will be used for the dual threshold comparison judgment in step S4 to drive the automatic generation of three types of processing instructions: learning, review, or rejection.

[0065] Step S4, automated decision output stage, receives the final comprehensive credibility score output from step S3. ,Will The system compares the data with preset dual thresholds, generates corresponding automated processing instructions based on the comparison results, and sends them to the downstream processing system. Simultaneously, a dynamic feedback optimization mechanism adaptively adjusts the weighting coefficients and decision thresholds. Figure 5 As shown, the specific steps include the following: S401, The automated decision-making module obtains the final comprehensive credibility score output from step S3. Simultaneously, it reads the dual threshold configuration parameters maintained within the automated decision-making module. The dual threshold configuration parameters are a preset high threshold and a preset high threshold. and preset low threshold , and All are preset constants in the interval between 0 and 1 and Greater than .

[0066] and The initial value is read by the automated decision-making module from a pre-set scenario parameter configuration table based on the current application scenario identifier. The scenario parameter configuration table pre-sets differentiated threshold combinations for different application scenarios, and the academic knowledge base constructs the scenario. Take 0.85 and Take 0.65, for news aggregation scenarios. Take 0.75 and Take 0.55, for social media monitoring scenarios Take 0.60 and Take 0.40.

[0067] The automated decision-making module allows administrators to manually adjust settings during runtime via a management interface. and The specific value can also be automatically adjusted through the threshold optimization engine in step S407.

[0068] S402, the automated decision-making module executes the threshold comparison and instruction judgment process, and the judgment rule consists of three mutually exclusive branches.

[0069] (1) The first branch is a high-confidence processing path: if Greater than or equal to If the condition is met, a learning instruction will be output. The learning instruction is a structured data packet and contains at least the following fields: an instruction type identifier (with a fixed value of LEARN), and a unique identifier for the information item to be evaluated. The specific numerical values, a complete list of score values ​​for each dimension, and timestamps.

[0070] The learning instruction is used to instruct the downstream processing system to directly include the information item to be evaluated into the knowledge base. After receiving the learning instruction, the downstream processing system automatically triggers the knowledge base update process. The knowledge base update process includes converting the format of the information item to adapt to the data structure of the knowledge base, writing the information item into the corresponding storage location of the knowledge base, updating the index of the knowledge base to support subsequent retrieval, and recording the entry operation log for audit traceability. The entire entry process does not require manual intervention.

[0071] The automated decision-making module and the downstream processing system adopt an asynchronous non-blocking communication mechanism. The overall time taken for the evaluation and warehousing operations to be completed is automatically adapted by the automated decision-making module according to the actual load capacity and network conditions of the downstream processing system, ensuring that the execution of subsequent evaluation tasks is not blocked under high load conditions.

[0072] (2) The second branch is a medium-confidence processing path: if Greater than or equal to And simultaneously less than If so, a review instruction will be output. The review instruction is a structured data packet and contains at least the following fields: an instruction type identifier (with a fixed value of REVIEW), and a unique identifier for the information item to be evaluated. The specific numerical values, a complete list of score values ​​for each dimension, a summary of the original feature data for each dimension, the reason tag that triggered the review order, and the timestamp.

[0073] The cause label is automatically generated by the automated decision-making module based on the gap between the scores of each dimension and the threshold. The source credibility score is lower than... But higher When the source credibility score is in the middle range, a label is generated; when the content quality score is below 0.5, a label is generated indicating low content quality; when the information freshness score is below 0.4, a label indicating insufficient information timeliness is generated; when multiple dimensions are triggered simultaneously, a combined label is generated.

[0074] The review instruction is used to instruct the downstream processing system to transfer the information item to be evaluated to the manual review queue for manual review. After receiving the review instruction, the downstream processing system automatically performs the following operations: writes the information item and its complete evaluation data to the designated storage location of the manual review queue, pushes the review task notification to the user interface of the manual review platform, and assigns a preset waiting priority to the information item in the manual review queue.

[0075] Waiting priority and The specific values ​​are positively correlated. The higher the value, the higher the priority of waiting.

[0076] After manual review is completed, the review conclusion is fed back to the dynamic feedback optimization module through a standardized interface. Before being fed back, the manual review conclusion is converted by the review conclusion standardization processor. The review conclusion standardization processor maps the qualitative judgment given by the reviewer to a preset standardized credibility score. In the qualitative judgment, credibility corresponds to a standardized score of 0.9, partial credibility corresponds to a standardized score of 0.6, and untrustworthiness corresponds to a standardized score of 0.3.

[0077] The review conclusion standardization processor also extracts the issue type tags marked by the reviewers. The issue type tags include at least the tags of questionable source, exaggerated content, and outdated timeliness. The review conclusion standardization processor packages the standardized score and issue type tags together into structured feedback data for the subsequent weight adjustment engine to call.

[0078] (3) The third branch is a low-confidence processing path: if Less than If the condition is met, a rejection instruction will be output. The rejection instruction is a structured data packet and contains at least the following fields: an instruction type identifier (with a fixed value of REJECT), and a unique identifier for the information item to be evaluated. The specific numerical value, the rejection reason label, and the timestamp.

[0079] The rejection reason label is automatically generated by the automated decision-making module based on the dimensions whose scores are below the preset single-item threshold. The preset single-item rejection threshold for each dimension is 0.3. When the source credibility score is below 0.3, a label is generated indicating that the source credibility score is below the single-item rejection threshold. When the content quality score is below 0.3 and the URL standardization score is below 0.3, a label is generated indicating that both the content quality score and the URL standardization score are below their respective single-item rejection thresholds.

[0080] The rejection instruction is used to instruct the downstream processing system to discard the information item to be evaluated directly without any further processing. After receiving the rejection instruction, the downstream processing system will automatically perform the following operations: move the information item to the recycle bin storage area. The data in the recycle bin storage area is managed according to the preset retention period. Within the retention period, the administrator can manually restore the information in the recycle bin. After the retention period expires, the downstream processing system will automatically trigger the permanent deletion process to release storage space.

[0081] The specific value of the retention period is dynamically set by the downstream processing system based on the current application scenario's requirements for fault tolerance and the sufficiency of storage resources, or it can be manually configured by the administrator through the management interface.

[0082] The downstream processing system also records a rejection operation log, which includes information item identifiers, rejection reason labels, scores for each dimension, and timestamps for subsequent auditing and traceability. The entire rejection process does not require manual intervention.

[0083] S403, the automated decision-making module sends the instructions generated in step S402 to the downstream processing system through the standardized instruction interface. The standardized instruction interface supports three communication methods: message queue communication protocol, remote procedure call protocol, and event bus communication protocol. The automated decision-making module selects one of the protocols for instruction transmission according to the actual deployment environment of the downstream processing system.

[0084] S404, the dynamic feedback optimization module continuously monitors the review result data of the manual review queue. Whenever a manual review is completed, the dynamic feedback optimization module reads the corresponding review result record. Each review result record contains at least the credibility judgment conclusion given by the manual reviewer for the information item and the judgment conclusion output by the automated decision module in step S402. The two sets of conclusions together constitute a comparison data.

[0085] S405, the weight adjustment engine within the dynamic feedback optimization module analyzes the comparison data obtained in step S404. The weight adjustment engine calculates the degree of deviation between the manual review conclusion and the evaluation conclusion. The degree of deviation is determined by comparing the manual score with... The absolute value of the difference between them is quantified, and the standardized credibility score output by the standardized processor is the result of manual scoring, which is the review conclusion.

[0086] The weight adjustment engine calculates the adjustment range of the weight coefficients of each dimension based on the degree of deviation. The adjustment rules are as follows: when the absolute value of the deviation is greater than the preset deviation threshold of 0.15, the adjustment step size of the weights of each dimension is the preset large step size value of 0.05; when the absolute value of the deviation is less than or equal to the preset deviation threshold of 0.15, the adjustment step size of the weights of each dimension is the preset small step size value of 0.01.

[0087] After the weight adjustment engine completes the calculation, it writes the updated weight coefficients to the weight configuration manager, which then records the version number and timestamp of this adjustment according to the versioned storage mechanism.

[0088] S406, the learning rate scheduler built into the dynamic feedback optimization module works synchronously during the operation of the weight adjustment engine. The learning rate scheduler dynamically adjusts the step size parameter used by the weight adjustment engine based on the cumulative runtime calculated from the initial startup time.

[0089] When the cumulative runtime is less than the preset initial runtime threshold of 168 hours, the learning rate scheduler sets the step size parameter to 0.05; when the cumulative runtime is greater than or equal to the preset initial runtime threshold of 168 hours, the learning rate scheduler switches the step size parameter to 0.01.

[0090] S407, The threshold optimization engine inside the dynamic feedback optimization module synchronously monitors the distribution of the number of various instructions output by the automated decision-making module in step S402. The threshold optimization engine calculates the proportion of learning instructions, review instructions and rejection instructions in a statistical period. A statistical period is 1000 evaluation records.

[0091] The expected distribution range is 40% to 60% for learning instructions, 20% to 40% for review instructions, and 10% to 30% for rejection instructions.

[0092] When the threshold optimization engine detects that the proportion of a certain type of instruction deviates from the expected distribution range, it automatically triggers a threshold adjustment process. The threshold adjustment process adjusts the threshold by a preset step size of 0.02. and Perform fine-tuning in the same or opposite direction; when the proportion of learning instructions is less than 40%, and At the same time, reduce by 0.02; when the percentage of rejected instructions exceeds 30%, and At the same time, the threshold is increased by 0.02, and the adjusted threshold parameter is written into the dual threshold configuration parameter of the automated decision-making module for subsequent decision-making.

[0093] The S408 automated decision-making module's built-in scene adaptation manager maintains a scene parameter configuration table, which stores different application scenarios in key-value pairs. and The recommended value combinations and the recommended weight coefficient combinations for each of the five dimensions.

[0094] The scene adaptation manager receives scene switching commands from external systems through a standardized interface. Upon receiving a scene switching command, it automatically performs the following operations: parses the target scene identifier carried in the command; queries the scene parameter configuration table based on the target scene identifier and reads the corresponding... , And recommended values ​​for the weight coefficients of the five dimensions; the read and Write the dual threshold configuration parameters into the automated decision-making module; write the read weight coefficients of the five dimensions into the current active configuration of the weight configuration manager; record the scene switching operation log, which includes the original scene identifier, the target scene identifier, the switching timestamp, and a comparison of the values ​​of each parameter before and after the switching.

[0095] After step S4 is completed, the entire processing flow of the multi-dimensional credibility dynamic evaluation method is finished, and the information items to be evaluated have been processed according to... The values ​​and their corresponding rating ranges are automatically routed to the knowledge base inclusion path, the manual review path, or the direct discard path.

[0096] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects.

[0097] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment includes only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A multi-dimensional dynamic credibility evaluation method based on an automated information processing system, characterized in that, include: Obtain the metadata of the information item to be evaluated, and simultaneously extract the original features of five dimensions from the metadata: source credibility features, information freshness features, author authority features, content quality features, and URL standardization features. Perform data standardization processing on the original features of each dimension to generate a structured dataset containing the standardized features of the five dimensions. Each of the five dimensions is configured with an independent scoring engine. Each scoring engine acquires the standardized features of the corresponding dimension in the structured dataset in parallel and outputs the score value of each dimension. Obtain the score values ​​of the five dimensions, configure the weight coefficients for each dimension, read the data quality identifiers of the original features of each dimension in the structured dataset and calculate the confidence factor of each dimension accordingly, and calculate the final comprehensive credibility score using a weighted summation method with confidence correction. The final overall credibility score is compared with preset dual thresholds, which include a preset high threshold and a preset low threshold. When the final overall credibility score is greater than or equal to the preset high threshold, a learning instruction is output. When the final overall credibility score is greater than or equal to the preset low threshold and less than the preset high threshold, a review instruction is output. When the final overall credibility score is less than the preset low threshold, a rejection instruction is output, and the instruction is sent to the downstream processing system.

2. The method according to claim 1, characterized in that, The source credibility feature includes the source domain name and the associated records of the source domain name in the public reputation query system; The information freshness feature includes a publication timestamp and the time zone attribute and synchronization calibration flag of the publication timestamp; The author authority feature includes the author name field and the author's cumulative credibility score in the internal historical evaluation records corresponding to the author name field; The content quality features include the main text and the effective total number of words and paragraph structure distribution data of the main text; The URL canonical characteristics include the complete path structure of the Uniform Resource Locator and the list of query parameters carried by the complete path structure.

3. The method according to claim 1, characterized in that, The information freshness scoring engine obtains the publication timestamp and the current evaluation time, calculates the day difference t between the publication timestamp and the current evaluation time, and calls the segmented decay function to calculate the score; The calculation rules for the piecewise decay function are as follows: when t≤1, output 1.0; when 1<t≤7, output 0.8; when 7<t≤30, output 0.6; when 30<t≤365, output 0.4; and when t>365, output 0.

2.

4. The method according to claim 1, characterized in that, When extracting the original features of the five dimensions simultaneously, a parallel processing architecture is adopted, which simultaneously starts five independent data parsing threads; After each of the five independent data parsing threads extracts the original feature data of its corresponding dimension, the extracted original feature data is uniformly sent to the data standardization processor. The data standardization processor converts heterogeneous data formats from different data sources into a unified internally predefined standard data structure according to preset standardization rules. The standard data structure includes five required fields: source domain name, publication timestamp, author name, text content, and Uniform Resource Locator.

5. The method according to claim 1, characterized in that, The content quality scoring engine obtains the main text content field, calls the word segmentation processor to count the total number of valid words w in the main text content field, and calls the paragraph separator recognizer to count the number of paragraphs p in the main text. The length score L is calculated using a piecewise linear function when w < 500. When 500 ≤ w ≤ 2000, L = 1.0; when w > 2000, L = 1.

0. ; Calculate paragraph score P. ; The average of the length score and the paragraph score is used as the content quality score.

6. The method according to claim 1, characterized in that, The URL canonical scoring engine performs pattern matching scans on Uniform Resource Locators (URLs) and uses a regular expression matching algorithm to traverse a preset list of suspicious keywords to detect whether the URLs contain preset path identifiers. The preset path identifiers include at least / sponsored, / ad, / clickbait, and / promo. If the Uniform Resource Locator contains any of the above identifiers, a preset low score is output; if the Uniform Resource Locator does not contain any suspicious patterns, a preset high score is output.

7. The method according to claim 1, characterized in that, The data quality identifier includes an integrity status and a credibility level. The integrity status is divided into a complete level, a partially missing level, and a completely missing level. The credibility level is divided into a high level, a medium level, and a low level. The confidence factor manager uses the combination of the integrity status and the confidence level as an index to query a pre-set confidence mapping table; In the confidence mapping table, the confidence factor corresponding to the combination of the partially missing level and the medium level is 0.7, the confidence factor corresponding to the combination of the partially missing level and the low level is 0.5, the confidence factor corresponding to the combination of the completely missing level and the high level is 0.4, the confidence factor corresponding to the combination of the completely missing level and the medium level is 0.2, and the confidence factor corresponding to the combination of the completely missing level and the low level is 0.

1.

8. The method according to claim 1, characterized in that, The review instruction includes a reason tag that triggers the review instruction. The reason tag is automatically generated based on the difference between the scores of each dimension and the threshold: when the source credibility score is lower than the preset high threshold but higher than the preset low threshold, a source credibility score in the middle range tag is generated; when the content quality score is lower than 0.5, a content quality score low tag is generated; when the information freshness score is lower than 0.4, an information timeliness insufficient tag is generated; when multiple dimensions are triggered at the same time, a combined tag is generated. The rejection instruction includes a rejection reason label, which is automatically generated based on the dimensions whose scores are below a preset single-item rejection threshold. The preset single-item rejection threshold for each dimension is 0.

3.

9. The method according to claim 1, characterized in that, It also includes dynamic feedback optimization: the weight adjustment engine listens to the review result data of the manual review queue, calculates the degree of deviation between the manual review conclusion and the evaluation conclusion, calculates the adjustment range of the weight coefficient of each dimension according to the degree of deviation, and writes the updated weight coefficient into the weight configuration manager. The threshold optimization engine calculates the proportion of learning instructions, review instructions, and rejection instructions within a statistical period of 1000 evaluation records. The preset expected distribution range is 40% to 60% for learning instructions, 20% to 40% for review instructions, and 10% to 30% for rejection instructions. When the proportion of a certain type of instruction deviates from the expected distribution range, the preset high threshold and the preset low threshold are finely adjusted in the same or opposite direction according to a preset step size of 0.

02.

10. The method according to claim 1, characterized in that, It also includes scene switching: The scene adaptation manager maintains a scene parameter configuration table, which stores the recommended value combinations of the preset high threshold and the preset low threshold under the academic knowledge base construction scene, the news aggregation scene, and the social media monitoring scene, as well as the recommended weight coefficient combinations of the five dimensions respectively. The preset high threshold for the academic knowledge base construction scenario is 0.85 and the preset low threshold is 0.65; the preset high threshold for the news aggregation scenario is 0.75 and the preset low threshold is 0.55; and the preset high threshold for the social media monitoring scenario is 0.60 and the preset low threshold is 0.

40. The scene adaptation manager queries the scene parameter configuration table based on the target scene identifier carried in the received scene switching instruction, writes the read threshold into the dual threshold configuration parameter, writes the read weight coefficient into the current active configuration of the weight configuration manager, and records the scene switching operation log.