Content auditing method and device, equipment, storage medium and computer program product

By employing a dual review mechanism of a pre-set proofreading engine and a large-scale quality inspection model, combined with a knowledge base and a deep research agent, the problem of low detection rate and low accuracy in traditional content review methods has been solved, achieving efficient and accurate content review results.

CN120974141APending Publication Date: 2025-11-18BEIJING QIHOOD TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511102047.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing content review methods suffer from low detection rates and low accuracy. In particular, when dealing with large volumes and diverse formats of content, they require significant human resources and rely on traditional keyword comparison systems that cannot effectively identify deep semantics and contextual dependencies.

Method used

A pre-set proofreading engine is used for preliminary review to generate an initial review report. Semantic quality inspection is then performed through a large quality inspection model, forming a dual review mechanism of preliminary review plus semantic verification. The pre-set proofreading engine is used to quickly filter the content to be reviewed, while the large quality inspection model performs in-depth semantic understanding and fact-checking. Supplementary checks are also performed by combining a knowledge base and a deep research agent.

Benefits of technology

It improves the detection rate and accuracy of content review, effectively identifies high-frequency, patterned errors and deep semantic issues, and enhances the efficiency and quality of content security review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974141A_ABST
    Figure CN120974141A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a content auditing method, device and equipment, a storage medium and a computer program product.The method comprises the steps that input to-be-audited content is responded, preliminary auditing is conducted on the to-be-audited content through a preset proofreading engine, a preliminary auditing report is obtained, the preliminary auditing report is input into a quality inspection large model, and a quality inspection result is obtained; semantic quality inspection is carried out on the preliminary examination report through a quality inspection large model, and a quality inspection report of the to-be-examined content is obtained; according to the method, the to-be-audited content is preliminarily audited through the preset proofreading engine, and semantic quality inspection is performed on the preliminary audit report through the quality inspection large model, so that a dual audit mechanism of preliminary audit and semantic verification is formed, and the detection rate and accuracy of content audit are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a content review method and device, equipment, a storage medium and a computer program product. BACKGROUND

[0002] At present, in the whole process of content production, content review is a very key link. However, the content is large in volume and various in form, and the review of content security needs a large number of human resources investment, and the review personnel needs to have rich experience and professional knowledge. Therefore, the related content review method usually has the defects of low detection rate and low accuracy. SUMMARY

[0003] The main purpose of the present application is to provide a content review method, device, equipment, storage medium and computer program product, which aims to solve the technical problem of low detection rate and low accuracy of the related content review method.

[0004] To achieve the above purpose, the present application provides a content review method, which comprises the following steps:

[0005] In response to the input of the content to be reviewed, the content to be reviewed is preliminarily reviewed by a preset proofreading engine to obtain a preliminary review report;

[0006] The preliminary review report is input into a quality inspection large model;

[0007] The semantic quality inspection of the preliminary review report is performed by the quality inspection large model to obtain a quality inspection report of the content to be reviewed.

[0008] In addition, to achieve the above purpose, the present application also provides a content review device, which comprises:

[0009] A preliminary review module is configured to, in response to the input of the content to be reviewed, preliminarily review the content to be reviewed by a preset proofreading engine to obtain a preliminary review report;

[0010] A model calling module is configured to input the preliminary review report into a quality inspection large model;

[0011] A model quality inspection module is configured to perform semantic quality inspection of the preliminary review report by the quality inspection large model to obtain a review report of the content to be reviewed.

[0012] In addition, to achieve the above purpose, the present application also provides a content review device, which comprises a storage, a processor and a content review program stored in the storage and executable on the processor, and the content review program is configured to implement the content review method as described above.

[0013] In addition, to achieve the above object, the application further provides a storage medium, wherein the storage medium stores a content review program, and the content review program is executed by a processor to implement the content review method.

[0014] In addition, to achieve the above object, the application further provides a computer program product, wherein the computer program product comprises a content review program, and the content review program is executed by a processor to implement the content review method.

[0015] The one or more technical solutions provided by the application have at least the following technical effects:

[0016] In the application, in response to inputted to-be-reviewed content, the to-be-reviewed content is preliminarily reviewed by a preset proofreading engine, an initial review report is obtained, the initial review report is inputted into a quality inspection large model, the initial review report is semantically reviewed by the quality inspection large model, and a quality inspection report of the to-be-reviewed content is obtained; since the to-be-reviewed content is preliminarily reviewed by the preset proofreading engine and the initial review report is semantically reviewed by the quality inspection large model, a double review mechanism of preliminary review plus semantic verification is formed, and thus the detection rate and accuracy of content review are improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the application and serve to explain the principles of the application together with the specification.

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced in the following. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without creative labor.

[0019] Figure 1 The flowchart of the first embodiment of the content review method of the application;

[0020] Figure 2 The flowchart of the second embodiment of the content review method of the application;

[0021] Figure 3 The flowchart of the third embodiment of the content review method of the application;

[0022] Figure 4 The flowchart of the fourth embodiment of the content review method of the application;

[0023] Figure 5 The flowchart of an embodiment of the content review method of the application;

[0024] Figure 6 Figure 1 is a schematic diagram of a quality inspection model training architecture for an embodiment of the content review method of the present application;

[0025] Figure 7 Figure 2 is a schematic diagram of a module structure of a content review device according to an embodiment of the present application;

[0026] Figure 8 Figure 3 is a schematic diagram of a device structure of a hardware operating environment involved in the content review method according to an embodiment of the present application.

[0027] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0028] It should be understood that the specific embodiments described herein are merely intended to explain the technical solutions of the present application, and are not intended to limit the present application.

[0029] In order to better understand the technical solutions of the present application, the following will be described in detail in combination with the accompanying drawings and specific embodiments.

[0030] At present, in the whole process of content production, content review is a very key link. However, the content volume is large, the form is various, and the content security review needs a large amount of human resources investment, and the review personnel needs to have rich experience and professional knowledge.

[0031] However, the content production party often does not have a matching content review team and sound technical support measures, the network platform is constantly innovating and iterating, various social creators publish content in various styles, professional content risk control technical means are needed to find illegal content and do a good job in review, and some network platforms or self-media produce or reprint political news reports and comments, and even outsource the content review to enterprises and institutions without content qualifications or experience, resulting in many illegal and bad content entering the public view and destroying the network ecology.

[0032] The related content review method is usually a traditional text proofreading and filtering system based on keyword comparison, and there are generally problems of not being able to detect and not being able to accurately detect, which cannot meet the needs of content security review in the rapidly developing mobile internet era.

[0033] Therefore, in order to overcome the above defects, the present application provides a solution, which comprises: in response to the input content to be audited, performing preliminary auditing on the content to be audited by a preset proofreading engine, obtaining a preliminary auditing report, inputting the preliminary auditing report into a quality inspection large model, performing semantic quality inspection on the preliminary auditing report by the quality inspection large model, and obtaining a quality inspection report of the content to be audited; since the content to be audited is preliminarily audited by the preset proofreading engine and the preliminary auditing report is subjected to semantic quality inspection by the quality inspection large model, a double auditing mechanism of preliminary auditing plus semantic verification is formed, thereby improving the detection rate and accuracy of content auditing.

[0034] It should be noted that the execution subject of the embodiment can be a content auditing device with data processing, network communication and program running functions, such as a server, a computer, etc., or other electronic devices capable of achieving the same or similar functions, which are not limited in the embodiment.

[0035] Based on this, the embodiment of the present application provides a content auditing method, which is described with reference to Figure 1 , Figure 1 The flowchart of the first embodiment of the content auditing method of the present application is shown in the figure.

[0036] In the first embodiment, the content auditing method comprises:

[0037] Step S10: in response to the input content to be audited, performing preliminary auditing on the content to be audited by a preset proofreading engine, and obtaining a preliminary auditing report.

[0038] It should be understood that, in order to quickly process a large amount of content, quickly eliminate high-frequency and patterned errors (such as term misspelling and format problems), and reduce the invalid workload of subsequent large models, in the embodiment, the content to be audited is first subjected to preliminary auditing by the preset proofreading engine to obtain a preliminary auditing report.

[0039] It can be understood that the content to be audited can refer to various types of content (such as text, images, and audio, etc., in this embodiment and other embodiments, taking the text to be audited as an example for description) that need to be audited. In the industry information scene, it is specifically the text content of industry-related news reports, information articles, analysis and comments, covering various information such as industry dynamics, enterprise information, policy interpretation, etc. The preset proofreading engine can refer to an automated preliminary review tool based on rule matching, relying on a cloud-based standardized vocabulary (such as an important expression library, an audit and proofreading rule library, and a historical error correction library) and a user-defined vocabulary, and through rule matching, regular expression and template comparison technology, the content is scanned and preliminarily filtered at high speed, and the engine focuses on processing high-frequency, clear and patterned errors. The preliminary review report can be the preliminary review result output by the preset proofreading engine, which includes all suspected errors marked by the engine through rule matching, such as fixed expression errors, blacklisted words, and formatting errors, and is the basis for subsequent in-depth quality inspection.

[0040] In a specific implementation, two types of vocabularies are relied on, one is a cloud-based standardized vocabulary (such as an industry terminology specification library and a historical common error library), and the other is a user-defined vocabulary (such as enterprise-specific expression rules and a product name specification library). Efficient rule matching (such as fixed combination verification), regular expressions (such as date format "YYYY-MM-DD" verification), and template comparison (such as product parameter expression templates) are used to scan the input content to be audited word by word and sentence by sentence. Focus on high-frequency, clear, and patterned errors, such as term misspelling ("block chain" is miswritten as "block chain"), format error ("2024.5.1" is not written according to the specification as "2024 May 1"), and blacklisted words (industry-banned false propaganda words "absolutely effective"). Generate a preliminary review report, clearly mark the location, type (such as "term error" and "format error") and preliminary judgment basis (such as "not in accordance with industry terminology specification library article XX") of all suspected errors.

[0041] Step S20: inputting the preliminary review report into the quality inspection large model.

[0042] It should be understood that the quality inspection large model can refer to a large model with semantic understanding, fact checking, and autonomous research capabilities, and can align with the preferences of human experts. Its core function is to deeply verify and supplement the preliminary review results, achieving "rule filtering + semantic verification" double insurance.

[0043] Step S30: performing semantic quality inspection on the preliminary review report by the quality inspection large model to obtain a quality inspection report of the content to be audited.

[0044] It can be understood that, in order to identify the preliminary review false positives through semantic understanding, find hidden errors through deep mining, and solve the problem of "not detected and not accurate" of traditional rule engine, in the embodiment, the semantic quality inspection of the preliminary review report is performed through the quality inspection large model to obtain the quality inspection report of the content to be audited.

[0045] It should be understood that semantic quality inspection can refer to a deep quality inspection process of context analysis, logical reasoning, and fact checking of content by the quality inspection large model through full-text semantic understanding, which can solve the problems of deep semantic, context dependence, and dynamic fact that cannot be handled by traditional rule engine. The quality inspection report can refer to the final report that combines the preliminary review results of the preset proofreading engine and the semantic quality inspection results of the quality inspection large model, which includes three types of errors: A type (high confidence errors confirmed by rule engine and large model), B type (newly found errors by large model only), and C type (preliminary review false positives and reasons determined by large model), and is accompanied by structured thinking chain analysis.

[0046] In a specific implementation, the quality inspection large model performs semantic quality inspection on the preliminary review report through at least one of the analysis methods of false positive identification, missing supplement inspection, context analysis, logical reasoning, and fact checking through full-text semantic understanding to obtain the quality inspection report of the content to be audited, which is not limited by the embodiment.

[0047] Further, in order to eliminate the false positives of rule engine due to "mechanical matching" and break through the limitation of traditional rule engine "only matching without understanding", the step S30 comprises: performing context understanding and fact checking on the suspected errors in the preliminary review report through the quality inspection large model to obtain false positive identification results; performing deep reading on the preliminary review report through the quality inspection large model to obtain missing supplement inspection results; and generating the quality inspection report of the content to be audited according to the false positive identification results and / or the missing supplement inspection results.

[0048] It can be understood that suspected errors can refer to content that may have problems marked by the preset proofreading engine through rule matching, template comparison and other techniques, which need to be further verified by the quality inspection large model whether it is a true error. Context understanding can refer to the process of the quality inspection large model judging whether the suspected error is reasonable in the specific context by analyzing the context logic, semantic association, expression scene, etc. of the content. Fact checking can refer to the process of the quality inspection large model verifying the authenticity of factual information (such as data, events, and associated relationships) in the content by calling internal knowledge bases (such as industry specification libraries) or external tools (such as web search engines). False positive identification results can refer to the results of the quality inspection large model determining which are the results of the rule engine misjudgment after the quality inspection large model performs context understanding and fact checking on the suspected errors in the preliminary review report, which can include "excluded errors" and detailed reasons. Deep reading can refer to the process of the quality inspection large model going beyond the scope of the preliminary review report and performing semantic analysis and logical analysis on the full text sentence by sentence and section by section, covering comprehensive checks in multiple dimensions such as language, facts, timeliness, and logic. Missing inspection results can refer to the results of the quality inspection large model finding hidden errors (such as factual errors, logical contradictions, and timeliness deviations) that are not marked by the preset proofreading engine through deep reading.

[0049] In a specific implementation, false positive identification: for suspected errors in the preliminary review report, combined with context semantics, industry knowledge and fact checking, determine whether it is a true error (such as the preliminary review marking "AI technology" is an incorrect expression, the model confirms that "AI technology" is a reasonable expression through semantic analysis, and determines it as a false positive). Missing inspection: deep reading of the full text to identify hidden errors that are not recognized by the rule engine, such as logical contradictions ("a product has a 10-hour battery life, but describes '12 hours of continuous use without charging'"), factual errors ("a company's 2023 revenue is 10 billion, but the actual public data is 8 billion").

[0050] For ease of understanding, the following examples are provided, but do not limit the present application. As an example, assume that the content to be audited is an industry information article about "2024 new energy vehicle technology trends", part of the content is as follows: "In 2024, a car company released a new electric vehicle equipped with A-type batteries, with a range of 800 kilometers, and a 10-minute charge to supplement 300 kilometers of range. The battery uses a silicon-based negative electrode material, with an energy density 50% higher than the previous generation, far exceeding the industry average of 30% improvement. It is understood that the technology is derived from the research results of the Nobel Prize in Chemistry winner in 2023."

[0051] The content auditing steps are:

[0052] I. Preliminary audit by the preset proofreading engine:

[0053] 1. Audit process: the engine calls "new energy vehicle terminology specification library", "date format library", etc. to scan the content through rule matching.

[0054] 2. Suspected errors are found:

[0055] (1) Term error: "Silicon-based negative electrode material" is the standard expression in the vocabulary library, and the engine marks it as "incomplete term expression."

[0056] (2) Format error: "2023 Nobel Prize in Chemistry" does not follow the standard notation for award time (it should be "2023 Nobel Prize in Chemistry," but the engine incorrectly judges that there is a missing "item" after "award" and marks it as "format not standardized."

[0057] 3. Output preliminary review report: includes the location, type, and preliminary basis of the two suspected errors.

[0058] II. Input quality inspection model for preliminary review report:

[0059] After receiving the preliminary review report, the model focuses on the two suspected errors, calls the internal "new energy battery technology library" and "Nobel Prize database," and prepares to conduct in-depth semantic analysis on the full text.

[0060] III. Semantic quality inspection by large model:

[0061] 1. False positive identification:

[0062] (1) For "incomplete term expression": through semantic analysis and in combination with the logical expression of "using silicon-based negative electrode material," the model confirms that "silicon-based negative electrode material" is a commonly used simplified expression in the industry and does not need to be forced to add "technology." It is determined to be a preliminary false positive (C-type observation item).

[0063] (2) For "format not standardized": the model searches the internal "award expression standard" and confirms that "2023 Nobel Prize in Chemistry" is the correct format and does not need to add "item" after "award." It is determined to be a preliminary false positive (C-type observation item).

[0064] 2. Missing supplement inspection:

[0065] The model conducts in-depth reading and finds that "the technology is derived from the research achievements of the winner of the 2023 Nobel Prize in Chemistry" has a factual error. Through the self-research agent calling the Nobel website, it is confirmed that the 2023 Nobel Prize in Chemistry focuses on "quantum dot technology" and is unrelated to the silicon-based negative electrode technology of Model A batteries, marked as "factual error" (B-type error).

[0066] It is found that "the energy density is 50% higher than the previous generation, far exceeding the industry average of 30% improvement" has a logical contradiction: through industry database retrieval, the model finds that the average improvement in the new energy battery industry in 2024 is about 40%, and the original "30%" data is incorrect, marked as "data accuracy error" (B-type error).

[0067] 3. Generate a quality inspection report:

[0068] Class A errors: none (no high-confidence errors agreed upon by the rule engine and the model).

[0069] Class B errors: 2 (factual errors, data accuracy errors), with modification suggestions ("2023 Nobel Prize in Chemistry" changed to "a university's new energy materials laboratory"; "30%" changed to "40%").

[0070] Class C observations: 2 (explain the reasons for the initial review mistakes).

[0071] Through the above process, the review of scientific and technological industry information not only quickly filters basic problems through the rule engine, but also solves deep semantic and factual problems through the large model, finally outputs accurate and transparent quality inspection results, and improves the efficiency and quality of content safety review.

[0072] This embodiment performs preliminary review on the content to be reviewed through the preset proofreading engine, and performs semantic quality inspection on the initial review report through the quality inspection large model, thereby forming a double review mechanism of preliminary review and semantic verification, and improving the detection rate and accuracy of content review.

[0073] Referring to Figure 2 , Figure 2 The second embodiment of the content review method of the present application is a flowchart, which is based on the first embodiment shown in the above Figure 1 The second embodiment of the content review method of the present application is a flowchart, which is based on the first embodiment shown in the above

[0074] In the second embodiment, before the step S30, it further includes:

[0075] Step S21: searching for relevant knowledge items related to the content to be reviewed in a knowledge base, wherein the knowledge base is pre-constructed based on a historical error correction library.

[0076] It should be understood that, in order to avoid missing key information by the model and improve the pertinence of quality inspection, in this embodiment, a prompt word is constructed according to at least one of the relevant knowledge items, the fact checking report and the initial review report, and the semantic quality inspection of the initial review report is performed by the quality inspection large model based on the prompt word, to obtain the review report of the content to be reviewed.

[0077] It can be understood that, in order to enable the quality inspection large model to follow the proofreading specification, a retrieval-augmented generation (RAG) process is adopted to query internal, deterministic and normative knowledge. This is the cornerstone of ensuring that the model understands the rules, and ensures that the quality inspection large model always takes the organization's "internal code" as the highest standard when making judgments, greatly improving the accuracy of detection of political and normative errors. The knowledge base can refer to a structured knowledge set constructed based on a historical error correction library, including internal specifications (such as industry terminology standards and important expression libraries), historical cases (such as typical error cases and false positive cases), and the like. Efficient retrieval is achieved through a vector database and a keyword index, and deterministic knowledge support is provided for quality inspection. The historical error correction library can refer to a database recording error types, error original texts, correct expressions and error correction bases in historical audits, and is the core basis for knowledge base construction, used to deposit proofreading experience and specifications. The relevant knowledge items can refer to knowledge content related to the content to be audited and the initial audit suspected errors retrieved from the knowledge base, including standard expressions, industry specifications, historical error correction cases, and the like, used to assist the quality inspection large model in judging content compliance. In specific implementation, the retrieval-augmented generation can be used to find relevant knowledge items related to the content to be audited in the knowledge base.

[0078] Further, in order to improve the accuracy of the relevant knowledge items and ensure that the query results are highly matched with the core theme of the content, and to avoid knowledge matching deviation, the step S21 includes: performing entity recognition and keyword extraction on the content to be audited to obtain target entities and target keywords; and performing similarity query in the knowledge base based on the target entities, the target keywords and the initial audit report to obtain relevant knowledge items related to the content to be audited.

[0079] It should be understood that entity recognition can refer to a process of automatically identifying and extracting entities with specific meanings from the content to be reviewed using natural language processing (NLP) technology, and the entities usually include proper nouns (such as technical names, product names), organization names, concept names, etc. Keyword extraction can refer to a process of extracting words or phrases that can reflect the core theme, key information or important content from the content to be reviewed, and the keywords are condensed expressions of the core meaning of the content. Target entities can refer to core entities extracted from the content to be reviewed through entity recognition, which are key elements with specific pointing in the content, such as technical names, product models, etc. Target keywords can refer to core words or phrases obtained from the content to be reviewed through keyword extraction, which are used to reflect the core theme and key information of the content. Similarity query can refer to a process of finding the most relevant knowledge items by calculating the semantic similarity between target entities, target keywords and knowledge items in the knowledge base, breaking through the limitations of traditional keyword exact matching, and realizing semantic-level association retrieval.

[0080] In a specific implementation, an NLP entity recognition algorithm based on a pre-trained language model (such as a BERT series model) is used to train the model to identify entities such as proper nouns, technical terms, product names, etc. in the content through labeled data. Entity recognition focuses on the words with specific pointing in the content, and preferentially extracts technical names (such as “A type battery”), material names (such as “silicon-based negative electrode material”), core concepts (such as “energy density”), etc. Combined with TF-IDF (Term Frequency-Inverse Document Frequency) algorithm, TextRank algorithm (a ranking algorithm based on text semantic association), and industry dictionary (such as new energy industry term library), the words or phrases reflecting the core of the content are extracted. Keyword extraction focuses on high-frequency words that play a decisive role in the theme of the content, such as “new energy vehicles”, “technology trends”, “improvement range”, etc., while referring to the words involved in suspected errors in the preliminary review report (such as “silicon-based negative electrode material”).

[0081] The knowledge base takes the historical error correction library as the core, integrates the industry term specification library, proofreading rule library, etc., and constructs a vector database (supporting semantic similarity retrieval) through vectorization processing, while retaining the keyword index (supporting exact match). The target entity, target keyword, and suspected errors in the preliminary review report (such as "silicon-based negative electrode material expression is incomplete") are converted into vectors as inputs for similarity queries. Through the cosine similarity algorithm of the vector database, the similarity between the input vector and the knowledge entry vector in the knowledge base is calculated, and the knowledge entries with a similarity higher than a preset threshold (such as 0.8) are selected. Combined with the suspected error types in the preliminary review report (such as "terminology expression problem"), the preliminarily selected knowledge entries are filtered again, and the industry specifications and historical cases related to the error types (such as "similar term simplification historical error correction cases") are preferentially retained. Obtain knowledge entries highly related to the content to be audited, including standard expression instructions, historical error correction cases, industry specification references, etc.

[0082] Step S22: calling a deep research agent to perform fact checking on the content to be audited, and obtaining a fact checking report.

[0083] It should be understood that in order to connect external real-time information sources, solve the problem of lagging knowledge base information (such as the latest industry data, current affairs dynamics), cover factual and timely errors, in this embodiment, a deep research agent is called to perform fact checking on the content to be audited, and a fact checking report is obtained. The deep research agent (Deep Research Agent) can refer to an intelligent agent that simulates the review of internal and external materials by experienced auditors and multiple verifications, has autonomous planning ability and a tool box (internal RAG query, network search engine, code interpreter, etc.), and is used for checking external dynamics, factual and timely problems. The fact checking report can be a report output by the deep research agent after verifying internal and external materials, containing the verification result of factual information, the source of the basis (such as authoritative source link, data report abstract), fact abstract, etc., used to support the factual judgment of the quality inspection large model.

[0084] In a specific implementation, the tool library of the deep research agent is called to perform fact checking on the content to be audited, and a fact checking report is obtained, wherein the preset tool library includes a real-time network search engine tool, a retrieval enhancement generation tool, and a code interpreter.

[0085] Further, in order to ensure systematic verification, the embodiment is designed by logical progressive steps to avoid randomness and omissions of fact verification, and to ensure the whole process specification from "information acquisition" to "conclusion verification". The step S22 includes: calling a deep research agent to plan research steps for the content to be audited, and obtaining a multi-step research plan, wherein the multi-step research plan includes a plurality of research steps with logical progressive relationship; calling a tool library of the deep research agent to perform fact verification on the content to be audited based on the multi-step research plan, and obtaining a fact verification report, wherein the preset tool library includes a real-time network search engine tool, a retrieval enhancement generation tool, and a code interpreter.

[0086] It can be understood that the research step plan can refer to an ordered verification step scheme autonomously designed by the deep research agent for factual problems in the content to be audited, for guiding subsequent tool calling and information verification. The multi-step research plan can refer to a verification plan generated by the deep research agent, containing a plurality of steps with logical progressive relationship, each step having a clear target, tool selection and expected result, ensuring the systematization of the verification process. The logical progressive relationship can refer to the sequential dependence relationship between steps in the multi-step research plan, i.e. the result of the previous step provides input or basis for the next step, such as "first search source → then extract information → finally compare and verify". The tool library can refer to a tool set that can be called by the deep research agent to support the whole process of fact verification, including a real-time network search engine tool, a retrieval enhancement generation tool (RAG tool), and a code interpreter. The real-time network search engine tool can refer to an API tool accessing a search engine, used to obtain the latest current information, authoritative source content, public data and other external dynamic information. The retrieval enhancement generation tool (RAG tool) can refer to a tool encapsulating the retrieval capability of the internal knowledge base, which can quickly call certain knowledge such as internal specifications and historical cases as the basis for fact verification. The code interpreter can refer to a tool running in a secure sandbox environment, supporting the execution of Python code, used for data calculation verification (such as growth rate, statistical data checking) and format conversion.

[0087] In a specific implementation, the in-depth research agent identifies the target for verification based on target entities in the content to be audited (such as "2023 Nobel Prize in Chemistry" and "energy density improvement range"), target keywords (such as "technology source" and "industry average data"), and suspected factual errors in the preliminary review report. The in-depth research agent designs logical progressive steps based on the verification target: Step 1: Clearly define information needs and determine core verification points (such as "research field of the 2023 Nobel Prize in Chemistry" and "average improvement range of new energy battery industry in 2024"); Step 2: Select tools according to the type of verification points (such as real-time network search engine for external dynamic information and RAG tool for internal specifications); Step 3: Design information extraction rules and clearly define the key elements to be extracted from the search results (such as "award field description" and "data report release agency and time"); Step 4: Plan verification logic and clearly define how to compare the extracted information with the content to be audited (such as "technology field match" and "data consistency"). The in-depth research agent calls tools in the order of multi-step research planning, executes the next step after completing the previous step, and ensures logical coherence. The role and execution process of each tool: Real-time network search engine tool: For external dynamic information (such as award field and latest industry report), input accurate keywords (such as "2023 Nobel Prize in Chemistry research field" and "2024 new energy battery energy density report"), and obtain search results from authoritative sources (such as official website and regular agency release page); Retrieval enhancement generation tool (RAG tool): For internal specifications or historical cases (such as "industry data expression standards"), call internal knowledge base to obtain relevant entries as verification reference; Code interpreter: For data calculation problems (such as "whether the improvement range conforms to statistical logic"), execute verification code (such as checking the sample range and calculation method of data source) to ensure data accuracy. Integrate the information obtained from each step to verify whether the factual statements in the content to be audited are accurate, and finally form a factual verification report containing "fact conclusion, source of evidence, and verification process".

[0088] For ease of understanding, the following examples are provided, but do not limit the present application. As an example, assume that the content to be audited is an industry information article about "2024 new energy vehicle technology trends", part of the content is as follows: "In 2024, a new electric vehicle released by a certain car company is equipped with A-type batteries, with a range of 800 kilometers and a 10-minute charge to supplement 300 kilometers of range. The battery uses silicon-based negative electrode materials, with an energy density improvement of 50% over the previous generation, far exceeding the industry average of 30% improvement. It is understood that this technology is derived from the research achievements of the winner of the 2023 Nobel Prize in Chemistry."

[0089] Factual issues to be verified: ① "technology derived from the 2023 Nobel Prize in Chemistry"; ② "industry average improvement range of 30%".

[0090] I. Call deep research agents to plan research steps, obtain multi-step research planning:

[0091] Trigger scenario: "Technical source" and "industry data" in the content to be reviewed belong to factual and timeliness issues, and the internal knowledge base is not fully covered, triggering research planning.

[0092] Multi-step research planning output:

[0093] Step 1: Goal: Confirm the research field of the 2023 Nobel Prize in Chemistry; Tool: Real-time web search engine tool; Expected result: Obtain the description of the research field of the 2023 Nobel Prize in Chemistry from the Nobel website or authoritative media; Dependency: None;

[0094] Step 2: Verify the relevance of A-type battery technology to the research field of the award; Tool: Real-time web search engine tool + RAG tool; Expected result: Retrieve the principle of A-type battery technology, compare with step 1 result, and determine whether it is related; Dependency: dependent on step 1 result;

[0095] Step 3: Obtain the average energy density improvement rate data of new energy batteries in 2024; Tool: Real-time web search engine tool; Expected result: Retrieve industry reports published by authoritative organizations (such as industry associations, third-party consulting companies) in 2024; Dependency: None;

[0096] Step 4: Verify data accuracy (such as report release time, sample range); Tool: Code interpreter + RAG tool; Expected result: Confirm data source compliance, and calculate logic consistent with industry standards; Dependency: dependent on step 3 result.

[0097] II. Call tools based on planning to conduct factual verification, obtain factual verification report:

[0098] Step 1 execution: Call the real-time web search engine tool to search the keyword "2023 Nobel Prize in Chemistry research field", and obtain the Nobel website information: "The 2023 Nobel Prize in Chemistry is awarded to quantum dot technology research, recognizing its contribution in the field of nanomaterials".

[0099] Step 2 execution: Call the real-time web search engine tool to retrieve "A-type battery technology principle", the result shows that its core is "structural innovation and silicon-based negative electrode material application"; Call RAG tool to obtain internal "technology field classification library", confirm that "quantum dot technology" and "battery structure material technology" belong to different fields; Compare the conclusion: A-type battery technology is not related to the 2023 Nobel Prize in Chemistry.

[0100] Step 3 execution: Call real-time web search engine tool to search "2024 new energy battery energy density improvement report", get "2024 new energy battery industry development white paper" published by XX industry association, which clearly states that "the industry average energy density improvement rate in 2024 is 40%".

[0101] Step 4 execution: Call code interpreter to verify data samples: white paper covers data of 30 major car companies, sample size meets industry standards; call RAG tool to confirm "industry average data expression standard", report release time is June 2024 (timeliness valid); verify conclusion: data is accurate and effective.

[0102] Fact-checking report output:

[0103] Fact 1 verification result: "The statement in the content to be audited that 'A type battery technology originated from the 2023 Nobel Prize in Chemistry' is incorrect. The research field of this award is quantum dot technology, which is unrelated to battery technology (based on the 2023 Nobel Prize announcement on the Nobel website).".

[0104] Fact 2 verification result: "The data in the content to be audited that 'industry average improvement rate is 30%' is incorrect. The average improvement rate of the new energy battery industry in 2024 is 40% (based on the XX Industry Association's '2024 New Energy Battery Industry Development White Paper').".

[0105] Through the above process, the deep research agent accurately identifies factual errors in the content to be audited through systematic step planning and multi-tool collaborative verification, providing reliable factual basis for subsequent semantic quality inspection of the quality inspection large model.

[0106] Further, in order to improve the success rate of research, the embodiment optimizes the planning by analyzing the failure reasons independently, solves the problems of invalid retrieval and tool mismatch, greatly improves the probability of obtaining effective information, and the tool library of the deep research agent based on the multi-step research planning is called to perform fact-checking on the content to be audited, and a fact-checking report is obtained, including: based on the multi-step research planning, the tool library of the deep research agent is called to perform progressive research on the content to be audited, and the research results of each research step are obtained; if the research result is research failure, the multi-step research planning is adjusted through the deep research agent, an adjusted multi-step research planning is obtained, and the step of returning to the tool library of the deep research agent based on the multi-step research planning to perform progressive research on the content to be audited and obtain the research results of each research step is performed.

[0107] It should be understood that the research results of each research step can refer to the specific achievements output by the deep research agent after executing each research step, including retrieved information, verified conclusions, data summaries, etc., which are the basis for supporting subsequent steps and final fact checking. Research failure can refer to the situation that the deep research agent fails to obtain effective information (such as no authoritative source in the search results, and data cannot be verified) or fails to achieve the goal of the step after executing a certain research step. Reflection and adjustment can refer to the process of the deep research agent autonomously analyzing the reasons for the failure (such as improper keywords, incorrect tool selection) and optimizing the research steps after the research failure, in order to improve the effectiveness of subsequent research. The adjusted multi-step research plan can refer to the new research plan generated after reflection and adjustment, which includes optimized step goals, tool selection, keywords or verification logic to solve the problems in the original plan.

[0108] In a specific implementation, the preset failure judgment criteria are, for example, “no authoritative source information is obtained after step execution”, “search results are irrelevant to the target”, “data verification fails and cannot be traced”, etc. When the research results meet any of the criteria, it is determined that the research has failed. The deep research agent autonomously analyzes the reasons for the failure, which can include: 1. Keyword problem: the search keywords are vague or imprecise (such as “industry data” without limiting “new energy battery”); 2. Incorrect tool selection: external dynamic information is required, but the internal RAG tool is called; 3. Step logic defect: necessary intermediate steps are missing (such as not first confirming the authority of the report publishing agency). The multi-step research plan is optimized according to the failure reasons, and the specific adjustment methods include: 1. Refining or replacing search keywords (such as changing “industry average increase rate” to “2024 new energy battery energy density average increase rate authoritative report”); 2. Tool replacement: adjusting the tool according to the information type (such as using a real-time web search engine tool for external dynamic information); 3. Adding auxiliary steps: supplementing source verification steps (such as first searching whether the report publishing agency is an industry-recognized authoritative agency). The adjusted multi-step research plan is returned to the progressive research step, and the research process is restarted until effective research results are obtained.

[0109] For ease of understanding, the following examples are provided, but do not limit the present application. As an example, assume that the fact-checking problem to be checked in the content to be audited is: ② “industry average increase rate of 30%”, and the original multi-step research plan for step 3 is “obtain 2024 new energy battery energy density average increase rate data, tool: real-time web search engine tool”.

[0110] I. Progressive research execution (first research failure scenario):

[0111] Step 3 execution: Call the real-time web search engine tool as planned, search for the keyword "industry average increase in 2024", and return results mostly as "general data of various industries" and "non-authoritative self-media articles". No authoritative report data on new energy battery field is obtained.

[0112] Result determination: Since no effective authoritative information is obtained, it is determined that the research is a failure.

[0113] II. Reflection, adjustment and re-research:

[0114] Failure analysis: The intelligent agent reflects that the reason for the failure is "imprecise keywords, not limited to the 'new energy battery' field and 'authoritative report' sources".

[0115] Adjusted multi-step research plan (step 3 optimization):

[0116] Step 3 (after optimization): Goal: Obtain authoritative data on the average increase in energy density of new energy batteries in 2024; Tool: Real-time web search engine tool; Expected result: Retrieve a 2024 special report published by an industry association or well-known consulting company; Dependency relationship: None;

[0117] Step 3.1 (new): Goal: Verify the authority of the report publishing agency; Tool: Real-time web search engine tool; Expected result: Confirm that the publishing agency is an industry-recognized authoritative agency (such as the China Automobile Industry Association); Dependency relationship: Dependent on Step 3 results.

[0118] III. Re-execute progressive research:

[0119] Step 3 execution: Call the real-time web search engine tool, search for the keyword "2024 new energy battery energy density average increase authoritative report industry association", and obtain the "2024 New Energy Vehicle Battery Technology Development Report" published by the China Automobile Industry Association.

[0120] Step 3.1 execution: Call the real-time web search engine tool to verify the publishing agency, and confirm that "the China Automobile Industry Association is an authoritative agency in the new energy industry, and the report has credibility".

[0121] Research results: The report clearly states that "the average increase in energy density of new energy batteries in 2024 is 40%", and the research is successful.

[0122] Through progressive research, the systematization of information acquisition is ensured, and the reflection adjustment mechanism solves the problem of imprecise keywords in the first search, ultimately successfully obtaining authoritative data, providing a reliable basis for fact-checking.

[0123] Further, in order to reduce the risk of missing detection, the embodiment ensures that hidden factual statements (such as expressions of implied causality) and time-sensitive information (such as content that requires the latest data without explicit time labeling) are fully captured through structured detection logic. Before the deep research agent is called to conduct fact checking on the content to be audited and a fact checking report is obtained, the following steps are further included: detecting whether there are factual statements in the content to be audited, and detecting whether there are time-sensitive information in the content to be audited; accordingly, the deep research agent is called to conduct fact checking on the content to be audited and a fact checking report is obtained, including: if there are factual statements in the content to be audited, or there are time-sensitive information in the content to be audited, or the problem cannot be solved through search enhancement generation, then the deep research agent is called to conduct fact checking on the content to be audited and a fact checking report is obtained.

[0124] It should be understood that a factual statement can refer to a statement or assertion of an objective fact in the content to be audited, including event correlation, data reference, technical source, causality, etc. Expressions that can be verified for authenticity, such as "a certain technology originates from a certain research achievement" and "a certain data is the industry average level". Time-sensitive information can refer to information in the content to be audited that is strongly related to time and dynamically changes over time, including the latest data, current events, job changes, policy updates, etc., the accuracy of which depends on the time node, such as "the industry average improvement rate in 2024" and "the latest research field of a certain award".

[0125] In a specific implementation, natural language processing (NLP) techniques are used in combination with pre-trained language models (such as BERT) and industry dictionaries to achieve automated detection. The factual statement detection logic: identifies assertion-type expressions in the content, focusing on sentences containing correlation words such as "derived from", "based on", "data is", "belongs to", etc. (such as "the technology is derived from the 2023 Nobel Prize in Chemistry"); through entity recognition, locate the core fact elements (such as technology name, event, data), and determine whether there is a verifiable objective fact correlation. The timeliness information detection logic: identifies time markers (such as "2024" and "latest") and dynamic attribute words (such as "industry average increase rate" and "position") in the content; combined with the internal knowledge base, determine whether the information is time-sensitive (such as "industry data" needs to refer to the latest report, not historical fixed values). The output result: mark the location of the factual statement (such as "Sentence 3: Technology derived from Nobel Prize in Chemistry") and the location of the timeliness information (such as "Sentence 3: 2024 industry average increase rate 30%") in the content to be reviewed. When any of the following conditions are met, trigger the call to the deep research agent to conduct fact checking on the content to be reviewed, and obtain a fact checking report: 1. There is a marked factual statement in the content to be reviewed; 2. There is a marked timeliness information in the content to be reviewed; 3. After searching the augmented generation (RAG) tool and internal knowledge base, insufficient information is obtained to answer the question (such as no latest industry data in 2024 in the knowledge base).

[0126] Step S23: constructing a prompt word according to at least one of the relevant knowledge item, the fact checking report, and the preliminary review report.

[0127] It can be understood that, in order to integrate scattered knowledge, verification results and preliminary review focuses into structured instructions, avoid missing key information by the model, and improve the relevance of quality inspection, in the embodiment, a prompt word is constructed according to at least one of the relevant knowledge item, the fact checking report, and the preliminary review report. The prompt word (Prompt) can be a structured text instruction guiding the quality inspection model to complete a specific task, including task objectives, background information, relevant knowledge, thinking chain requirements, etc., used to precisely constrain the output direction and content of the model.

[0128] In a specific implementation, the construction is based on: related knowledge items: provide internal norms and historical cases (such as "Simplified representation specification of silicon-based negative electrode materials"); fact checking report: provide external dynamic fact basis (such as Nobel Prize in Chemistry field, industry data); preliminary review report: clearly indicate suspected errors that need to be analyzed in depth (such as "Incomplete representation of silicon-based negative electrode materials" and "Non-standard format of Nobel Prize in Chemistry"). Prompt word structure: task instruction: "Identify false positives for suspected errors in the preliminary review report, and read the full text in depth to check for missing errors"; background information: content segment to be audited and core content of the preliminary review report; knowledge injection: related knowledge items (such as "Knowledge base shows that'silicon-based negative electrode materials' are industry-standard representations") and fact checking report abstract (such as "Nobel Prize website verifies that the 2023 Chemistry Prize focuses on quantum dot technology").

[0129] Correspondingly, the step S30 comprises:

[0130] Step S30': based on the prompt word, performing semantic quality inspection on the preliminary review report by the quality inspection large model to obtain an audit report of the content to be audited.

[0131] In a specific implementation, the quality inspection large model combines the knowledge and information in the prompt word to start semantic understanding (analyze context logic), logical reasoning (judge the rationality of the representation), and fact comparison (match knowledge items and verification results) capabilities. Based on the knowledge base items, determine whether the preliminary suspected errors meet the industry standards (such as "Silicon-based negative electrode material representation is reasonable, which is a false positive"); combined with the fact checking report, find errors that are not marked by the preliminary review (such as "Incorrect technology source, incorrect industry data"). Classify A / B / C type errors, with the analysis of the thought chain of each error (such as "According to knowledge base item X and fact checking report Y, XX is determined to be a false positive").

[0132] For ease of understanding, the following is an example, but does not limit the present application. As an example, assume that the content to be audited is an industry information article about "2024 new energy vehicle technology trends", part of the content is as follows: "In 2024, a certain car company released a new electric vehicle equipped with A type battery, with a range of 800 kilometers, and a 10-minute charge to supplement 300 kilometers of range. The battery uses silicon-based negative electrode materials, with an energy density 50% higher than the previous generation, far exceeding the industry average of 30% increase. According to reports, the technology is derived from the research of the winner of the 2023 Nobel Prize in Chemistry."

[0133] The preliminary review report marks two suspected errors: ① "Silicon-based negative electrode materials" are not fully represented; ② "2023 Nobel Prize in Chemistry" is not in the standard format. The specific steps of the content auditing method include:

[0134] I. Knowledge base search for related knowledge items:

[0135] Search process: Extract keywords "silicon-based anode materials", "terminology", and suspected errors in the initial review ①, and search for similar knowledge entries in the knowledge base.

[0136] Results: The relevant knowledge entry shows that "'silicon-based anode material' is a commonly used simplified term in the new energy industry. Similar expressions in historical cases (a news item in October 2023) were not judged as errors, and there is no need to force the addition of the word 'technology'."

[0137] II. Invoking deep research agents for fact-checking

[0138] Triggering scenario: The statement "the technology originates from the 2023 Nobel Prize in Chemistry" is a factual statement, and the statement "the industry average improvement rate is 30%" is time-sensitive data that is not covered by the knowledge base, thus triggering the deep research agent.

[0139] Execution process:

[0140] Planning steps: "Search the Nobel Prize in Chemistry 2023 winning fields on the official website → Verify the relevance of technologies → Search the 2024 battery industry report to obtain average improvement data";

[0141] Tool usage: Obtain information from the Nobel Prize website (2023 Chemistry Prize focuses on quantum dot technology) through online search, and access the "2024 New Energy Battery Industry Development Report" (released by XX institution);

[0142] Overall assessment: The source of the technology has been confirmed to be incorrect; the industry average improvement should be 40%.

[0143] Fact check report: Includes "The 2023 Nobel Prize in Chemistry did not involve 4680 battery technology" and "The average energy density of new energy batteries increased by 40% in 2024" and source links.

[0144] III. Constructing prompt words:

[0145] Example of prompt words:

[0146] "Task: Identify two suspected errors in the preliminary review report as false positives, and thoroughly review the entire report to fill in any gaps. Background: The content to be reviewed mentions 'silicon-based anode materials,' 'technology originates from the 2023 Nobel Prize in Chemistry,' and 'industry average improvement of 30%.' The preliminary review flags ① incomplete description and ② non-standard format. Knowledge Support: 1. The knowledge base shows that 'silicon-based anode materials' is a commonly used simplified expression in the industry (refer to the 2023 case); 2. Fact verification confirms: The 2023 Nobel Prize in Chemistry focused on quantum dot technology, which is unrelated to 4680 batteries; the industry average improvement in 2024 was 40%. Requirements: Clearly identify the error type and provide suggested corrections."

[0147] IV. Semantic quality inspection based on prompt words to obtain an audit report:

[0148] Model analysis process:

[0149] False positive identification: combined with knowledge base entries, it is determined that "silicon-based negative electrode material expression is incomplete" and "Nobel Prize in Chemistry format is not standard" are false positives of the preliminary review (consistent with industry expression habits);

[0150] Missing supplement inspection: according to the fact checking report, it is found that "technology source error" and "industry data error" are two newly added errors.

[0151] V. Audit report output:

[0152] Class A errors: none (rules engine and model have no common high confidence errors);

[0153] Class B errors (missing supplement inspection): 2, respectively, "technology source fact error" (suggested to be modified to "certain university new energy material laboratory") and "data accuracy error" (suggested to be modified to "40%") ;

[0154] Class C observation items (false positives): 2, reasons are "industry general simplified expression is reasonable" and "award format is in line with the standard".

[0155] Through the above process, the system realizes the deep integration of "internal specifications + external facts + intelligent reasoning", providing precise and transparent support for content security audit.

[0156] In this embodiment, the prompt words are constructed according to at least one of the relevant knowledge entries, the fact checking report and the preliminary review report, and the semantic quality inspection of the preliminary review report is performed based on the prompt words by the quality inspection large model to obtain the audit report of the content to be audited, so as to avoid missing key information by the model and improve the pertinence of quality inspection.

[0157] Reference Figure 3 , Figure 3 This is a flowchart of the third embodiment of the content audit method of the present application, based on the above embodiments, the third embodiment of the content audit method of the present application is proposed.

[0158] In the third embodiment, after step S30, it further includes:

[0159] Step S40: generating a model internal thinking chain corresponding to the quality inspection report, and displaying the quality inspection report and the model internal thinking chain, wherein the model internal thinking chain is used to represent the review and analysis process of the quality inspection large model generating the quality inspection report.

[0160] It should be understood that, in order to improve the credibility of the quality inspection large model, the embodiment makes the "black box" decision of the quality inspection large model transparent through the thinking chain, so that the user understands the judgment logic and enhances the trust in the suggestion of the quality inspection large model. Among them, the internal thinking chain of the model can refer to the structured review analysis process output by the quality inspection large model in the process of generating the quality inspection report, which includes five core links of "task identification, error positioning and classification, analysis basis statement, reasoning and argumentation, conclusion and suggestion", which is used to transparent AI decision logic and ensure traceability of judgment.

[0161] In a specific implementation, in the semantic quality inspection process, the quality inspection large model is forced to output the analysis process according to the preset structured framework: task identification: clearly define the current analysis target (such as "review the suspected errors of preliminary review + supplement the missing errors"); error positioning and classification: mark the location of suspected errors in the text and preliminarily classify them into one of the 25 error types (such as "fact error" "data accuracy error"); analysis basis statement: explain the source of the judgment basis (such as "internal knowledge base entry, fact checking report source"); reasoning and argumentation: elaborate the analysis logic (such as "verify the technical field mismatch through the Nobel official website, so determine that the fact is wrong"); conclusion and suggestion: give the final judgment (such as "confirm the error" "exclude the error") and specific modification suggestions. The large model integrates the error classification into three categories of A, B and C based on the conclusion of the thinking chain, and clearly defines the location, type, modification suggestion and basis summary of each error. The structured quality inspection report (error list in table form) and the internal thinking chain of the model (analysis process in text form) are output synchronously for manual review and reference.

[0162] For ease of understanding, the following examples are given, but do not limit the present application. As an example, assume that the content to be audited is an industry information about "2024 new energy vehicle technology trends", part of the content is as follows: "In 2024, a car company released a new electric vehicle equipped with A type battery, with a range of 800 kilometers, and a 10-minute charge to supplement 300 kilometers of range. The battery uses silicon-based negative electrode material, with an energy density 50% higher than the previous generation, far exceeding the industry average of 30% increase. It is understood that the technology comes from the research results of the Nobel Prize in Chemistry winner in 2023." The internal thinking chain of the model can be:

[0163]

Task identification

[0164]

Error positioning and classification

[0165] Preliminary suspected error 1: the 2nd sentence "silicon-based negative electrode material" is marked as "incomplete term expression";

[0166] Preliminary suspected error 2: the 3rd sentence "2023 Nobel Prize in Chemistry" is marked as "unspecified format";

[0167] Newly discovered suspected error 1: the third sentence "The technology originated from the 2023 Nobel Prize in Chemistry" (factual error);

[0168] Newly discovered suspected error 2: the third sentence "The industry average improvement rate is 30%" (data accuracy error).

[0169]

Analysis according to the statement

[0170] According to the internal knowledge base "Silicon-based negative electrode material is a general simplified expression in the industry";

[0171] According to the fact-checking report: the Nobel website focuses on quantum dot technology in the 2023 Nobel Prize in Chemistry, and the China Automobile Industry Association's 2024 report shows that the industry average improvement is 40%.

[0172]

Reasoning and argumentation

[0173] Suspected error 1: combined with the knowledge base, "silicon-based negative electrode material" is the industry's conventional expression, and there is no need to add "technology", and the judgment is false;

[0174] Suspected error 2: search for award format specifications, "2023 Nobel Prize in Chemistry" is the standard expression, and the judgment is false;

[0175] Newly discovered error 1: the technology field is unrelated to the research direction of the award, and the judgment is a factual error;

[0176] Newly discovered error 2: the industry data does not match the authoritative report, and the judgment is a data error.

[0177]

Conclusion and suggestion

[0178] C-class observation items (false positives): 2, reasons are shown in the above reasoning;

[0179] B-class errors (newly discovered): 2, suggestions are "change the technology source to a certain university laboratory" and "change the data to 40%".

[0180] Step S50: Responding to the user's artificial final audit decision based on the quality inspection report and the model's internal thought chain feedback.

[0181] It can be understood that the artificial final audit decision can refer to the user's (senior reviewer) final decision on the judgment result of the quality inspection large model based on the quality inspection report and the model's internal thought chain, which is the key feedback data for model reinforcement learning.

[0182] In a specific implementation, an interactive interface is provided, and a user selects a ruling result for each error and thought chain analysis in the quality inspection report and can supplement a textual description (such as a modification reason or an ignored basis). The association between the ruling result and the quality inspection report and the thought chain is automatically recorded to form structured feedback data (such as "Class B error 1 is adopted, and a reward of +1.0 is given").

[0183] Step S60: performing final review on the quality inspection report based on the artificial final review ruling to obtain a final review report.

[0184] It should be understood that, in order to ensure the authority of the review, the artificial final review is taken as the final basis in the embodiment, and it is ensured that the output result meets the industry specifications and expert standards. The final review report can be a final report formed by adjusting the quality inspection report based on the artificial final review ruling, and the final error type, modification suggestion, and ruling basis are explicitly confirmed, which is the final output result of the content review.

[0185] In a specific implementation, the quality inspection report is updated according to the artificial ruling result: the adopted error is kept in the original type and suggestion and is marked as "final review confirmation"; the modified error is updated in the suggestion content and is marked as "final review confirmation after modification"; and the ignored error is removed from the report or is classified as "final review confirmation false alarm" with an ignored reason. The final output file is formed by integrating the adjusted error list, the artificial ruling basis, and the core conclusion of the thought chain, and it is ensured that the content is traceable and auditable.

[0186] The embodiment makes the black-box decision of the quality inspection large model transparent through the thought chain, enables a user to understand the judgment logic, enhances the trust in the suggestion of the quality inspection large model, and takes the artificial final review as the final basis to ensure that the output result meets the industry specifications and expert standards, thereby ensuring the authority of the review.

[0187] Further, after the step S60, the method further includes:

[0188] Step S70: obtaining ruling data of the artificial final review ruling, and updating the knowledge base according to the ruling data.

[0189] It should be understood that, in order to improve the timeliness and accuracy of knowledge, the embodiment incorporates the latest norms and correct expressions confirmed by artificial final review into the knowledge base, solves the problem of knowledge lag, and ensures that the most authoritative reference can be called during subsequent quality inspection. Among them, the decision data can refer to the structured decision data generated by the user (senior reviewer) after the final review of the quality inspection report, including the decision result (adopt, modify, ignore), error type label, modification suggestion adjustment and decision reason, which is the core feedback information for continuous optimization of the system. The knowledge base can refer to a structured knowledge set pre-constructed based on historical error correction library, industry standard library, proofreading rule library, etc. After vectorization processing, it supports efficient retrieval and provides deterministic knowledge support for quality inspection, which can be continuously updated through external feedback.

[0190] In a specific implementation, the complete decision of artificial final review is recorded through an interactive interface, including: the decision result of each error (such as “adopt B class error 1” “confirm C class observation item as false positive”), the modified correct expression (such as “industry average improvement rate 40% (according to authoritative report)”), and the decision reason (such as “data source is verified as the latest report of industry association”). The decision data is stored in a structured manner as a standardized format of “original text fragment + error type + correct expression + decision basis + timestamp”, ensuring traceability. The correct expression confirmed by artificial (such as “2024 new energy battery energy density average improvement rate is 40%”) and industry norm supplement (such as “data reference needs to mark authoritative source”) are added to the knowledge base, and the vector database is updated synchronously to ensure the accuracy of semantic retrieval. The typical error cases confirmed by the decision (such as “fact-based error unrelated to technology source and award”) are classified into the corresponding error type library, supplemented with “error characteristics + correct logic + decision basis”, and the historical error correction case library is enriched. The frequently adopted knowledge items (such as authoritative data report reference specification) by artificial are given priority in retrieval to ensure that high credibility knowledge is called preferentially during subsequent quality inspection.

[0191] Step S80: constructing a training corpus based on the decision data, and training the quality inspection large model according to the training corpus.

[0192] It can be understood that, in order to align the expert preferences, the embodiment converts the implicit knowledge of artificial adjudication (such as risk judgment scale, suggestion expression style) into model parameters through reinforcement learning, so that the model output is more in line with the intuition and standards of experienced auditors. Among them, the training corpus can refer to the data set used for training the quality inspection large model, which contains structured information such as error original text, correct expression, audit thinking chain, and adjudication feedback, and is divided into positive example data (correct identification and suggestion), negative example data (false positive / missed case) and feedback data (artificial adjudication result). Quality inspection large model training can refer to the process of optimizing model parameters by using training corpus through supervised fine-tuning (SFT) and reinforcement learning based on human feedback (GRPO), so that the model improves the error identification ability and aligns the expert preferences.

[0193] In a specific implementation, the training corpus construction includes: 1. Positive example data expansion: converting "artificially adopted AI suggestions" into positive examples, including "error original text + AI correct identification process + artificial adoption mark" (such as "30% increase amplitude" → AI identifies data error and suggests "40%" → artificial adoption), with thinking chain and adjudication basis. 2. Negative example data supplement: converting "artificially marked false positives / missed" into negative examples, including "AI misjudgment original text + error analysis + artificial correction reason" (such as "AI does not identify 'technical source error' → missed analysis → artificial mark as factual error"). 3. Feedback data integration: dataize the reward signal required by GRPO reinforcement learning, such as "accurate detection (+1.0)" "effective clues (+0.5)" "missed (-1.0)" "false positives (-0.2)", and associate the corresponding adjudication results.

[0194] Quality inspection large model training optimization includes: 1. Supervised fine-tuning (SFT) update: add new positive and negative example data to the SFT training set, and optimize the model's identification ability for specific error types (such as data accuracy, factual correlation) through multiple iterations to ensure that the model masters the latest audit standards. 2. Reinforcement learning (GRPO) optimization: adjust model parameters based on reward signals in adjudication data. For example, increase the penalty weight for "missed factual error" cases to encourage the model to enhance its sensitivity to hidden errors such as technical sources and data authenticity; increase the reward for "artificially adopted modification suggestions" to make the model's suggestions more in line with expert preferences. 3. Model verification and deployment: after training, verify the model performance (such as detection rate and accuracy improvement) through the test set, update the online model after meeting the standards, and realize the closed loop of "feedback-training-optimization".

[0195] The embodiment incorporates the latest norms and correct expressions confirmed by artificial final review into the knowledge base, solves the problem of knowledge lag, ensures that the most authoritative reference basis can be called during subsequent quality inspection, and thus improves the timeliness and accuracy of knowledge. The embodiment also converts implicit knowledge of artificial adjudication into model parameters through reinforcement learning, so that the model output is more in line with the intuition and standards of experienced auditors, and thus can align with expert preferences.

[0196] Reference Figure 4 , Figure 4 FIG. 4 is a flowchart of a fourth embodiment of a content review method according to the present application. The fourth embodiment of the content review method according to the present application is based on the above embodiments.

[0197] In the fourth embodiment, before the step S20, the method further comprises:

[0198] Step S11: training a preset large model through supervised fine-tuning based on a preset data set to obtain a quality inspection basic large model.

[0199] It should be understood that, in order to enable the model to quickly master the basic knowledge and norms of proofreading, recognize high-frequency and explicit errors (such as term errors and format problems), generate a structured report, and solve the problem of “whether it can be done”, the embodiment trains a preset large model through supervised fine-tuning based on a preset data set to obtain a quality inspection basic large model. The preset data set can be a structured data set used for supervised fine-tuning, and contains at least one of positive example data, negative example data, synthetic data, and feedback data, and covers proofreading knowledge, error cases, and norm standards, which is the core material for the model to learn basic proofreading ability. Supervised fine-tuning can be a training method for guiding the model to learn a specific task through labeled data, which is used in the embodiment to inject basic proofreading knowledge, norms, and thinking logic into the preset large model, so that it has basic error recognition ability. The preset large model can be a general large model (such as a pre-trained model with basic natural language understanding ability) that has not been trained for a specific proofreading task, which serves as the starting point for training the quality inspection large model and provides basic language understanding and reasoning ability. The quality inspection basic large model can be a model obtained through supervised fine-tuning, which has mastered the basic knowledge, process norms, and standard output format of proofreading work, can recognize common errors and generate a structured quality inspection report, but has not yet aligned with human expert preferences.

[0200] In a specific implementation, starting with a preset large model with basic language capabilities, the preset data set is input into the model in batches, and the model parameters are adjusted through the back propagation algorithm. The training target focuses on "basic capability shaping", enabling the model to master the definitions and characteristics of 25 types of errors, the structured output logic of proofreading thinking chain (task identification → error positioning → statement according to → reasoning → conclusion), and industry standards (such as term expression, format standard). Output requirements: Force the model to output results with structured "thinking chain", ensuring that the judgment process is interpretable, for example, "identify'silicon-based negative electrode material' expression → search industry term library → confirm that the simplified expression is reasonable → conclusion: exclude errors". The obtained quality inspection basic large model can stably identify common errors and generate structured quality inspection reports (such as JSON format) containing thinking chains, but may deviate from human expert intuition in judging fuzzy errors in complex contexts.

[0201] Further, in order to provide high-quality, multi-dimensional training materials for supervised fine-tuning, and to ensure that the model can master the basic knowledge of proofreading, specifications and case logic from the early stage of learning, before the step S11, the method further comprises: generating generation data corresponding to the content review task by the general large model; obtaining feedback data from historical decision data, wherein the historical decision data is decision data before the current review; and constructing a preset data set based on at least one of positive example data, negative example data, the generation data, and the feedback data.

[0202] It can be understood that the general large model can refer to a pre-trained large model with strong natural language understanding and generation capability (such as GPT-4o), which is used in this embodiment to generate simulated training data according to the content review task requirements, and provide diversified learning samples for the quality inspection model. The generated data can refer to the simulated data generated by the general large model according to the error type definition of the content review task, which includes error original text, error type label, review thinking chain and modification suggestion, and is used to enrich the diversity and coverage of the training data. The historical decision data can refer to the final decision records made by the artificial review link on the past contents to be reviewed before the current review task, including decision results (adopt, modify, ignore), error type judgment, modification suggestion and decision reason, which is the core source of feedback data. The feedback data can refer to the data extracted from the historical decision data and cleaned and labeled, including “error original text + AI preliminary suggestion + artificial decision result + decision basis”, which is used to reflect the artificial review preference and optimize the judgment logic of the model. The positive example data can refer to the data containing real errors and correct review results, with the structure of “error original text + error type + correct modification suggestion + review thinking chain”, which is used to train the model to identify and correct correct errors. The negative example data can refer to the data containing false positive cases of the rule engine or the model, with the structure of “suspected error original text + context analysis + artificial exclusion reason”, which is used to train the model to identify false positives and avoid over-sensitivity. The preset data set can refer to the structured training data set constructed for the supervised fine-tuning (SFT) of the quality inspection large model, which integrates positive example data, negative example data, generated data and feedback data, and is used to inject review knowledge, specifications and case experience into the model.

[0203] In a specific implementation, 25 common error types (such as data accuracy error, factual error, terminology expression error, etc.) defined in the content review task are taken as the core, and the characteristics, scenarios and judgment standards of each error type are determined as the instruction basis for the general large model to generate data. The structured instructions are input to the general large model to clearly generate the target, for example: “Please generate ‘data accuracy error’ case: contains a piece of technology industry information segment, which implicitly contains an error in the energy density improvement amplitude data; and output ‘error original text + error type + correct modification suggestion + review thinking chain analysis’”. After the general large model generates a large number of simulated cases according to the instructions, artificial or automatic tools filter the generated data to select cases that meet the real review scenarios and clear error characteristics, to ensure data quality. The generated data is uniformly structured and stored as “error original text segment + error type label + correct expression + thinking chain analysis (including judgment basis and reasoning process)”, which is consistent with the format of real cases.

[0204] Extract the manual final review decision records before the current review task, covering different content types (such as technical information, industry reports), and the decision results of different error types, to ensure data coverage. Remove invalid or duplicate decision records (such as invalid annotations with format errors), and correct ambiguous expressions (such as supplementing unclear decision reasons), to ensure data integrity. Convert historical decision data into a structured format of "to-be-reviewed original text segment + AI initial quality inspection suggestion + manual decision result (adopted / modified / ignored) + decision reason + error type final label", highlighting the modification logic of manual correction of AI suggestions. Extract key information from structured historical decision data, focusing on retaining "AI misjudgment cases (such as missed reports, false positives)" and "manual modification suggestion cases" with learning value as key feedback for model optimization.

[0205] The preset data set needs to include at least one or more of the following data types to achieve complementary coverage: positive example data: selected from real proofreading cases, "AI correctly identifies errors and suggestions are adopted", reflecting correct proofreading logic; negative example data: "rule engine or AI false positives" cases from historical cases, with context analysis and manual exclusion reasons, used to train false positive identification ability; generated data: simulated error cases generated by general large models, supplementing error types or scenarios not covered in real data; feedback data: manual correction cases extracted from historical decision data, reflecting expert proofreading preferences and implicit knowledge. Format all integrated data uniformly to ensure that each piece of data contains "input (error original text / suspected error)" "output (correct result / judgment reason)" "label (error type / decision result)" three core elements for model learning. Divide the preset data set into training set (used for model parameter optimization) and validation set (used for performance evaluation during training) according to the proportion, usually training set accounts for 70%-80%, validation set accounts for 20%-30%, to ensure the model training effect is verifiable.

[0206] Step S12: Adjust the internal parameters of the quality inspection basic large model through reinforcement learning training to obtain a quality inspection large model.

[0207] It should be understood that, in order to align the expert preferences, the embodiment optimizes the model parameters through artificial feedback, so that the decision of the model in the fuzzy scene (such as the rationality of term simplification and the judgment of data accuracy) is consistent with that of the experienced expert, and the problem of “whether to do well” is solved. The reinforcement learning training can refer to a training mode of adjusting model parameters based on human feedback. In the embodiment, a GRPO (Generative Reinforcement Learning with Policy Optimization) framework is adopted to convert the artificial final review decision into a reward signal, and the judgment logic and output preference of the model are optimized. The quality inspection large model can be a final model obtained after two-stage training of supervised fine-tuning and reinforcement learning, which has deep semantic understanding, fact checking, autonomous research ability, and can align with human expert preferences, and can output accurate and transparent quality inspection results.

[0208] In a specific implementation, the historical decision data is converted into a reward score, wherein the historical decision data is decision data before the current review; and the internal parameters of the quality inspection basic large model are adjusted based on the reward score through reinforcement learning training to obtain the quality inspection large model. The specific reinforcement learning training method is as follows:

[0209] 1. Reinforcement learning framework design:

[0210] Environment (Environment): The manuscript to be reviewed, the preliminary review report, the suggestion report generated by the quality inspection basic large model, and the artificial review interaction interface constitute the scene background of model training.

[0211] Agent (Agent): The quality inspection basic large model is the object of optimization for accepting the reward signal.

[0212] Action (Action): The complete quality inspection report (including error identification, modification suggestion and thinking chain) generated by the model for the content to be reviewed.

[0213] Reward function (Reward Function): The scoring mechanism designed based on the artificial final review decision is the core of reinforcement learning, including:

[0214] Accurate detection (+1.0): The model finds true errors and the suggestion is adopted by the artificial one-key;

[0215] Effective clues (+0.5): The model finds true errors, but the suggestion is adopted after being modified by the artificial;

[0216] Major mistakes-miss (-1.0): The model does not find the errors manually found by the artificial;

[0217] Minor mistakes-misreport (-0.2): The model suggestion is marked as a misreport by the artificial.

[0218] 2. Strengthen the implementation of learning:

[0219] Collect data from human final judgments and convert it into reward points for corresponding "actions".

[0220] The GRPO algorithm is used to adjust the neural network weights of the basic quality inspection model according to the reward score: strengthen the corresponding parameters for high reward behaviors (such as accurate detection) and weaken the corresponding parameters for penalty behaviors (such as false negatives).

[0221] The training objective focuses on "preference alignment": to make the model's judgment criteria, risk sensitivity, and modification suggestion style as close as possible to the intuition of senior reviewers, achieving a leap from "following the rules" to "understanding the intent".

[0222] 3. Training Output: A large-scale quality inspection model was obtained, which significantly improved the accuracy of error identification and the rate of suggestion adoption, and the judgment logic was highly consistent with that of human experts.

[0223] This embodiment trains a pre-set large model based on a pre-set dataset through supervised fine-tuning to obtain a basic large model for quality inspection. This enables the model to quickly master the basic knowledge and standards of review and verification, and to identify high-frequency and clear errors. Furthermore, this embodiment optimizes the model parameters through human feedback, so that the model's decision-making in ambiguous scenarios is consistent with that of senior experts.

[0224] For ease of understanding, please refer to Figure 5 This explanation is provided, but does not limit the scope of this application. Figure 5 This is a flowchart illustrating one embodiment of the content review method for this application. Figure 5 The overall approach to the content review method in this application is to construct a two-stage, closed-loop feedback collaborative workflow with the ability to self-evolve and continuously learn. The verification engine, as the first line of defense, achieves millisecond-level initial review through rule matching and template comparison, handling common errors. The large-scale model review and quality inspection, as the second line of defense, dynamically backtests the review results through full-text semantic understanding, forming a dual-insurance mechanism of "rule filtering + semantic verification" to improve the detection rate and accuracy of the review.

[0225] 1. First stage: Initial review by millisecond-level rule engine:

[0226] The verification engine serves as the first line of defense.

[0227] Relying on standardized thesaurus (important expression library, proofreading rule library, and historical error correction library) operated in the cloud and user-defined thesaurus (such as enterprise-specific expression rules and product name specification library), the document is scanned and initially filtered at high speed through efficient rule matching, regular expression and template comparison technology.

[0228] The core advantage of this stage is speed and certainty. It can process high-frequency, explicit, and patterned errors in massive texts at millisecond level, such as fixed collocations of important expressions, absolute errors in blacklists, and formatted number usage, etc., to complete the basic and large-scale "cleaning" work. The essence of this stage is "matching" rather than "understanding". Therefore, it cannot handle errors that require deep semantic understanding, contextual logical reasoning, and cross-domain background knowledge.

[0229] 2. Second stage: Large model deep semantic quality inspection:

[0230] Introduce a "final review expert" large model trained by multiple stages and multiple paradigms, with autonomous research ability and alignment with human preferences (hereinafter referred to as quality inspection large model), as the second and most critical line of defense for quality control. Its core tasks include:

[0231] (1) False report identification: All suspected errors marked by preliminary review are subjected to accurate context understanding and fact checking to eliminate false positives.

[0232] (2) Missing inspection: Deeply read the entire text, not only check the language level, but also cross-verify the facts, timeliness, and logic to uncover the most hidden errors.

[0233] The integration of the two stages will ultimately present an "intelligent proofreading report" to the artificial review link, which combines the results of the two stages. The report will clearly mark:

[0234] A class of errors: high confidence errors confirmed by rule engine and "quality inspection large model".

[0235] B class of errors: newly discovered errors by "quality inspection large model" through deep semantic understanding and autonomous research.

[0236] C class of observations: "quality inspection large model" believes that the preliminary review result has false positives, and provides detailed analysis reasons.

[0237] 3. Synergy and self-evolution closed loop.

[0238] From "rule filtering + semantic verification" to "rule preliminary screening + autonomous research verification + alignment of human preferences", a three-fold insurance mechanism. Each final decision (adopt, ignore or modify suggestions) of artificial review will be captured by the system as the most valuable feedback data, which will be used to train and shape the "quality inspection large model" in real time and continuously through reinforcement learning mechanism, so that its judgment criteria are closer to or even surpass the personal preferences and intuition of experienced editors, forming a self-evolution closed loop from "following rules" to "understanding intentions".

[0239] For ease of understanding, please refer to Figure 6The following description is provided to assist in understanding the present application and is not intended to limit the application. Figure 6 The training architecture diagram of the quality inspection large model for an embodiment of the content review method is as follows, Figure 6 In the present application, the construction of the quality inspection large model adopts a hierarchical and progressive technology stack, each layer solves a specific problem, and together builds the professional quality inspection capability, and the specific steps are as follows:

[0240] I. Quality inspection large model training: SFT and GRPO double-stage evolution paradigm.

[0241] The present application divides the training of the model into two core stages, so that it grows from an "apprentice" with basic instruction following ability to a "master" who understands the intentions of human experts.

[0242] Stage 1: Supervised fine-tuning (SFT), injecting expert knowledge and specifications. The goal is to enable the model to master the basic knowledge, workflow and standard output format of the review work. This is a solid foundation for the model's ability to ensure that it "can do" the work.

[0243] 1. High-quality dataset construction:

[0244] (1) Positive data: Use existing historical decision data. "Error original text + error type" as input, "correct modification suggestion + corresponding review thinking chain analysis" as expected output.

[0245] (2) Negative data (false positive data): Artificially sorted a batch of typical false positive cases generated by the rules engine in history, and train the model to learn to make "exclusion error" judgments in specific contexts combined with the context.

[0246] (3) Synthetic data generation: Use higher-order general large models (such as GPT-4o) to generate large-scale and diversified synthetic data according to detailed 25 error definitions. For example, give the model a correct sentence, instruct it to "generate 'data accuracy error' cases: include a piece of technology industry information, which contains a hidden energy density improvement data error; At the same time output 'error original text + error type + correct modification suggestion + review thinking chain analysis'".

[0247] (4) Feedback data closed loop: After the initial system goes online, the human review link to the AI recommendation decision data is continuously cleaned, labeled, and added to the SFT training set.

[0248] Training output: Through SFT, a quality inspection basic model is obtained that can skillfully identify various errors and stably generate structured quality inspection reports (usually in JSON format) containing thinking chains.

[0249] Phase 2: Reinforcement Learning Based on Human Feedback, Aligning Human Final Review Preferences. The goal is to make the model not just "do it," but "do it well" and "do it cleverly." Make its judgment criteria, risk preferences, and even the "tone" of modification suggestions infinitely close to or even surpass the personal preferences and intuition of experienced auditors.

[0250] Introduce the GRPO reinforcement learning framework to convert the final decision of the artificial review link into the most direct and effective reward signal. Environment: Manuscript to be reviewed + preliminary review results + "quality inspection big model" generated report + interaction interface of artificial review link. Agent: Quality inspection big model. Action: The model generates a complete review report for the manuscript (including all errors it considers and modification suggestions). Reward function: This is the core of the entire evolution mechanism. When the artificial review link completes the final decision, the system will automatically calculate a comprehensive reward score for this "action." For example:

[0251] Precise detection (high positive score + 1.0): The model finds a true error and the modification suggestion is directly adopted by the artificial review link.

[0252] Effective clues (medium positive score + 0.5): The model finds a true error, but the modification suggestion is adopted after being modified by the artificial review link. This indicates that the model has found the problem, but the proposed solution is not perfect.

[0253] Major mistake - missed report (high negative score -1.0): The model fails to find an error (False Negative) manually found by the artificial review link. This is the most punishable behavior because it directly affects the bottom line of the review quality.

[0254] Minor mistake - false report (low negative score -0.2): The model's suggestion is directly ignored or marked as false positive (False Positive) by the artificial review link. Although it is not as serious as a missed report, it also wastes the time and attention of the artificial review.

[0255] Training process: The GRPO algorithm will adjust the internal parameters (neural network weights) of the "quality inspection big model" based on the comprehensive reward score of each manuscript. This process will encourage the model to generate review reports that can obtain high rewards (i.e. more accurately find true errors and propose more recognized modification suggestions), and strive to avoid actions that result in punishment (especially missed reports).

[0256] To ensure the model output is interpretable, traceable, and reliable, during training, the "quality inspection large model" must generate a structured "proofreading analysis process", i.e. "thinking chain", before giving any conclusion. For example, in the prompt design, it is clearly required that the model strictly follow the following thinking path:

[0257]

Task identification

[0258]

Error positioning and classification

[0259]

Analysis basis statement

[0260]

Reasoning and argumentation

[0261]

Conclusion and suggestion

[0262] The "thinking chain" not only greatly improves the accuracy of the judgment, but more importantly, it completely transparentizes the "black box" decision-making process of AI, making the human review process understand, trust and efficiently use the suggestions of AI.

[0263] II. Knowledge enhancement: knowledge retrieval engine

[0264] To make the "quality inspection large model" comply with the proofreading specifications, the internal, deterministic, and normative knowledge query is processed by the retrieval enhancement generation (RAG). This is the cornerstone of ensuring that the model "understands the rules", ensuring that the "quality inspection large model" always takes the organization's "internal code" as the highest standard when making judgments, greatly improving the accuracy of political and normative error detection. The implementation scheme is as follows:

[0265] 1. Knowledge base preprocessing: The historical decision data is comprehensively vectorized and constructed into an efficient vector database. For structured rules, key word indexes can also be retained.

[0266] 2. Real-time retrieval: When processing a text, the system performs entity recognition (personal names, place names, organizations, and proper nouns) and key word extraction on the text content.

[0267] 3. Triggered query: Use the extracted keywords, entities, and suspected errors marked in the first stage to perform similarity retrieval in the vector database, and query the most relevant knowledge entries.

[0268] 4. Dynamic injection prompt: Dynamically and structurally inject the retrieved knowledge (e.g., "The standard expression should be XXX", "According to the whitelist, this word is allowed to use here", "Historical cases show that this type of expression is prone to errors, pay attention to XXX") into the prompt word (Prompt) given to the quality inspection large model.

[0269] III. External fact difficult problem verification:

[0270] To handle external, dynamic, and cross-verified fact-based and time-sensitive problems in the content security review process, a self-developed deep research agent (Deep Research Agent) is designed, which simulates the research process of a senior reviewer when encountering difficult problems, consulting internal and external materials, and seeking evidence from multiple parties. The implementation scheme is as follows:

[0271] Build an intelligent agent (Agent) based on a large model, and give it autonomous planning capabilities and a powerful toolbox.

[0272] 1. Toolbox:

[0273] (1) Internal RAG query tool: encapsulate the RAG capabilities of scheme one as the preferred query tool.

[0274] (2) Real-time network search engine tool: API access to search engines to obtain the latest current information, authoritative source content, public data, and other external dynamic information.

[0275] (3) Code interpreter: a safe sandbox environment for executing Python code for simple numerical calculations, data verification (such as whether the growth rate mentioned in the article and statistical data are calculated correctly), and format conversion.

[0276] 2. Workflow:

[0277] (1) Triggering condition: When the content to be reviewed is a fact-based or time-sensitive problem, or the internal knowledge base is not fully covered, the deep research agent will be automatically triggered.

[0278] (2) Autonomous planning: It can autonomously generate a multi-step research plan based on the problem. For example, for an industry information article on "2024 new energy vehicle technology trends", the plan can be: Step 1: Goal: Identify the research field of the 2023 Nobel Prize in Chemistry; Tool: Real-time web search engine tool; Expected result: Obtain the description of the research field of the 2023 Nobel Prize from the official website or authoritative media; Dependency: None; Step 2: Verify the relevance of A-type battery technology to the research field of the award; Tool: Real-time web search engine tool + RAG tool; Expected result: Retrieve the principles of A-type battery technology and compare them with the results of Step 1 to determine if they are related; Dependency: dependent on Step 1 result; Step 3: Obtain the average increase in energy density of new energy batteries in 2024; Tool: Real-time web search engine tool; Expected result: Retrieve industry reports published by authoritative organizations (such as industry associations, third-party consulting companies) for 2024; Dependency: None; Step 4: Verify data accuracy (such as report release time, sample range); Tool: Code interpreter + RAG tool; Expected result: Confirm that the data source is compliant and the calculation logic conforms to industry standards; Dependency: dependent on Step 3 result.

[0279] (3) Multi-step execution and reflection: The agent sequentially calls tools according to the plan. If a step fails (such as a fruitless search), it can reflect and adjust the plan (such as replacing keywords and searching again).

[0280] (4) Comprehensive research and judgment: Integrate multiple sources of information (web information, internal specifications, calculation results) to form a high-confidence "fact-checking report", then compare it with the original text to accurately determine whether there are "fact errors" or "current information expression errors".

[0281] The technical solution proposed in this application can achieve the following effects:

[0282] 1. Deep fact-checking and accurate error mining capability: This application goes beyond simple text matching by integrating the semantic understanding of large models and "autonomous research agents" (Deep Research Agent) to actively connect internal and external knowledge bases for fact and timeliness verification, thereby accurately mining hidden, logical, and factual errors that traditional tools cannot detect, and intelligently identifying false positives.

[0283] 2. Transparent and traceable human-machine collaboration capability: With the "review thinking chain" mechanism, each judgment of the system is accompanied by clear decision-making basis and reasoning process. This makes the "black box" of AI completely transparent, allowing human editors to understand, trust, and efficiently adopt AI suggestions, achieving smooth and efficient human-machine collaboration.

[0284] 3. Adaptive evolution ability to align expert preferences: Through human feedback-based reinforcement learning (GRPO), the system can continuously learn and evolve from each final review decision made by human editors. This enables its judgment criteria and proofreading "intuition" to constantly align with the real preferences of senior experts, realizing the transition from "following rules" to "understanding intentions", and becoming smarter with use.

[0285] It should be noted that the above examples are only for understanding the present application and do not constitute a limitation on the content review method of the present application. Based on this technical concept, more forms of simple transformation are within the protection scope of the present application.

[0286] The present application also provides a content review device, please refer to Figure 7 , the content review device comprises:

[0287] The preliminary review module 10 is configured to perform preliminary review on the to-be-reviewed content through a preset proofreading engine in response to the input to-be-reviewed content, and obtain a preliminary review report.

[0288] The model calling module 20 is configured to input the preliminary review report into a quality inspection large model.

[0289] The model quality inspection module 30 is configured to perform semantic quality inspection on the preliminary review report through the quality inspection large model, and obtain a review report of the to-be-reviewed content.

[0290] The content review device provided by the present application adopts the content review method in the above embodiments, which can solve the technical problems of low detection rate and low accuracy of related content review methods. Compared with the prior art, the content review device provided by the present application has the same beneficial effects as the content review method provided by the above embodiments, and other technical features in the content review device are the same as the features disclosed in the above embodiments, which will not be repeated here.

[0291] The present application provides a content review device, which comprises at least one processor and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the content review method in the above embodiment one.

[0292] The following will be described with reference to Figure 8The diagram illustrates a structural schematic of a content moderation device suitable for implementing embodiments of this application. The content moderation device in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 8 The content moderation device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0293] like Figure 8 As shown, the content moderation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the content moderation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touch screens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, LCDs (Liquid Crystal Displays), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the content moderation device to communicate wirelessly or wiredly with other devices to exchange data. While the figures show content moderation devices with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.

[0294] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program codes for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network through a communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiments disclosed in the present application are executed.

[0295] The content auditing device provided by the present application adopts the content auditing method in the above-mentioned embodiments, and can solve the technical problems of low detection rate and low accuracy rate of related content auditing methods. Compared with the prior art, the content auditing device provided by the present application has the same beneficial effects as the content auditing method provided by the above-mentioned embodiments, and other technical features in the content auditing device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.

[0296] It should be understood that various parts of the present application can be realized by hardware, software, firmware or a combination thereof. In the description of the above-mentioned embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0297] The above is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0298] The present application provides a computer readable storage medium having stored thereon computer readable program instructions (i.e. computer program) for executing the content auditing method in the above-mentioned embodiments.

[0299] The computer readable storage medium provided in the present application may, for example, be a U disk, but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, system, or device, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more conductive wires, a portable computer diskette, a hard disk, a RAM (Random Access Memory), a ROM (Read Only Memory), an EPROM (Erasable Programmable Read Only Memory) or a flash memory, an optical fiber, a CD-ROM (CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present embodiment, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer readable storage medium can be transmitted in any suitable medium, including but not limited to an electrical wire, an optical cable, an RF (Radio Frequency), and the like, or any suitable combination of the above.

[0300] The above computer readable storage medium can be contained in the content review device, or can exist separately and not be assembled into the content review device.

[0301] The above computer readable storage medium carries one or more programs, which, when executed by the content review device, cause the content review device to perform the above content review method.

[0302] Computer program code for carrying out operations of the present application can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0303] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0304] The modules involved in the embodiments of the present application can be implemented in the form of software or in the form of hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.

[0305] The readable storage medium provided by the present application is a computer readable storage medium, which stores computer readable program instructions (i.e. computer programs) for executing the content review method described above, and can solve the technical problems of low detection rate and low accuracy of related content review methods. Compared with the prior art, the computer readable storage medium provided by the present application has the same beneficial effects as the content review method provided by the above-mentioned embodiments, and will not be described here.

[0306] The present application also provides a computer program product comprising a computer program, which, when executed by a processor, implements the content review method as described above.

[0307] The computer program product provided by the present application can solve the technical problems of low detection rate and low accuracy of related content review methods. Compared with the prior art, the computer program product provided by the present application has the same beneficial effects as the content review method provided by the above-mentioned embodiments, and will not be described here.

[0308] The above only describes some embodiments of the present application, and does not limit the scope of the present application. Any equivalent structural transformation made by using the content of the present application specification and drawings, or direct / indirect application in other related technical fields is included in the protection scope of the present application.

[0309] The present application discloses A1, a content review method, the content review method comprising:

[0310] perform preliminary review on the to-be-reviewed content through a preset proofreading engine to obtain a preliminary review report;

[0311] input the preliminary review report into a quality inspection large model;

[0312] perform semantic quality inspection on the preliminary review report through the quality inspection large model to obtain a quality inspection report of the to-be-reviewed content.

[0313] A2, the content review method of A1, before the semantic quality inspection on the preliminary review report through the quality inspection large model to obtain the review report of the to-be-reviewed content, further comprising:

[0314] search for a relevant knowledge item related to the to-be-reviewed content in a knowledge base, wherein the knowledge base is pre-constructed based on a historical error correction library;

[0315] invoke a deep research agent to perform fact checking on the to-be-reviewed content to obtain a fact checking report;

[0316] construct a prompt word according to at least one of the relevant knowledge item, the fact checking report, and the preliminary review report;

[0317] Accordingly, the semantic quality inspection on the preliminary review report through the quality inspection large model to obtain the review report of the to-be-reviewed content, comprising:

[0318] perform semantic quality inspection on the preliminary review report through the quality inspection large model based on the prompt word to obtain the review report of the to-be-reviewed content.

[0319] A3, the content review method of A2, the invocation of the deep research agent to perform fact checking on the to-be-reviewed content to obtain a fact checking report, comprising:

[0320] invoke a deep research agent to perform research step planning on the to-be-reviewed content to obtain a multi-step research planning, wherein the multi-step research planning comprises a plurality of research steps having a logical progressive relationship;

[0321] invoke a tool library of the deep research agent based on the multi-step research planning to perform fact checking on the to-be-reviewed content to obtain a fact checking report, wherein the preset tool library comprises a real-time network search engine tool, a retrieval enhancement generation tool, and a code interpreter.

[0322] A4, the content review method of A3, the invocation of the deep research agent based on the multi-step research planning to perform fact checking on the to-be-reviewed content to obtain a fact checking report, comprising:

[0323] Based on the multi-step research plan, the tool library of the deep research agent is used to perform progressive research on the content to be reviewed, and research results of each research step are obtained.

[0324] If the research result is research failure, the multi-step research plan is adjusted by the deep research agent, an adjusted multi-step research plan is obtained, and the step of performing progressive research on the content to be reviewed by the tool library of the deep research agent based on the multi-step research plan is returned to obtain research results of each research step.

[0325] A5, the content review method of A2, before the step of calling the deep research agent to perform fact checking on the content to be reviewed and obtaining a fact checking report, further comprising:

[0326] Detecting whether there is a factual statement in the content to be reviewed and detecting whether there is time-sensitive information in the content to be reviewed;

[0327] Correspondingly, the step of calling the deep research agent to perform fact checking on the content to be reviewed and obtaining a fact checking report comprises:

[0328] If there is a factual statement in the content to be reviewed, or there is time-sensitive information in the content to be reviewed, or the problem cannot be answered through retrieval enhancement, the deep research agent is called to perform fact checking on the content to be reviewed, and a fact checking report is obtained.

[0329] A6, the content review method of A2, wherein the step of searching for relevant knowledge items related to the content to be reviewed in the knowledge base comprises:

[0330] Performing entity recognition and keyword extraction on the content to be reviewed to obtain target entities and target keywords;

[0331] Based on the target entities, the target keywords, and the preliminary review report, a similarity query is performed in the knowledge base to obtain relevant knowledge items related to the content to be reviewed.

[0332] A7, the content review method of any one of A1 to A6, after the step of performing semantic quality inspection on the preliminary review report by the quality inspection large model based on the prompt word to obtain the review report of the content to be reviewed, further comprising:

[0333] Generating a model internal thinking chain corresponding to the quality inspection report, and displaying the quality inspection report and the model internal thinking chain, wherein the model internal thinking chain is used to represent the review and analysis process of the quality inspection large model for generating the quality inspection report;

[0334] an artificial final review decision of the user according to the quality inspection report and the model internal thinking chain feedback;

[0335] final review of the quality inspection report based on the artificial final review decision, to obtain a final review report.

[0336] A8. The content review method of A7, after the final review of the quality inspection report based on the artificial final review decision, to obtain a final review report, further comprising:

[0337] obtaining decision data of the artificial final review decision, and updating a knowledge base according to the decision data;

[0338] constructing a training corpus based on the decision data, and training the quality inspection large model according to the training corpus.

[0339] A9. The content review method of any one of A1 to A6, before the inputting of the preliminary review report into the quality inspection large model, further comprising:

[0340] training a preset large model based on a preset data set through supervised fine-tuning to obtain a quality inspection basic large model;

[0341] adjusting internal parameters of the quality inspection basic large model through reinforcement learning training to obtain a quality inspection large model.

[0342] A10. The content review method of A9, the adjusting of the internal parameters of the quality inspection basic large model through reinforcement learning training to obtain a quality inspection large model, comprising:

[0343] converting historical decision data into a reward score, wherein the historical decision data is decision data before the current review;

[0344] adjusting internal parameters of the quality inspection basic large model based on the reward score through reinforcement learning training to obtain a quality inspection large model.

[0345] A11. The content review method of A9, before the training of the preset large model based on the preset data set through supervised fine-tuning to obtain a quality inspection basic large model, further comprising:

[0346] generating generation data corresponding to the content review task through a general large model;

[0347] obtaining feedback data from historical decision data, wherein the historical decision data is decision data before the current review;

[0348] constructing a preset data set based on at least one of positive example data, negative example data, the generation data, and the feedback data.

[0349] A12. The content review method of any one of A1-A6, wherein the semantic review of the preliminary review report by the quality review large model comprises:

[0350] contextual understanding and fact checking of the suspected errors in the preliminary review report by the quality review large model to obtain a false positive identification result;

[0351] deep reading of the preliminary review report by the quality review large model to obtain a missing supplement review result;

[0352] generating the quality review report of the content to be reviewed according to the false positive identification result and / or the missing supplement review result.

[0353] The present application also discloses B13, a content review device, comprising:

[0354] a preliminary review module configured to, in response to input of content to be reviewed, perform preliminary review of the content to be reviewed by a preset proofreading engine to obtain a preliminary review report;

[0355] a model calling module configured to input the preliminary review report into a quality review large model;

[0356] a model quality review module configured to perform semantic review of the preliminary review report by the quality review large model to obtain a review report of the content to be reviewed.

[0357] B14. The content review device of B13, further comprising:

[0358] a prompt word construction module configured to find relevant knowledge items related to the content to be reviewed in a knowledge base, wherein the knowledge base is pre-constructed based on a historical correction word library; call a deep research intelligent agent to perform fact checking on the content to be reviewed to obtain a fact checking report; and construct a prompt word according to at least one of the relevant knowledge items, the fact checking report, and the preliminary review report;

[0359] Correspondingly, the model quality review module is further configured to perform semantic review of the preliminary review report by the quality review large model based on the prompt word to obtain a review report of the content to be reviewed.

[0360] B15. The content review apparatus of B14, wherein the cue word construction module is further configured to invoke a deep research agent to plan a multi-step research plan for the content to be reviewed, wherein the multi-step research plan comprises a plurality of research steps having a logical progression relationship; and invoke a tool library of the deep research agent to conduct fact-checking on the content to be reviewed based on the multi-step research plan, and obtain a fact-checking report, wherein the preset tool library comprises a real-time web search engine tool, a retrieval enhancement generation tool, and a code interpreter.

[0361] B16. The content review apparatus of B15, wherein the cue word construction module is further configured to conduct progressive research on the content to be reviewed by the tool library of the deep research agent based on the multi-step research plan, and obtain a research result of each research step; and if the research result is a research failure, adjust the multi-step research plan by the deep research agent, obtain an adjusted multi-step research plan, and return to the step of conducting progressive research on the content to be reviewed by the tool library of the deep research agent based on the multi-step research plan, and obtain a research result of each research step.

[0362] B17. The content review apparatus of B14, wherein the cue word construction module is further configured to detect whether there is a factual statement in the content to be reviewed, and detect whether there is time-sensitive information in the content to be reviewed; and if there is a factual statement in the content to be reviewed, or there is time-sensitive information in the content to be reviewed, or the problem cannot be answered by retrieval enhancement generation, invoke a deep research agent to conduct fact-checking on the content to be reviewed, and obtain a fact-checking report.

[0363] The application further discloses C18, a content review device, comprising a memory, a processor, and a content review program stored in the memory and executable on the processor, wherein the content review program is executed by the processor to implement the content review method.

[0364] The application further discloses D19, a storage medium, wherein the storage medium stores a content review program, and the content review program is executed by a processor to implement the content review method.

[0365] The application further discloses E20, a computer program product, comprising a content review program, and the content review program is executed by a processor to implement the content review method.

Claims

1. A content moderation method, characterized in that, The content review methods include: In response to the input content to be reviewed, a preliminary review is performed on the content to be reviewed through a preset proofreading engine to obtain a preliminary review report; Input the preliminary review report into the quality inspection model; The preliminary review report is semantically inspected using the aforementioned quality inspection model to obtain a quality inspection report for the content to be reviewed.

2. The content moderation method as described in claim 1, characterized in that, Before performing semantic quality inspection on the preliminary review report using the quality inspection big data model to obtain the review report for the content to be reviewed, the process further includes: Search for relevant knowledge entries related to the content to be reviewed in the knowledge base, wherein the knowledge base is pre-built based on a historical error correction thesaurus; The deep research agent is invoked to perform fact-checking on the content to be reviewed, and a fact-checking report is obtained; Construct prompt words based on at least one of the relevant knowledge entries, the fact-checking report, and the preliminary review report; Accordingly, the step of performing semantic quality checks on the preliminary review report using the quality inspection model to obtain the review report for the content to be reviewed includes: Based on the prompt words, the preliminary review report is semantically inspected using the quality inspection model to obtain the review report for the content to be reviewed.

3. The content moderation method as described in claim 2, characterized in that, The process of invoking a deep research agent to perform fact-checking on the content to be reviewed and obtaining a fact-checking report includes: The deep research agent is invoked to plan the research steps for the content to be reviewed, resulting in a multi-step research plan, which includes multiple research steps with a logically progressive relationship. Based on the multi-step research plan, the tool library of the deep research agent is invoked to perform fact-checking on the content to be reviewed and obtain a fact-checking report. The preset tool library includes a real-time web search engine tool, a search enhancement generation tool, and a code interpreter.

4. The content review method as described in claim 3, characterized in that, The process involves calling the tool library of the deep research agent based on the multi-step research plan to perform fact-checking on the content to be reviewed, and obtaining a fact-checking report, including: Based on the tool library of the deep research agent in the multi-step research plan, the content to be reviewed is studied progressively to obtain the research results of each research step; If the research result is a research failure, the deep research agent will reflect on and adjust the multi-step research plan to obtain an adjusted multi-step research plan. Then, the tool library of the deep research agent based on the multi-step research plan will be returned to conduct progressive research on the content to be reviewed, and the research results of each research step will be obtained.

5. The content moderation method as described in claim 2, characterized in that, Before invoking the deep research agent to perform fact-checking on the content to be reviewed and obtaining a fact-checking report, the process also includes: The system detects whether the content to be reviewed contains factual statements and whether it contains time-sensitive information. Accordingly, the invocation of the deep research agent to perform fact-checking on the content to be reviewed and obtain a fact-checking report includes: If the content to be reviewed contains factual statements, outdated information, or questions that cannot be answered by retrieval enhancement, then a deep research agent is invoked to perform fact-checking on the content to be reviewed and obtain a fact-checking report.

6. The content moderation method as described in claim 2, characterized in that, The step of searching for relevant knowledge entries related to the content to be reviewed in the knowledge base, wherein the knowledge base is pre-built based on a historical error correction lexicon and includes: The content to be reviewed is subjected to entity recognition and keyword extraction to obtain the target entity and target keywords; Based on the target entity, the target keywords, and the preliminary review report, a similarity query is performed in the knowledge base to obtain relevant knowledge entries related to the content to be reviewed.

7. A content moderation device, characterized in that, The content moderation device includes: The preliminary review module is used to respond to the input content to be reviewed, and to perform a preliminary review of the content to be reviewed through a preset proofreading engine to obtain a preliminary review report; The model invocation module is used to input the preliminary review report into the quality inspection big model; The model quality inspection module is used to perform semantic quality inspection on the preliminary review report using the quality inspection big model, and obtain the review report of the content to be reviewed.

8. A content moderation device, characterized in that, The content moderation device includes: a memory, a processor, and a content moderation program stored in the memory and executable on the processor, wherein the content moderation program, when executed by the processor, implements the content moderation method as described in any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium stores a content moderation program, which, when executed by a processor, implements the content moderation method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a content moderation program, which, when executed by a processor, implements the content moderation method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Content auditing model training method and device, equipment and storage medium

    CN116502003A

  • Document auditing method, device and equipment and storage medium

    CN116663525A

  • Question answering system construction method based on large language model and question answering system

    CN118964587A

  • Large language model intelligent contract code auditing method and system based on retrieval enhancement generation and back one-step cue word

    CN119416217A

  • Answer auditing method and related equipment

    CN119988586A

Cited By

  • Enabling decision assistance method, device and equipment for data security control

    CN121256830A

  • Data value evaluation method and system based on knowledge mining large model and analogue simulation agent

    CN121526720A

  • Content generation method and device based on man-machine mixed feedback, equipment and medium

    CN121638305A