A content review quality monitoring system and method based on real-time dialing
By utilizing the task scheduling, test case switching, parameter supplementation, and feedback collection modules of the real-time monitoring system, the system addresses the sluggishness of traditional systems in scenario adaptation, enabling efficient monitoring of complex content and sensitive scenarios, and improving the coverage accuracy and adaptability of the review quality.
Patent Information
- Application Number
- CN202511487588.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-10-17
AI Technical Summary
In existing technologies, traditional real-time content review quality monitoring systems are slow to adapt to different scenarios and task nodes, and have difficulty responding in real time to content category expansion or node changes. This results in limited coverage of review quality monitoring under complex business needs and dynamic content distribution conditions, making it impossible to collect high-risk or new content in a timely manner, and lacking sufficient archiving and retrospective capabilities.
A content review quality monitoring system based on real-time probing is adopted, including a task scheduling module, a test case switching module, a parameter supplementation module, and a feedback collection module. By analyzing the consistency between scenario status tags and sensitive period tags, review tasks are dynamically allocated, enabling multi-scenario switching and flexible collaboration of task nodes. Feedback data covering blind spots and high-risk links is continuously collected to optimize the distribution of content category tasks and improve the traceability and scenario adaptability of the review process.
It enables linked monitoring of complex content types and sensitive scenarios, improves the coverage accuracy and scenario adaptability of the entire content review process, supports differentiated test case management for different periods, dynamically supplements sample categories, and enhances the monitoring capabilities for abnormal distribution and node changes.
Smart Images

Figure CN120973689B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of Internet content security, and in particular to a content review quality monitoring system and method based on real-time dialing. BACKGROUND
[0002] The field of Internet content security involves detecting, identifying, analyzing and managing information content spread on the network to protect the health and legality of the network environment. The traditional real-time dialing content review quality monitoring system refers to continuously submitting sample data of known categories to the content review system at different time nodes through a pre-set test account or identity, and collecting and counting review behaviors according to the sample review results to monitor the review process execution and review standard consistency. It usually uses periodic sample submission, records review response and results, compares sample expected labels with actual review labels, and other means to complete the tracking and quality collection of the review quality in the content review link.
[0003] The prior art uses fixed use cases and periodic sample detection, lacks real-time adaptation to different scene states and task nodes, responds slowly when facing content category expansion or node changes, has gaps in sample coverage in sensitive periods and variable scenarios, and it is difficult to effectively track the execution of review nodes under multi-dimensional content categories, which may cause some high-risk or new content to be missed, and the archiving and backtracking capabilities are insufficient, resulting in limited coverage and delayed response of the review quality monitoring under complex business requirements and dynamic content distribution conditions. SUMMARY
[0004] The purpose of the present application is to solve the shortcomings in the prior art and to provide a content review quality monitoring system and method based on real-time dialing.
[0005] In order to achieve the above purpose, the present application adopts the following technical scheme: a content review quality monitoring system based on real-time dialing, comprising a task scheduling module, a use case switching module, a parameter supplement module and a feedback collection module;
[0006] The task scheduling module analyzes the consistency of the scene state label and the sensitive period label based on the dialing start time parameter, judges the correspondence between the scene label and the scheduling node number, filters the use case library number, determines the association between the use case item set and the node, and obtains the dialing distribution sequence parameter;
[0007] The use case switching module judges the correspondence between the use case number and the service type based on the dialing distribution sequence parameter, analyzes the matching of the application identifier and the channel identifier, compares the content text and the expected review result, identifies the matching of the use case item and the node, and obtains the use case deviation comparison details;
[0008] The parameter supplement module analyzes inconsistent text and audit identity coverage based on the use case bias comparison details, judges content category task distribution frequency, screens risk level occurrence frequency, compares submission time task distribution, optimizes short video violation identification task supplement, and obtains coverage gap distribution characteristics.
[0009] The feedback collection module judges sample supplement task allocation based on the coverage gap distribution characteristics, analyzes node number and audit result association, compares interface return state and original task data, screens matched and newly added sample category items, and obtains sample feedback structure fluctuation.
[0010] The application improves that the test distribution sequence parameters include task distribution identifier, execution priority, and sequence number, the use case bias comparison details include difference type identifier, influence range mark, and difference source mark, the coverage gap distribution characteristics include gap category identifier, gap distribution proportion, and gap influence level, and the sample feedback structure fluctuation includes structure change category, change node distribution, and change influence degree.
[0011] The application improves that the task scheduling module includes a time label matching sub-module, a node number screening sub-module, and a sequence parameter generation sub-module.
[0012] The time label matching sub-module judges the overlap state of the time parameter and the sensitive period label based on the test start time parameter, combines the set sensitive period label and the current time characteristics, analyzes the applicable classification of the business scene label under this time node, determines the mapping relationship between the time and the scene state, obtains the time scene corresponding type, and determines the mapping relationship between the time and the scene state.
[0013] The node number screening sub-module screens the scene labels involved based on the time scene corresponding type, compares the business scene categories corresponding to each scheduling node number, judges the consistency of the node number and the scene label, optimizes the matching situation of the node number set and the use case library number, identifies the node number set of the same scene type, and obtains the node and use case matching range.
[0014] The sequence parameter generation sub-module adjusts the arrangement order of each use case item according to the node and use case matching range, analyzes the distribution of the scheduling node number and the use case item combination, optimizes the relationship between the task distribution identifier and the execution priority, integrates the scheduling task parameters in turn, and obtains the test distribution sequence parameters.
[0015] The application improves that the use case switching module includes a service matching sub-module, a difference node identification sub-module, and a result adjustment sub-module.
[0016] The service matching submodule analyzes the correspondence between the use case number and the service type based on the dial test allocation sequence parameter, judges whether the type of the service associated with each use case is accurate, compares whether the application identifier and the channel identifier are consistent with the parameters associated with the scheduling node, identifies the matching situation of the content text and the review state, and obtains the service state matching degree;
[0017] The difference node identification submodule analyzes the review state returned by multiple scheduling nodes based on the service state matching degree, judges the similarities and differences between the feedback result of each node and the expected review state, filters the nodes with different review states, optimizes the feedback distribution between the node and the majority nodes, and obtains the feedback node deviation amplitude.
[0018] The result adjustment submodule compares the feedback node deviation amplitude with the correspondence between the use case number and the node number, analyzes the state distribution of the same use case on each node, calculates the concentration degree of the state distribution of the difference node, obtains the use case data difference breadth, adjusts the state distribution of each node, filters the abnormal fluctuation items, and obtains the use case deviation comparison details.
[0019] The parameter supplement module comprises a text coverage identification submodule, a task distribution filtering submodule and a content category optimization submodule.
[0020] The text coverage identification submodule judges the distribution of the text content and the review identity in all tasks based on the use case deviation comparison details, filters the text not covered by manual or machine review, compares the coverage difference under each identity, optimizes the identity allocation rule, and obtains the review coverage missing proportion.
[0021] The task distribution filtering submodule filters the distribution characteristics of the tasks under the content category and the time interval based on the review coverage missing proportion, judges the distribution difference between the high-frequency category task and the low-frequency category task, compares the proportion of the risk level in each content category, optimizes the category distribution, and obtains the risk frequency distribution structure.
[0022] The content category optimization submodule filters the coverage of the short video content category under the difference period and the distribution intensity based on the risk frequency distribution structure, compares the relationship between the distribution of the illegal category task and the content supplement demand, calculates the supplement priority of each category task, adjusts the content supplement node distribution, and obtains the coverage gap distribution characteristics.
[0023] The feedback collection module comprises a task allocation identification submodule, a state comparison determination submodule and a structure fluctuation calculation submodule.
[0024] The task allocation identification submodule analyzes the matching between the node information corresponding to the supplementary task and the task allocation record based on the coverage gap distribution characteristics, judges whether the task instruction and the node distribution are consistent, filters the task path associated with the structure, and gradually optimizes the task trajectory matching combined with the allocation sequence to obtain a sample node matching path set;
[0025] The state comparison and determination submodule compares whether the audit state of each task and the state of the original record are consistent based on the sample node matching path set, filters the task categories and corresponding nodes with differences, optimizes the grouping statistics of the state difference content, and obtains the audit state difference distribution;
[0026] The structure fluctuation calculation submodule analyzes the task structure involved in the audit state difference distribution, judges the change of the sample category and the node distribution, compares the structure combination difference between the sample category and the node, judges the structure change trend, and obtains the sample feedback structure fluctuation.
[0027] The system further comprises a report archiving module, which analyzes the test task data and the scheduling parameters based on the sample feedback structure fluctuation, optimizes the sample distribution labeling and the scene switching time matching, compares the differences between manual auditing and machine auditing, judges the comparison of the marked items, and obtains a multi-scene coverage distribution result.
[0028] The multi-scene coverage distribution result includes scene coverage range, coverage structure composition, and coverage distribution relationship.
[0029] The report archiving module comprises a feedback analysis submodule, a consistency discrimination submodule, and a coverage distribution construction submodule.
[0030] The feedback analysis submodule analyzes the feedback content in the test task data based on the sample feedback structure fluctuation, judges the change category of the task parameters and the use case parameters between nodes, compares the difference between the interface feedback data and the original audit result, filters the node allocation in the scheduling stage parameters, and obtains a structure difference identification result.
[0031] The consistency discrimination submodule compares the synchronization between the scene switching time and the sample distribution content labeling time based on the structure difference identification result, judges the consistency of the feedback of each node in the artificial auditing difference clustering scene, optimizes the matching relationship between the feedback marking content and the node result, and obtains a feedback result matching feature.
[0032] The coverage distribution construction submodule determines the coverage between node information and scene labels based on the feedback result matching features, compares the correlation between the distribution features of task nodes and the scene label configuration, identifies the node and label configuration in the task archiving parameters during the scheduling phase, optimizes the attribution correspondence between scenes and node tasks, and obtains multi-scene coverage distribution results.
[0033] A content moderation quality monitoring method based on real-time testing, executed based on the aforementioned content moderation quality monitoring system based on real-time testing, includes the following steps:
[0034] S1: Based on the test start time parameter, analyze the matching relationship between the time and the scene status label, determine the consistency between the current time and the preset sensitive period label, compare the correspondence between the scene label and the scheduling node number, filter the corresponding use case library number, determine the association between the use case item set and the scheduling node number, determine the scheduling task order, and obtain the test allocation sequence parameter.
[0035] S2: Based on the test allocation sequence parameters, determine the correspondence between the test case number and the service type, analyze the matching of the application identifier and the channel identifier, compare the content text with the expected review result, determine the difference between the difference node and the content returned by the feedback interface, identify the matching of the test case item and the scheduling node, and adjust the difference in the data performance of the test case item to obtain the test case deviation comparison details.
[0036] S3: Based on the deviation comparison details of the use cases, analyze the coverage of inconsistent text and review identity, determine the frequency of task distribution in the content category, filter the frequency of occurrence of risk level items, compare the distribution of tasks within the submission time range, optimize the content category supplementary tasks in the short video violation identification scenario, calculate the differences between task types, and obtain the coverage gap distribution characteristics.
[0037] S4: Based on the distribution characteristics of the coverage gap, determine the allocation of sample supplementation tasks, analyze the correlation between node number and audit result item, compare the interface return status with the original task data, filter the sample category items of matching tasks and new tasks, and obtain the sample feedback structure fluctuation.
[0038] S5: Based on the fluctuations in the sample feedback structure, analyze the test task data and scheduling stage parameters, optimize the matching of sample distribution content labeling and scene switching time, compare the feedback consistency under manual review difference clustering scenario with the differences in machine review items, and obtain multi-scenario coverage distribution results.
[0039] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0040] This invention automatically allocates review tasks based on different business scenarios and scheduling points, enabling multi-scenario switching of test cases and flexible collaboration of task nodes. It supports differentiated test case management for sensitive periods, non-sensitive periods, and temporary periods. By analyzing differences in node feedback and content category distribution, it dynamically supplements sample categories, continuously collects feedback data from coverage blind spots and high-risk links, constructs a multi-dimensional coverage structure based on task execution and scenario tags, and aggregates abnormal distributions and node changes in the review process. This improves the ability to monitor the linkage of complex content types, sensitive scenarios, and review node status, and promotes the traceability, coverage accuracy, and scenario adaptability of the entire content review process. Attached Figure Description
[0041] Figure 1 This is a system flowchart of the present invention;
[0042] Figure 2 This is a flowchart of the task scheduling module in this invention;
[0043] Figure 3 This is a flowchart of the use case switching module in this invention;
[0044] Figure 4 This is a flowchart of the parameter supplementation module in this invention;
[0045] Figure 5 This is a flowchart of the feedback acquisition module in this invention;
[0046] Figure 6 This is a flowchart of the report archiving module in this invention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0048] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0049] Example
[0050] Please see Figure 1The present invention provides a technical solution: a content review quality monitoring system based on real-time probing, including a task scheduling module, a test case switching module, a parameter supplementation module and a feedback collection module;
[0051] The task scheduling module analyzes the matching relationship between the test start time parameter and the scene status label, judges the consistency between the current time and the preset sensitive period label, compares the correspondence between the scene label and the scheduling node number, filters the corresponding use case library number, judges the association between the use case item set and the scheduling node number, determines the scheduling task order, and obtains the test allocation sequence parameter.
[0052] The test case switching module determines the correspondence between test case numbers and service types based on the test allocation sequence parameters, analyzes the matching of application identifiers and channel identifiers, compares the content text with the expected review results, determines the differences between the difference nodes and the content returned by the feedback interface, identifies the matching of test case entries with scheduling nodes, and adjusts the differences in the data performance of test case entries to obtain a detailed comparison of test case deviations.
[0053] The parameter supplementation module analyzes the coverage of inconsistent text and review identity based on the test case deviation comparison details, determines the frequency of task distribution in content categories, filters the frequency of occurrence of risk level items, compares the distribution of tasks within the submission time range, optimizes the content category supplementation task in the short video violation identification scenario, calculates the differences between task types, and obtains the distribution characteristics of coverage gaps.
[0054] Based on the distribution characteristics of the coverage gap, the feedback collection module determines the allocation of sample supplementation tasks, analyzes the correlation between node numbers and audit result items, compares the interface return status with the original task data, filters sample category items of matching tasks and newly added tasks, optimizes the structural differences between tasks, and obtains sample feedback structure fluctuations.
[0055] The report archiving module analyzes the fluctuations in sample feedback structure, examines test task data and scheduling parameters, optimizes the matching of sample distribution content labeling and scene switching time, compares the consistency of feedback under manual review difference clustering scenarios with the differences in machine review items, judges the comparison of labeled items, and obtains multi-scenario coverage distribution results.
[0056] The test allocation sequence parameters include task allocation identifier, execution priority, and sequence number; the test case deviation comparison details include difference type identifier, impact range marker, and difference source marker; the coverage gap distribution characteristics include gap category identifier, gap distribution ratio, and gap impact level; the sample feedback structure fluctuation includes structure change category, change node distribution, and change impact degree; and the multi-scenario coverage distribution results include scenario coverage range, coverage structure composition, and coverage distribution relationship.
[0057] The system supports three types of test case libraries:
[0058] Sensitive Period Use Case Library: Content review standards for special periods (such as anniversaries and sensitive dates);
[0059] Non-sensitive period use case library: Content review standards for regular periods;
[0060] Temporary test case library: Used to address temporary testing needs in response to unexpected events.
[0061] Each test case in the test case library contains the following attributes:
[0062] server: Service type (web or media, corresponding to web page review API and converged media API respectively);
[0063] appId: Application identifier;
[0064] channel: Channel identifier;
[0065] query: Test content text;
[0066] expected_result: Expected audit result (PASS or REJECT);
[0067] The system performs the dial-up test through the following steps:
[0068] (1) Determine the sensitive / non-sensitive period based on the current time and load the corresponding test case library;
[0069] (2) For each test case, call the corresponding content moderation API according to the service type;
[0070] (3) Compare the audit results returned by the API with the expected results of the use cases;
[0071] (4) Record test results, including execution time, API return value, matching status, etc.;
[0072] (5) Generate and archive the test report;
[0073] (6) When a mismatch is found, an alarm notification is triggered.
[0074] Employs a flexible scheduling strategy based on Cron expressions:
[0075] (1) Supports setting the execution frequency via a web interface or configuration file;
[0076] (2) Supports start, pause, and resume scheduling operations;
[0077] (3) Supports single manual execution;
[0078] (4) Automatic retry when scheduling fails.
[0079] When the test results do not match expectations, the system triggers an alarm based on the following rules:
[0080] (1) Set different alarm levels according to the failure rate: all failures are critical alarms, failure rate ≥ 50% is high alarm, and failure rate < 50% is medium alarm;
[0081] (2) Alarm messages are pushed through enterprise instant messaging tools (such as WeChat for Business);
[0082] (3) The alarm message includes details of the failed test case, the test time, and the expected results;
[0083] (4) Supports setting alarm silence period to avoid repeated alarms in a short period of time.
[0084] To improve execution efficiency, the system supports concurrent execution mode, and the maximum number of concurrent users and the timeout time for a single test case can be configured.
[0085] In addition, the system supports the generation and management of reports in multiple formats:
[0086] (1) JSON format report: facilitates program parsing and secondary processing;
[0087] (2) HTML format report: easy for manual viewing and analysis;
[0088] (3) Supports historical report archiving and querying;
[0089] (4) Supports report export and sharing.
[0090] In the task scheduling module, the test start time parameter refers to the specific time (e.g., year, month, day, hour, minute) when the system automatically initiates a content review test (test) task, used to determine which type of test period (sensitive period / non-sensitive period / temporary); the scenario status label refers to the preset scenario category label, such as "daily scenario," "major event scenario," and "temporary event scenario," used to distinguish the current review business background; the preset sensitive period label refers to the predefined special period markers (e.g., "anniversary," "holiday," "public emergency"), which affect test case selection and scheduling logic; the scenario label refers to the identifier that specifically reflects the review business scenario, such as "short video review," "image and text content review," and "community message review," etc.; the scheduling section The node number refers to the unique number of each content review interface test node (e.g., Node01, Node02), indicating which execution unit the task will be assigned to; the test case library number refers to the unique identifier of each test case library, such as the sensitive period test case library, the non-sensitive period test case library, and the temporary test case library (corresponding to the test case library management in patent 2.1); the test case item set refers to the set of all individual test cases contained under a specific test case library number (each test case has independent parameters); the scheduling task order refers to the specific order in which the test tasks to be executed are sorted according to the scheduling strategy and number, ensuring efficient scheduling and allocation of the system; the test allocation sequence parameter refers to the number string that uniquely identifies all task scheduling and allocation in this session after sorting, used for subsequent task tracking and archiving.
[0091] In the test case switching module, service type refers to the type of review interface targeted by each test case, such as Web content review API or converged media API; application identifier and channel identifier matching refers to analyzing whether the "appId" (application identifier) and "channel" (channel identifier) in the test case match the actual application and channel of the target review service; difference node refers to the difference in feedback results between nodes when multiple nodes (multiple content review API instances) are tested in parallel, and the node that produces the difference is the "difference node"; difference in returned content refers to the difference between the actual API return result and the "expected_result" (expected review result) in the test case, including whether the review result is passed or rejected; matching status refers to whether the API return content meets expectations (whether it is completely consistent) during the actual execution of the test case item and the scheduling node. A match indicates consistency, and a mismatch requires further analysis; difference in data performance refers to the difference in the numerical and status performance of the feedback review results after the test data is executed on different nodes, different test cases, or different channels (e.g., one node judges it as PASS, and another judges it as REJECT).
[0092] In the parameter supplementation module, inconsistent text and review identity coverage refer to analyzing content text that is inconsistent with the API review results and test case expectations, and statistically analyzing the coverage of review identities (such as manual review and machine review) to determine if there are any omissions or blind spots; task distribution frequency refers to statistically analyzing the occurrence frequency of various content review tasks (such as different content categories and different review identities) to detect distribution balance or identify deviations; the occurrence frequency of risk level items refers to analyzing the distribution frequency of test cases of different risk levels (such as "high-risk content" and "general content") in probing tasks, and supplementing them in a timely manner when uneven distribution is found; content category supplementation tasks refer to generating supplementary test tasks for content categories with insufficient sample coverage found in the test to ensure that all categories are monitored; differences between task types refer to analyzing the differences in distribution, performance, and results of different task types (such as text and image review, short video review, community comments, etc.) after test execution.
[0093] In the feedback collection module, the allocation of supplementary test tasks refers to the specific review nodes (i.e., execution units) to which the generated supplementary test tasks have been assigned, and the allocation results; the interface return status refers to the status code, status flag, or review result field returned by the review API during test execution, directly reflecting the API's running status and review conclusion; the original task data refers to all original test information recorded by the system at the start of the test task, such as test case parameters, execution nodes, and expected results; the sample category entries refer to samples refined to specific content categories (such as pornography, violence, advertising, normal, etc.), used for feedback collection statistics and comparison; the structural differences between tasks refer to the structural differences between the original test tasks and supplementary tasks in terms of content category, risk level, execution node, review conclusion, etc.
[0094] In the report archiving module, the test task data refers to all raw and processed data generated by the test system during task execution, including task parameters, test case parameters, review results, node information, etc.; scheduling phase parameters refer to the parameters involved in each stage of the entire process from scheduling to allocation, execution, and feedback, such as scheduling time, node allocation information, task archiving number, etc.; matching status refers to the consistency status description between the actual feedback result and the expected result for each task or test case, used for final report archiving and problem localization; manual review difference clustering scenario refers to the statistical analysis of differences between manual review conclusions and machine judgments for review tasks with discrepancies, which facilitates tracking consistency and abnormal distribution; the comparison status of marked items refers to the one-to-one comparison of the marked content in the actual test feedback (such as pass, rejection, abnormal, etc.) with the expected marked items, statistically analyzing and archiving all differences and consistency items to support multi-scenario report analysis.
[0095] Please see Figure 2 The task scheduling module includes a time tag matching submodule, a node number filtering submodule, and a sequence parameter generation submodule;
[0096] The time tag matching submodule, based on the test start time parameter, combined with the pre-set sensitive period tags and current time characteristics, determines the overlap between the time parameter and the sensitive period tag, analyzes the applicable classification of the business scenario tag at this time node, determines the mapping relationship between time and scenario status, and obtains the corresponding type of time scenario.
[0097] For example, if the current test time is "March 5, 2025, 14:30", the time is first parsed into five items in the standard structure: year, month, day, hour, and minute. Then, each record is read from the pre-defined sensitive period label table. Each record in the table consists of three parts: label name, start time, and end time. For example, the label "Anniversary" records from March 3, 00:00 to March 10, 23:59. The time range of each label is compared field by field. The test start time is compared with the start and end times of each label to determine if it is greater than or equal to, or less than or equal to. If the time range of a label meets the condition, the current time is considered to be within the sensitive period. In this example, the result is that it is within the "Anniversary" sensitive period. The system first assigns a "Time Period Tag" period, then requests a set of scenario categories from the mapping configuration table based on the current time stamp. This table has a one-to-one correspondence between "Time Period Tag" and "Scenario Tag List". Scenario tags matching the current task scheduling priority are selected as the preferred match according to preset priority rules. If multiple tag candidates exist, they are sorted according to the task weight value of the tag. The task weight value of the tag is manually set by the administrator. For example, if the weight of "Image and Text Hotspot Review" is 100 and the weight of "Sensitive Comment Review" is 80, the former is selected as the final business scenario. This scenario tag is then combined with the current testing time to generate a time scenario type in the form of "Sensitive Period - Image and Text Hotspot Review", which serves as the core basis for matching subsequent scheduling nodes with the test case library.
[0098] The node number filtering submodule filters the scenario tags involved based on the time scenario type, compares the business scenario categories corresponding to each scheduling node number, judges the consistency between the node number and the scenario tag, optimizes the matching between the node number set and the test case library number, identifies the node number set with the same scenario type, and obtains the matching range between nodes and test cases.
[0099] Upon receiving the time scenario type "Sensitive Period - Hot Topic Review," the system extracts the set of all registered scheduling node numbers from the node information table. For example, there are 10 registered nodes, numbered Node01 to Node10. Each node is configured with business capability fields. For instance, Node01 is bound to the tags "Hot Topic Review, Timeliness Inspection," and Node02 is bound to "Short Video Recognition, Ad Detection." The current target scenario tag "Hot Topic Review" is broken down into two keywords: "Image and Text" and "Hot Topic." A precise keyword matching operation is performed on the tag fields of each node. The matching rule is that both keywords must appear in the node tag to be considered a match. For example, Node01's capability tag field contains both "Image and Text" and "Hot Topic," so it is considered a complete match. Node03, which only contains the "Hot Topic" field, is discarded. The final set of matching node numbers is Node01, N... Node04 and Node06 are then used. The test case library number configuration table is then called to find the test case library number corresponding to the current scenario tag, such as "CaseSet_A_001". Fields such as "server type" and "content tag" are read and compared again with the node capability fields. For example, if the test case library server field is "Web interface", Node04 supports both Web and converged media interfaces, Node06 only supports Web, and Node01 only supports media interfaces. Therefore, only Node04 and Node06 meet the interface requirements. After this round of comparison, the node number set is further converged to Node04 and Node06. Finally, the node number set "Node04, Node06" is marked as the scheduling candidate node number set for the current task, and the binding mapping between the current scenario type "Sensitive Period - Image and Text Hotspot Review" and the test case library "CaseSet_A_001" is recorded.
[0100] The sequence parameter generation submodule adjusts the order of each test case entry based on the matching range of nodes and test cases, analyzes the distribution of scheduling node numbers and test case entry combinations, optimizes the relationship between task allocation identifiers and execution priorities, and integrates scheduling task parameters in sequence to obtain the test allocation sequence parameters.
[0101] After obtaining the matching range of nodes and test cases, for example, if the current node set is Node04 and Node06, and the test case entries are 5 numbered Case001 to Case005, each test case is combined with two nodes to form a total of 10 test tasks. First, the review service type, risk level, and expected review result fields of each test case are read. For example, Case001 is set to "Reject" with an expected result of "High" and a risk level of "High", and the content category is "Sensitive Comments". Case002 is set to "Pass" with a risk level of "Low". According to the preset priority strategy, tasks with a risk level of "High" and an expected rejection are marked as priority 1, and others are marked as priority 2. Then, each task item is assigned an execution order number. Under the condition of the same priority, the nodes are sorted in ascending order of their node numbers. For example, Node04 executes Case001 task number Seq001, Node06 executes Case001 task number Seq002, and so on, ultimately forming 10 scheduling entries. Each entry consists of "task identifier, node number, test case number, priority, and sequence number". For example, the content of a task entry is: "Task_Node04_Case001, priority 1, sequence number 001". The scheduling entries are integrated to form a task allocation parameter structure. The internal field structure is uniformly "task ID, execution node, task sequence number, priority label, and corresponding test case number", and it is stored in the allocation buffer pool as the sole basis for subsequent task distribution.
[0102] Please see Figure 3 The use case switching module includes a service matching submodule, a difference node identification submodule, and a result adjustment submodule;
[0103] The service matching submodule analyzes the correspondence between test case numbers and service types based on the test allocation sequence parameters, determines whether the type of service associated with each test case is accurate, compares whether the application identifier and channel identifier are consistent with the parameters associated with the scheduling node, identifies the matching status of content text and review status, and obtains the service status matching degree.
[0104] Read the task information structure in the allocation sequence one by one, and extract the test case number, service type field, application identifier field, channel identifier field, and scheduling node number field. Based on the first step's analysis of the correspondence between the test case number and service type, query the service type field stored in the test case library to see if it matches the service call interface type field recorded in the scheduling task. For example, if the service type corresponding to test case number Case102 is "Web content review interface", and the service call type recorded in the scheduling node is "converged media review interface", then it is identified as a service type mismatch, and such tasks are marked as erroneous tasks. If the test case number matches the service field, proceed to the next verification step, extract the application identifier and channel identifier values from the task item, for example, "appId is App_001, channel is Channel_01". Then, retrieve the service adaptation parameter field recorded in the scheduling node registration information, find the application scope and channel list supported by the node, and perform a string value exact match operation. For example, Node01 supports App_001 and Channel_001. If the result is 01, it is considered a complete match; otherwise, it is marked as an inconsistent task. The audit result field returned by the scheduling node is compared with the expected audit status field of the test case in the test case library. For example, if the node returns an audit result of "REJECT" after the task is executed, but the expected result of the test case is "PASS", it is recorded as an inconsistent audit status item. All comparison operations are recorded in the feedback log table and a comparison flag field is generated. Matching items are marked as "matched", and non-matching items are marked as "non-matched". Then, according to the three types of matching field status of each task item, its matching value is calculated. For example, if a task has two matching items in the service type, application identifier, and audit status, the service matching degree is "2 / 3" and the matching rate is 66.67%. The matching values of all task items are collected and classified according to the following standards: when the matching rate is greater than or equal to 80%, it is marked as "high matching", between 50% and 80% is "medium matching", and below 50% is "low matching". A list of service status matching degrees of all test task items is generated as the basis for subsequent feedback node identification.
[0105] The difference node identification submodule analyzes the review status returned by multiple scheduling nodes based on the service status matching degree, judges the similarity and difference between the feedback result of each node and the expected review status, filters the nodes with different review statuses, optimizes the feedback distribution between the node and the majority of nodes, and obtains the deviation of the feedback node.
[0106] First, group all execution tasks by task number and extract the review status return values of all scheduling nodes involved in each task number. For task number T001, if it involves three nodes, Node02, Node05, and Node06, and their review statuses are "PASS", "REJECT", and "PASS" respectively, then collect this group of statuses into a status sequence and count their frequencies. In this example, "PASS" appears twice and "REJECT" appears once. First, determine that the main feedback status of this task number is "PASS", which has the highest frequency, and use it as the reference status for the current task number. Then, compare the return status of each node with the reference status. If they are consistent, mark it as "main status consistent"; otherwise, mark it as "deviation node". In this example, Node05 is a deviation node. Continue to perform the same operation on all task numbers and count how many times each node is marked as "deviation node" in all tasks. The total number of times a node is marked as "off-node" is calculated, followed by the deviation ratio of each node. The calculation method is to divide the number of tasks marked as off-nodes by the total number of tasks executed. For example, if Node05 executes 30 tasks, and 10 of them are inconsistent with the main state, the deviation ratio is 33.33%. A preset threshold is used to determine whether a node needs further attention. The threshold is set to 20%. That is, if the deviation ratio of a node exceeds 20%, it is considered an unstable feedback state node. If the ratio is between 10% and 20%, it is considered a slight deviation. If it is below 10%, it is marked as a stable feedback node, and so on. After classifying all nodes by deviation level, the deviation ratio results are output in a list. For example, the deviation ratio of Node01 is 5.4%, Node04 is 21.7%, and Node05 is 33.3%. The task feedback logs of all off-nodes are classified and archived as input for subsequent feedback fluctuation analysis to obtain the deviation magnitude of the feedback nodes.
[0107] The result adjustment submodule compares the deviation magnitude of feedback nodes with the correspondence between use case numbers and node numbers, analyzes the state distribution of the same use case on each node, and calculates the concentration of the state distribution of the differing nodes using the formula:
[0108] ;
[0109] Breadth of differences in test case data Adjust the state distribution of each node, filter out abnormal fluctuation items, and obtain a detailed comparison of test case deviations. Indicates the first The audit status number returned by the scheduling node. Indicates the first The expected state number set in the use case. This indicates the number of use case entries included in the comparison. Indicates the first The difference between the feedback state of each node and the average state of the nodes in the same group. Indicates the first Fluctuations in the state distribution of individual differing nodes This indicates the number of different nodes.
[0110] The breadth of test case data variation reflects the degree of difference between the actual and expected review status of test cases across multiple scheduling nodes. It measures the distribution difference of review results between different nodes, i.e., the degree of deviation of the review conclusions of each node for the same test case. It provides an indicator to measure the consistency and degree of variation of review status. The larger the value, the greater the difference in the review results of the test cases across nodes, and the worse the review consistency and stability of the system.
[0111] Based on the deviation magnitude of feedback nodes, the correspondence between use case numbers and node numbers is compared. The state distribution of the same use case across different nodes is analyzed one by one, especially when the state differences between feedback nodes are significant. The focus is on calculating the concentration of state distribution among the differing nodes. A weighted average of the state differences for each node is applied using a formula. First, the difference between the actual feedback state and the expected state of each node is calculated, and the absolute values of the differences for each node are accumulated and averaged as the first part of the error. Second, the difference between the feedback state of each differing node and the average state of the nodes in the same group, along with the degree of fluctuation of the difference, is calculated, and a weighted sum is used to obtain the second part. Finally, the breadth of use case data differences is obtained by summing and weighting the two parts. This value is used for subsequent adjustments to the state distribution between nodes, filtering out abnormal fluctuations and optimizing them. In a certain test case, the following is the situation of node feedback and expected results (normalized):
[0112] Node 1: Actual Feedback Status (Original state "REJECT"), Expected state (Original expectation: "PASS");
[0113] Node 2: Actual Feedback Status (Original state "PASS"), Expected state (Original expectation: "PASS")
[0114] Node 3: Actual Feedback Status (Original state "PASS"), Expected state (Original expectation: "PASS").
[0115] To calculate the breadth of difference in use case data First, calculate the difference for each node:
[0116] For node 1, the difference is ;
[0117] For node 2, the difference is: ;
[0118] For node 3, the difference is: ;
[0119] Then, average the absolute values of the differences:
[0120] ;
[0121] Calculate the fluctuation of each dissimilar node, if the fluctuation of each node differs from the average state of other nodes as follows:
[0122] Node 1: State fluctuation Fluctuation range ;
[0123] Node 2: State Fluctuation Fluctuation range ;
[0124] Node 3: State Fluctuation Fluctuation range ;
[0125] Sum the fluctuations and amplitudes with weights, and then take the average:
[0126] ;
[0127] Calculate the breadth of difference in use case data :
[0128] ;
[0129] use case data diversity breadth This represents the degree of difference between the feedback state and the expected state of the use case at different nodes. This value can be used to adjust the node state distribution in the future, thereby optimizing the execution order and accuracy of the test task.
[0130] Set the following interval range:
[0131] Interval 1 (low difference interval): Between 0 and 0.2, when Within this range, it indicates that the state differences between the scheduling nodes are very small, the audit results are relatively consistent, the system audit accuracy is high, and the consistency of state feedback is good.
[0132] Interval 2 (Interval of Difference): Between 0.2 and 0.4, when This range indicates that there are certain differences between the scheduling nodes, and the system needs further optimization to reduce the deviation between nodes and ensure higher consistency.
[0133] Interval 3 (High Difference Interval): Above 0.4, when When the value exceeds 0.4, it indicates that there are significant differences in the status between the scheduling nodes, resulting in poor consistency in the review process. The system needs to be significantly adjusted to optimize node configuration, reduce feedback discrepancies, and improve review stability.
[0134] This indicates that the data falls within interval 2 (medium difference interval), meaning that there are some differences in the feedback states among multiple scheduling nodes, but the "high difference" standard has not yet been reached. Optimization and adjustment should be considered: adjust the configuration of scheduling nodes to reduce the number of nodes with large feedback differences; redistribute review tasks to ensure a more even distribution of node states in the review process; and optimize the review rules for nodes with large feedback differences to reduce deviations.
[0135] Please see Figure 4 The parameter supplementation module includes a text coverage recognition submodule, a task distribution filtering submodule, and a content category optimization submodule;
[0136] The text coverage recognition submodule determines the distribution of text content and review identity across all tasks based on the use case deviation comparison details, filters out texts that have not been covered by human or machine review, compares the coverage differences under each identity, optimizes the identity allocation rules, and obtains the review coverage missing ratio.
[0137] First, iterate through each deviation detail record, extracting the test case number, corresponding text content field, expected review result, actual return result, and review identity type field. The review identity field indicates whether the task was reviewed manually or by machine. Count the number of tasks completed manually and the number completed by machine across all test cases, and construct an identity distribution mapping table based on this. For example, if there are 1000 tasks in total, with 340 manually reviewed and 660 machine reviewed, the initial identity distribution is 34% manual and 66% machine. Then, identify task items in the deviation comparison details whose review status is inconsistent with expectations, extract their corresponding text content field, and search the historical task records to see if this text content has ever been reviewed manually or by machine. The criterion is whether there are any records of manual or machine review in the last five tasks with the same text content. If none are found, the text content is considered unrecorded. If a text is not covered, all text content not covered by human or machine review is recorded in a text blind spot list. This list is then categorized by role type. For example, if 70 out of 100 uncovered texts were not reviewed by a human reviewer and 30 by a machine reviewer, the human coverage gap rate is determined to be 70%, and the machine coverage gap rate is 30%. The difference between the two is then compared, and optimization actions are performed based on the minimum coverage benchmark value set in the role rules. For example, if the role benchmark coverage ratio is set to 40%, and the human review coverage rate is lower than this benchmark, it is determined that the human role needs to be adjusted and more high-risk category tasks should be assigned. The coverage statistics for the two review roles, the number of blind spot texts, the actual difference, and other parameters are output. The overall text coverage gap ratio during the review process is also calculated. For example, if 90 out of 1000 tasks are not reviewed by any role in all tasks, the review coverage gap ratio is calculated to be 9%, and the review coverage gap ratio is output.
[0138] The task distribution filtering submodule filters the distribution characteristics of tasks under different content categories and time intervals based on the proportion of missing review coverage. It judges the distribution differences between high-frequency and low-frequency tasks, compares the proportion of risk levels in each content category, optimizes the category distribution, and obtains the risk frequency distribution structure.
[0139] All tasks are filtered and categorized based on time and content category. First, task data is grouped according to the content category field, which includes, but is not limited to, types such as "violent," "pornographic," "advertising," and "normal." For example, if there are 1000 tasks in total, 250 are in the violent category, 300 in the pornographic category, 150 in the advertising category, and 300 in the normal category. The number of tasks in each category is recorded and converted into a distribution percentage, which is 25%, 30%, 15%, and 30% respectively. Then, tasks are divided into different time periods based on the task creation time field, such as January to March 2025 and April to June 2025. The frequency of each category of tasks is counted in each period, and the fluctuations in the number of content categories appearing in different time periods are compared. For example, if there are 200 violent tasks in the first quarter and 50 in the second quarter, the distribution is considered to have changed drastically, thus identifying high-frequency and low-frequency task categories. The proportion difference between categories is used to define high-frequency tasks as those that account for more than 25% of the total number of tasks in a quarter, and low-frequency tasks as those that account for less than 10%. For example, if "advertising" accounts for only 6% in the second quarter, it is marked as low-frequency. Then, the risk level distribution of each content task category is statistically analyzed. For example, in 300 pornographic tasks, there are 210 high-risk tasks, 60 medium-risk tasks, and 30 low-risk tasks. Their risk proportion distribution is calculated to be 70%, 20%, and 10%, respectively. Based on the preset category risk distribution standard, it is judged whether there is a distribution imbalance. If the high-risk proportion exceeds the 60% threshold, the scheduling frequency of this type of task needs to be increased and the review priority needs to be raised. The task frequency structure and risk level composition of each category in the time interval are summarized. For example, the output is "Pornographic tasks are high-frequency tasks in the second quarter, with a high-risk proportion of 70%", and a risk frequency distribution structure dataset is formed.
[0140] The content category optimization submodule, based on the risk frequency distribution structure, filters the coverage of short video content categories under different time periods and distribution intensities, compares the relationship between the distribution of violation category tasks and content supplementation needs, and calculates the supplementation priority of each category task using the formula:
[0141] ;
[0142] Adjusting the content and supplementing the node distribution yields the characteristics of the coverage gap distribution. ,in, Indicates the first Number of short video tasks Indicates the first Coverage density of short video-like tasks Indicates the first Risk coverage intensity corresponding to short video tasks Indicates the first The number of covered nodes for each type of task. Indicates the first Historical variation rate of violations in short video-like tasks. Indicates the first Content category granularity for short video tasks.
[0143] Coverage gap distribution characteristics refer to the distribution of short video tasks under different categories, times, risk intensity, and node distribution in content review testing scenarios. Through comprehensive analysis using formulas, it measures the aggregate performance of the system in terms of uneven sample coverage, sparse task distribution, and risk detection blind spots in content category monitoring. It reflects the "shortcomings" and "gaps" of the system's current review coverage for various short video tasks. For example, in certain high-risk categories or specific time periods, when the distribution of task review nodes is too concentrated or the distribution density is insufficient, monitoring coverage blind spots are likely to occur, which in turn lead to content review quality risks.
[0144] No. The number of samples for each type of task is denoted as Cover density is denoted as The risk coverage intensity is The number of covered nodes is The historical difference rate is Content category granularity is First, by comparing the "content_type" and "channel_id" fields of each sample record in the task log, a category mapping group is established. Task content belonging to the same z category is extracted, and the number of such tasks is counted. For example, if the z-th task is "entertainment short videos", the original number of tasks is 65, which is normalized to 0.81. Then, the coverage density is determined by the ratio of the number of records in this type of task that have undergone at least one identity verification process to the total number of tasks. The original coverage ratio was set to 0.72, corresponding to a normalized value of 0.68. Then, based on the proportion of samples marked as high-risk in this type of task, the risk coverage strength was extracted. Set to 0.43, normalized to 0.59, representing the number of review nodes corresponding to the task. Set to 4, the normalized value is 0.57, historical variance rate. By comparing the review status returned by humans and machines for the same sample to see if they are consistent, the percentage of inconsistent records is calculated and set at 0.25, with a normalized value of 0.42. (Granularity parameter) The number of unique categories of this type of task content tag is determined by 5, with a normalized value of 0.64. Substituting these values into the formula, the calculation is as follows:
[0145] ;
[0146] Based on the parameter settings in the coverage gap task scheduling. The scoring range is divided as follows:
[0147] Interval A (low priority): ;
[0148] If the coverage density is high or the task risk level is low, there is no need to add nodes; simply include them in the regular monitoring scope.
[0149] Interval B (medium priority): ;
[0150] The tasks have some differences or uneven coverage, but the supplementary needs are in an adjustable phase and are generally not given priority.
[0151] Interval C (high priority): ;
[0152] Significant coverage gaps or concentrated historical discrepancies necessitate the deployment of supplementary review nodes in the short term.
[0153] Interval D (very high priority): ;
[0154] There are serious coverage blind spots and task distribution imbalances, requiring immediate task supplementation. These tasks have the highest scheduling priority.
[0155] The result indicates that the calculated It belongs to interval D (extremely high priority), which has exceeded the upper limit of the set high priority threshold, indicating that the first... Short video-like tasks have significant shortcomings in terms of coverage, node distribution, and review consistency. These tasks should be prioritized in the supplementary scheduling and sent to multi-node interfaces for supplementary testing.
[0156] Please see Figure 5 The feedback acquisition module includes a task allocation and identification submodule, a state comparison and judgment submodule, and a structural fluctuation calculation submodule.
[0157] The task allocation and identification submodule analyzes the matching between the node information corresponding to the supplementary task and the task allocation record based on the distribution characteristics of the coverage gap, determines whether the task instructions and node distribution are consistent, filters the structurally related task paths, and gradually optimizes the task trajectory matching by combining the allocation sequence to obtain the sample node matching path set.
[0158] Based on the distribution characteristics of coverage gaps, the execution path of supplementary tasks is identified. First, the scheduling parameter fields of each supplementary task are extracted, including task number, target review node number, task issuance timestamp, review content category, and supplementary tag. Then, the scheduling details of historical tasks are extracted from the task allocation record table and compared by task number to confirm whether the supplementary task has been actually executed by that node. When a supplementary task is scheduled for the first time and does not match a historical node record, it is determined that the task instruction and node distribution are inconsistent, and the task is recorded in the allocation anomaly list. Next, the review content category of the current task is compared with the service capability range of the node. For example, if the supplementary task is a violence-related review task, the scheduling node's capability tag must contain a violence identification tag. If the node's service range does not have this tag, its service capability is considered to be mismatched with the task. After identifying inconsistent nodes, the task number structure before and after the current task in the scheduling sequence is read and constructed according to the time sequence. The task chain path is marked with node distribution and task category fields in each path. Then, the correlation strength between paths is confirmed by filtering the node numbers of tasks with the same structure in the path. The judgment is based on whether the consecutive node numbers in the path appear repeatedly and the corresponding review categories are consistent. For example, if path 1 contains tasks T100, T101, and T102, and its nodes are Node03, Node03, and Node05, all of which are related to pornography, the path is marked as a stable structure. Conversely, if the nodes in the path change frequently and the task category distribution is unstable, the path is judged as an unstructured correlation path. Finally, after removing all unstructured paths, the remaining path set is renumbered and compared with the existing node trajectory information in the original scheduling sequence record. The sequential numbering of task trajectories in the path is gradually corrected and supplemented, and the order of task paths with skipped numbers or breaks is adjusted to maximize the continuity of nodes and review categories in the path, forming a sample node matching path set.
[0159] The status comparison and judgment submodule compares the review status of each task with the status of the original record based on the sample node matching path set, filters out the task categories and corresponding nodes with differences, optimizes the grouping statistics of the status difference content, and obtains the distribution of review status differences.
[0160] After receiving the sample node matching path set, the review status field is extracted for each task within each path. This status value is then compared item by item with the expected status field in the original task record. The comparison is based on an equality judgment operation of the field values. For example, a complete equality match is achieved between review statuses of "PASS" and "REJECT". If they are not equal, it is judged as an inconsistent status. For instance, in the path of task number T105, Node04 returns a status of "PASS", while the original record status is "REJECT". This is marked as a status difference task item. All status difference task items are categorized and summarized by task category field. For example, among all status differences, tasks related to pornography account for 40%, those related to violence account for 30%, those related to advertising account for 20%, and those related to normal accounts for 10%. The node numbers of the difference tasks are then statistically analyzed. The system counts the number of times each node appears in state difference tasks and determines whether the frequency exceeds a set baseline threshold. The default baseline value is that the proportion of nodes hitting difference tasks does not exceed 15%. If a node participates in 100 tasks in all tasks and is marked as a state difference task 20 times, the deviation ratio is 20%. If it exceeds the threshold, it is marked as a state fluctuation node. Then, all nodes are sorted from high to low according to the difference frequency to construct a node difference distribution sequence. The system then performs group statistics on the state difference of multiple nodes under the same category and summarizes the difference concentration under different review content types. For example, Node06 has 8 differences in advertising tasks, accounting for 25% of its total advertising tasks. It is marked as an advertising state difference node. Finally, the system outputs the review state difference distribution of each node and task category.
[0161] The structural fluctuation calculation submodule analyzes the task structure involved based on the distribution of differences in the audit status, judges the changes in sample categories and node distribution, compares the differences in structural combinations between sample categories and nodes, judges the trend of structural changes, and obtains the sample feedback structural fluctuation.
[0162] Read the sample category field and node distribution field from the differential task structure, and analyze the sample content type tags corresponding to each task, such as "pornographic images and text," "advertisement comments," and "short videos." Extract the structural fields and perform grouping operations. Then, extract the node number field to which the task belongs, and calculate the distribution density of the node numbers under the structural category. For example, in the sample type "short videos," there are 200 differential tasks, of which Node01 undertakes 80 tasks, Node02 undertakes 60 tasks, and Node03 undertakes 60 tasks. Their node distribution ratios are 40%, 30%, and 30%, respectively. Subsequently, compare the differences in node composition between the current task structure and the historical task structure, calculate the changes in node composition under the same category, and determine whether there are any new nodes or existing nodes leaving the task. The system marks the degree of change in the structure combination. For example, if the number of new nodes exceeds 50% of the original number of nodes, it is marked as a serious structural fluctuation. Then, it judges the trend of the increase or decrease of the number of tasks in the structural path. For example, if the number of violent crime-related paths increases from 120 last month to 200 this month, it is marked as an upward trend in the structure. Next, it judges whether there is a change in the concentrated or balanced distribution of nodes in the path. If the difference in the number of tasks between two nodes exceeds 30, it is considered as a structural tilt. Finally, it outputs a list of changes in the combination of the current sample category and node number under the multi-task structural path, and marks the degree of fluctuation of each type of task structure. For example, the fluctuation level is divided into four levels: "stable", "slight fluctuation", "moderate fluctuation" and "serious fluctuation". The level is determined based on the statistical results of the change rate of the number of tasks and the number of changes in the number of nodes, and the sample feedback structure fluctuation is output.
[0163] Please see Figure 6 The report archiving module includes a feedback parsing submodule, a consistency judgment submodule, and a coverage distribution construction submodule;
[0164] The feedback analysis submodule analyzes the feedback content in the test task data based on the sample feedback structure fluctuation, determines the category of changes in task parameters and test case parameters between nodes, compares the differences between interface feedback data and original review results, filters the node allocation in scheduling phase parameters, and obtains structural difference identification results.
[0165] The feedback records in the test task data are read one by one, and the field content is extracted, including feedback text, review result, task number, execution node, execution time, review identity type, etc. Then, the original test case parameter information corresponding to the task is extracted, including preset review result, review type, content tag, risk level, channel identifier, service type, etc. Field-level comparison operation is performed on the task parameters and test case parameters to determine their changes in the execution version at each node. The judgment standard is the number of fields with different values in the original test case parameters at a certain node. For example, if the service type in the original test case is "Web", but the service type is recorded as "converged media" when Node06 is executed, it is a type of change. If more than two parameter field values are different, the task is marked as "multi-field change", otherwise it is "single-field change" or "no change". Then, the interface feedback data fields are extracted, including interface status code, review conclusion, rejection reason, etc. The process involves comparing the task's parameters, such as risk labels, with the expected audit results of the original use cases. For example, if the original expectation for task number T205 is "REJECT", and the feedback after execution by Node03 is "PASS", the difference fields are the audit conclusion and risk label. This is recorded as a feedback difference task. At the same time, the node allocation information of the task is read from the scheduling phase parameter record. Fields include allocation order number, node binding time, execution round, etc., to determine whether there is rescheduling or node switching between the tasks. If a task is scheduled from Node02 to Node04 and the execution rounds are different, the number of node switching for that task is recorded. Tasks with more than 1 node allocation change are identified as "multi-round node allocation" tasks. Finally, the above task parameter change categories, feedback content difference items, and node allocation records are merged and sorted to form a complete record line of task structure differences. The summary results form the structure difference identification results.
[0166] The consistency discrimination submodule compares the synchronization between scene switching time and sample distribution content labeling time based on the structural difference identification results, judges the consistency of feedback from each node in the manual review difference clustering scenario, optimizes the matching relationship between feedback labeling content and node results, and obtains feedback result matching features.
[0167] Extract the task scheduling time and sample annotation time fields from each structural difference record. The sample annotation time refers to the time when the text content was manually or automatically reviewed and annotated. Calculate the minute-level time difference between the scheduling time and the annotation time, and determine if the time difference is less than a set synchronization threshold. The default synchronization threshold is 15 minutes. If the task scheduling time is 10:20 AM on August 3, 2025, and the sample annotation time is 10:30 AM, the time difference is 10 minutes, and it is determined to be a synchronous matching task. Conversely, if the difference exceeds 15 minutes, it is marked as an "asynchronous task." Then, in the scenario of manual review difference clustering, extract the feedback content of all nodes that executed the task. Perform a consistency judgment operation on the review results of the same content under different nodes. For example, if task number T501 is executed on Node01, Node02, and Node04 respectively, where Node0... The review conclusions for Node 1 and Node 04 are "PASS", and for Node 02 it is "REJECT". The system checks if more than 50% of nodes report the same status. If so, they are considered consistent; otherwise, they are considered inconsistent node groups. Differences are then categorized and statistically analyzed by task type and node number. For example, in advertising tasks, the proportion of inconsistent node combinations is 60%, while in normal tasks it is only 20%. This proportion is output by task type. Then, the feedback tag field is extracted and compared with the actual node result field. For example, if a task feedback tag is "requires review", but the node returns "PASS", this difference is recorded, and the tag inconsistency is added to the task tag comparison table. Finally, the data is aggregated by node number, scenario type, and content category, and the matching degree data between the feedback results and tags for all feedback items is output. The final output is the feedback result matching feature.
[0168] The coverage distribution construction submodule determines the coverage between node information and scene labels based on the feedback result matching features, compares the correlation between the distribution features of task nodes and the scene label configuration, identifies the node and label configuration in the task archiving parameters during the scheduling phase, optimizes the correspondence between scene and node task, and obtains multi-scene coverage distribution results.
[0169] First, obtain the node ID and scene label fields from all scheduled tasks. Then, construct a set of "node-scene label" tuples, forming a matching pair in each scheduling record. For example, task T300 is executed by Node05, and the scene is labeled "short video pornography review," so the extracted tuple is (Node05, short video pornography). Count the number of node IDs under the same label in all tasks to determine if the node covers the current scene label. The criterion is that a node appears in the scene with the label at least 5 times, which is considered a valid covering node. If Node05 executes tasks 6 times under the label, it is considered to cover the scene. Then, compare the node distribution characteristics, i.e., whether the number of tasks each node participates in is evenly distributed across multiple labels. Calculate the standard deviation of the number of label types each node participates in and the task distribution. If the standard deviation exceeds the set benchmark value of 10, it is considered... If the distribution of nodes is uneven, the tag-node preset mapping relationship is extracted from the scene tag configuration table. The field is the set of recommended nodes under the scene tag. For example, the recommended nodes for the "community message scene" are Node02, Node04, and Node08, while the main execution nodes in the actual task are Node01 and Node06. This indicates a deviation between the scheduling execution and the configuration, and is marked as "tag configuration deviates from task". Then, the archive field in the scheduling stage record is extracted, and the final archive node and scene tag for each task are extracted one by one to build an archive mapping table. The archive data is then cross-matched with the tag configuration preset table to calculate the task attribution accuracy. If the matching rate between the node and configuration in the archive of a certain scene tag is less than 60%, it is determined that the mapping relationship between the scene and the node needs to be adjusted. Finally, the execution mapping results of the task nodes under all scene tags are combined to output the multi-scene coverage distribution results.
[0170] A content moderation quality monitoring method based on real-time probing includes the following steps:
[0171] S1: Based on the test start time parameter, analyze the matching relationship between the time and the scene status label, determine the consistency between the current time and the preset sensitive period label, compare the correspondence between the scene label and the scheduling node number, filter the corresponding use case library number, determine the association between the use case item set and the scheduling node number, determine the scheduling task order, and obtain the test allocation sequence parameter.
[0172] S2: Based on the test allocation sequence parameters, determine the correspondence between test case number and service type, analyze the matching of application identifier and channel identifier, compare the content text with the expected review result, determine the difference between the difference node and the content returned by the feedback interface, identify the matching of test case items and scheduling nodes, and adjust the differences in the data performance of test case items to obtain a detailed comparison of test case deviations.
[0173] S3: Based on the use case deviation comparison details, analyze the inconsistency text and the coverage of the review identity, determine the frequency of task distribution in the content category, filter the frequency of occurrence of risk level items, compare the distribution of tasks within the submission time range, optimize the content category supplementary tasks in the short video violation identification scenario, calculate the differences between task types, and obtain the distribution characteristics of the coverage gap.
[0174] S4: Based on the distribution characteristics of the coverage gap, determine the allocation of sample supplementation tasks, analyze the correlation between node number and audit result item, compare the interface return status with the original task data, filter the sample category items of matching tasks and new tasks, and obtain the sample feedback structure fluctuation.
[0175] S5: Based on the fluctuations in the sample feedback structure, analyze the test task data and scheduling stage parameters, optimize the matching of sample distribution content labeling and scene switching time, compare the feedback consistency under the manual review difference clustering scenario with the differences in machine review items, and obtain multi-scenario coverage distribution results.
[0176] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A content review quality monitoring system based on real-time probing, characterized in that, The task scheduling module, the use case switching module, the parameter supplement module and the feedback collection module are comprised. The task scheduling module analyzes the consistency of the scene state label and the sensitive period label based on the dial test starting time parameter, judges the correspondence relationship between the scene label and the scheduling node number, screens the use case library number, determines the association of the use case item set and the node, obtains the dial test distribution sequence parameter, judges the correspondence relationship between the use case number and the service type based on the dial test distribution sequence parameter, analyzes the matching of the application identifier and the channel identifier, compares the content text and the expected audit result, identifies the use case item and the node matching condition, and obtains the use case deviation comparison details. The parameter supplement module analyzes the inconsistent text and the audit identity coverage based on the use case deviation comparison details, judges the content category task distribution frequency, screens the risk level occurrence frequency, compares the submission time task distribution, optimizes the short video violation identification task supplement, and obtains the coverage gap distribution characteristics. The feedback collection module judges the sample supplement task distribution based on the coverage gap distribution characteristics, analyzes the node number and the audit result association, compares the interface return state and the original task data, screens the matching and the new sample category item, and obtains the sample feedback structure fluctuation. The dial test distribution sequence parameter includes task distribution identifier, execution priority, sequence number, the use case deviation comparison details include difference type identifier, influence range mark, difference source mark, the coverage gap distribution characteristics include gap category identifier, gap distribution proportion, gap influence level, and the sample feedback structure fluctuation includes structure change category, change node distribution and change influence degree. 2.The real-time dial test based content review quality monitoring system of claim 1, wherein, The task scheduling module includes a time label matching sub-module, a node number screening sub-module and a sequence parameter generation sub-module. 3.The real-time probe-based content review quality monitoring system of claim 1, wherein, The time label matching sub-module judges the overlap state of the time parameter and the sensitive period label based on the dial test starting time parameter, combines the set sensitive period label and the current time characteristics, analyzes the applicable classification of the business scene label at this time node, determines the mapping relationship between time and scene state, and obtains the time scene corresponding type. The node number screening sub-module screens the scene labels involved based on the time scene corresponding type, compares the business scene categories corresponding to each scheduling node number, judges the consistency of the node number and the scene label, optimizes the matching situation of the node number set and the use case library number, identifies the node number set of the same scene type, and obtains the node and use case matching range. The sequence parameter generation sub-module adjusts the arrangement order of each use case item according to the node and use case matching range, analyzes the distribution of the scheduling node number and the use case item combination, optimizes the relationship between the task distribution identifier and the execution priority, integrates the scheduling task parameters in turn, and obtains the dial test distribution sequence parameter. The use case switching module includes a service matching sub-module, a difference node identification sub-module and a result adjustment sub-module. 4.The real-time dial test based content review quality monitoring system of claim 1, wherein, The service matching submodule analyzes the correspondence between the use case number and the service type based on the dial test allocation sequence parameter, judges whether the type of the service associated with each use case is accurate, compares whether the application identifier and the channel identifier are consistent with the parameters associated with the scheduling node, identifies the matching situation of the content text and the review state, and obtains the service state matching degree; The difference node identification submodule analyzes the review state returned by multiple scheduling nodes based on the service state matching degree, judges the similarities and differences between the feedback results of each node and the expected review state, filters the nodes with different review states, optimizes the feedback distribution between the node and the majority nodes, and obtains the feedback node deviation amplitude; The result adjustment submodule compares the feedback node deviation amplitude with the correspondence between the use case number and the node number, analyzes the state distribution of the same use case on each node, calculates the concentration degree of the state distribution of the difference nodes, obtains the use case data difference breadth, adjusts the state distribution of each node, filters the abnormal fluctuation items, and obtains the use case deviation comparison details. 5.The real-time probe-based content review quality monitoring system of claim 1, wherein, The parameter supplement module includes a text coverage identification submodule, a task distribution filtering submodule, and a content category optimization submodule; The text coverage identification submodule judges the distribution of the text content and the review identity in all tasks based on the use case deviation comparison details, filters the text not covered by manual or machine review, compares the coverage differences of each identity, optimizes the identity allocation rules, and obtains the review coverage missing proportion; The task distribution filtering submodule filters the distribution characteristics of the tasks under the content category and the time interval based on the review coverage missing proportion, judges the distribution differences between high-frequency category tasks and low-frequency category tasks, compares the proportion of risk levels in each content category, optimizes the category distribution, and obtains the risk frequency distribution structure; The content category optimization submodule filters the coverage of short video content categories under different time periods and distribution intensities based on the risk frequency distribution structure, compares the relationship between the distribution of violation category tasks and the content supplement demand, calculates the supplement priority of each category task, adjusts the content supplement node distribution, and obtains the coverage gap distribution characteristics. 6.The real-time probe-based content review quality monitoring system of claim 1, wherein, The feedback collection module includes a task allocation identification submodule, a state comparison determination submodule, and a structure fluctuation calculation submodule; The task allocation identification submodule analyzes the matching situation between the node information corresponding to the supplement task and the task allocation record based on the coverage gap distribution characteristics, judges whether the task instruction and the node distribution are consistent, filters the task path associated with the structure, and gradually optimizes the task trajectory matching combined with the allocation sequence, and obtains the sample node matching path set; The state comparison determination submodule compares whether the review state of each task and the original record state are consistent based on the sample node matching path set, filters the task categories and corresponding nodes with differences, optimizes the grouping statistics of the state difference content, and obtains the review state difference distribution situation; The structural fluctuation calculation sub-module analyzes the task structure involved based on the audit state difference distribution situation, judges the change situation of sample categories and node distribution, compares the structural combination difference between sample categories and nodes, judges the structural change trend, and obtains sample feedback structural fluctuation.
7. The real-time probe-based content review quality monitoring system of claim 1, wherein, The report archiving module also includes a feedback analysis sub-module, a consistency discrimination sub-module, and a coverage distribution construction sub-module. The report archiving module includes a feedback analysis sub-module, a consistency discrimination sub-module, and a coverage distribution construction sub-module. 8.The real-time dial test based content review quality monitoring system of claim 7, wherein, The feedback analysis sub-module analyzes the feedback content in the dial test task data based on the sample feedback structural fluctuation, judges the change category of task parameters and use case parameters between nodes, compares the difference between interface feedback data and original audit results, screens the node allocation in the scheduling stage parameter, and obtains the structural difference identification result. The consistency discrimination sub-module compares the synchronization between scene switching time and sample distribution content labeling time based on the structural difference identification result, judges the consistency of each node feedback in the artificial audit difference clustering scene, optimizes the matching relationship between feedback labeling content and node results, and obtains the feedback result matching feature. The coverage distribution construction sub-module judges the coverage between node information and scene labels based on the feedback result matching feature, compares the correlation between the distribution characteristics of task nodes and scene label configurations, identifies the node and label configuration in the scheduling stage task archiving parameter, optimizes the ownership correspondence of scenes and node tasks, and obtains the multi-scene coverage distribution result. The report archiving module also includes a feedback analysis sub-module, a consistency discrimination sub-module, and a coverage distribution construction sub-module. 9.A method for monitoring content review quality based on real-time dial testing, characterized in that, S1: Based on the dial test start time parameter, analyze the matching relationship between the time and the scene state label, judge the consistency between the current time and the preset sensitive period label, compare the corresponding relationship between the scene label and the scheduling node number, screen the application case library number, judge the association between the use case item set and the scheduling node number, determine the scheduling task order, and obtain the dial test distribution sequence parameter; S2: Based on the dial test distribution sequence parameter, judge the corresponding relationship between the use case number and the service type, analyze the matching situation of the application identifier and the channel identifier, compare the content text and the expected audit result, judge the difference between the difference node and the feedback interface return content, identify the use case item and the scheduling node matching situation, and adjust the difference of the use case item in the data performance, and obtain the use case deviation comparison details; S3: Based on the use case bias comparison details, analyze the inconsistent text and audit identity coverage, judge the task distribution frequency in the content category, filter the appearance frequency of risk level items, compare the distribution of tasks within the submission time range, optimize the content category supplement task in the short video violation recognition scene, calculate the difference between task types, and obtain the coverage gap distribution characteristics; S4: Based on the coverage gap distribution characteristics, judge the distribution of sample supplement tasks, analyze the association between node number and audit result items, compare the interface return state and original task data, filter the sample category items of matching tasks and new tasks, and obtain the sample feedback structure fluctuation; S5: Based on the sample feedback structure fluctuation, analyze the test task data and scheduling stage parameters, optimize the matching of sample distribution content labeling and scene switching time, compare the feedback consistency in the artificial audit difference clustering scene and the difference of machine audit items, and obtain the multi-scene coverage distribution result.
Citation Information
Patent Citations
Video content auditing method and system
CN117173608A
Test method and test management system for intelligent PON gateway test
CN120614542A