Video auditing method and device, electronic equipment and robot

By analyzing video data from multiple perspectives and conducting human-machine collaborative review, the problems of low efficiency and insufficient depth in existing video review technologies have been solved. This has enabled the automation and closed-loop learning of machine review, thereby improving review efficiency and accuracy.

CN121921697APending Publication Date: 2026-04-24PAXINI TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PAXINI TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2025-12-15
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing video review methods are inefficient and lack depth. They cannot perform feedback loop learning for models, the review standards are inconsistent, and reliance on manual review is subjective and costly.

Method used

Multi-view operation video data is analyzed, action unit sequences and key features are extracted through a large video language model, and review is conducted by combining explicit and implicit knowledge bases. In-depth review is carried out using a human-machine collaborative review module, and the model is optimized by updating the video data knowledge base through feedback.

Benefits of technology

It has achieved automated processing of machine review, improved review efficiency and accuracy, reduced costs, built a closed-loop architecture of review, training and learning, ensured that the review results are verifiable, and improved data credibility and review depth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921697A_ABST
    Figure CN121921697A_ABST
Patent Text Reader

Abstract

The invention discloses a video auditing method and device, electronic equipment and a robot. The method comprises the following steps: acquiring multi-view operation video data; analyzing the multi-view operation video data through a preset video language large model to obtain an action unit sequence; key feature extraction is carried out on the action unit sequence to obtain a feature vector corresponding to the action unit sequence, and retrieval is carried out in a preset video data knowledge base based on the feature vector to obtain target knowledge data matched with the feature vector; performing fusion analysis on the multi-view operation video data and the target knowledge data to generate a first auditing report, and sending the first auditing report to a man-machine collaborative auditing module; and receiving an auditing result returned by the man-machine collaborative auditing module, and generating a target auditing report according to the auditing result. According to the method, a three-in-one closed-loop structure of auditing, training and learning is constructed, so that the auditing efficiency and accuracy are improved, the auditing standard is unified, the video data knowledge base is updated in a closed-loop manner, and the video language large model is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of robot learning and data processing technology, specifically relating to a video review method, device and electronic equipment, and robot. Background Technology

[0002] Currently, in the field of advanced robot imitation learning technology, the core objective is to enable robots to learn and master complex skills by observing the meticulous operations of human experts. By learning directly from human demonstrations, robots can gain greater flexibility and adaptability. To acquire video data of human operations as "fuel" for robot imitation learning, a data acquisition environment is typically set up, and domain professionals are invited to strictly follow standard operating procedures to perform tasks, recording the entire operation process from multiple viewpoints.

[0003] After obtaining the raw multi-view video data, quality review and data annotation are performed to extract accurate and error-free demonstration data from the massive amount of raw data. However, existing review processes mostly rely on teams of human reviewers. In implementing the embodiments of this application, the inventors discovered that the prior art has at least the following problems: low review efficiency, inconsistent review standards, insufficient review depth, and the inability of the reviewed data to be used for feedback loop learning of the model. Summary of the Invention

[0004] The implementation method of this application mainly addresses the technical problems of low efficiency and insufficient depth in video review, which prevents further feedback and closed-loop learning of the model.

[0005] To address the aforementioned technical problems, one technical solution adopted in this application is to provide a video review method, comprising: acquiring multi-view operation video data; the multi-view operation video data including video data acquired simultaneously from multiple different shooting angles for recording the same operation process; parsing the multi-view operation video data using a preset video language big data model to obtain an action unit sequence; extracting key features from the action unit sequence to obtain a feature vector corresponding to the action unit sequence, and retrieving target knowledge data matching the feature vector from a preset video data knowledge base based on the feature vector; performing fusion analysis on the multi-view operation video data and the target knowledge data using the video language big data model to generate a first review report; sending the first review report to a human-machine collaborative review module; the human-machine collaborative review module performing in-depth review operations on the first review report and outputting review results; receiving the review results returned by the human-machine collaborative review module, and generating a target review report based on the review results.

[0006] Optionally, the step of parsing the multi-view operation video data using a preset video language model to obtain an action unit sequence includes: analyzing the multi-view operation video data using the video language model to identify human posture, hand movements, object interactions, and scene information in the multi-view operation video data; summarizing the identified human posture, hand movements, object interactions, and scene information to form content information of the multi-view operation video data; and segmenting the content information to generate independent action unit sequences.

[0007] Optionally, the method further includes constructing the video data knowledge base, comprising: acquiring objective knowledge data, digitizing and structuring the objective knowledge data to generate minimum knowledge units, and summarizing the minimum knowledge units to obtain an explicit rule base; acquiring subjective experience data from human experts, vectorizing the subjective experience data to generate soft knowledge units, and summarizing the soft knowledge units to obtain an implicit experience base; and integrating the explicit rule base and the implicit experience base to generate the video data knowledge base.

[0008] Optionally, the step of fusing and analyzing the multi-view operation video data and the target knowledge data through the video language big model to generate a first review report includes: matching the action unit sequence of the multi-view operation video data with the knowledge units of the target knowledge data to generate contextual prompts; the contextual prompts include operation actions, object interactions, scene information, and preset rule matching results; inputting the contextual prompts into the video language big model, and the video language big model performing semantic analysis, anomaly detection, and compliance judgment on the action unit sequence based on the contextual information to generate a first review report.

[0009] Optionally, the human-machine collaborative review module is specifically used for: receiving the review results of human experts judging each item in the first review report, confirming whether each review result matches the preset conditions, and if not, modifying and improving the first review report to generate a second review report; receiving the review results returned by the human-machine collaborative review module and generating a target review report based on the review results includes: if the review results match the preset conditions, setting the first review report as the target review report and increasing the confidence of the multi-view operation video data in knowledge base matching and deep reasoning analysis; if the review results do not match the preset conditions, obtaining the data of the correction operation performed by human experts on the first review report, generating a second review report based on the correction operation data and the first review report, setting the second review report as the target review report, and recording the correction operation data in the video data knowledge base to realize the content update of the video data knowledge base.

[0010] Optionally, the method further includes: inputting the target review report into the video data knowledge base to update the video data knowledge base; wherein, updating the video data knowledge base includes: updating the smallest knowledge unit in the explicit rule base and updating the soft knowledge unit in the implicit experience base; inputting the target review report and the updated video data knowledge base into the video language big model, and iteratively learning the video language big model based on the review results in the target review report and the updated video data knowledge base to adjust the parameters of the video language big model, thereby obtaining an adjusted video language big model; the adjusted video language big model is used to parse subsequently acquired multi-view operation video data.

[0011] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a video review device, comprising: a data acquisition module for acquiring multi-view operation video data; the multi-view operation video data includes video data acquired simultaneously from multiple different shooting angles for recording the same operation process; a data parsing module for parsing the multi-view operation video data using a preset video language model to obtain action unit sequences; a data retrieval module for extracting key features from the action unit sequences to obtain feature vectors corresponding to the action unit sequences, and retrieving target knowledge data matching the feature vectors from a preset video data knowledge base; a first review module for fusing and analyzing the multi-view operation video data and the target knowledge data using the video language model to generate a first review report; a second review module for sending the first review report to a human-machine collaborative review module; the human-machine collaborative review module is used to perform in-depth review operations on the first review report and output review results; and a report output module for receiving the review results returned by the human-machine collaborative review module and generating a target review report based on the review results.

[0012] To solve the above-mentioned technical problems, another technical solution adopted in the embodiments of this application is: to provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.

[0013] To solve the above-mentioned technical problems, another technical solution adopted in the embodiments of this application is to provide a robot that includes the above-mentioned electronic equipment.

[0014] To solve the above-mentioned technical problems, another technical solution adopted in the embodiments of this application is: providing a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by an electronic device, the electronic device performs the method described in any of the above-mentioned embodiments.

[0015] Unlike related technologies, this application provides a video review method, apparatus, electronic device, and robot. It proposes automated machine review to replace frame-by-frame comparison in manual review, improving efficiency, reducing costs, and freeing human experts from repetitive tasks. By utilizing knowledge retrieval and matching from a video data knowledge base, machine judgment is grounded in evidence and depth, accurately matching review content with knowledge data to ensure the review process is based on evidence and the results are verifiable, thus enhancing data credibility. Furthermore, a closed-loop architecture integrating review, training, and learning is constructed. Review feedback updates the video data knowledge base and optimizes the video language model, enabling the system to learn independently. Review accuracy and depth continuously improve with use, allowing issues discovered during review to be fed back to upstream processes, driving iterative optimization of operational standards. Moreover, the real-time closed-loop feedback and explicit reports of deep knowledge from machine review provide more standardized review knowledge. Attached Figure Description

[0016] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0017] Figure 1 This is a flowchart illustrating a video review method provided in an embodiment of this application.

[0018] Figure 2 This is a flowchart illustrating a method for receiving the audit results returned by a human-machine collaborative audit module and generating a target audit report based on the audit results in a video audit method provided in this application embodiment.

[0019] Figure 3 This is a flowchart illustrating a video review method provided in another embodiment of this application.

[0020] Figure 4 This is a schematic diagram of a video review device provided in an embodiment of this application.

[0021] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0023] It should be noted that, unless otherwise specified, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device schematic diagram or the order in the flowchart.

[0024] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.

[0025] In the field of advanced machine learning, especially in the cutting-edge area of ​​imitation learning, the core objective is to enable machines to learn and master complex skills by observing the precise operations of human experts. To make machine imitation learning more accurate and efficient, it is necessary to collect massive amounts of high-quality and highly consistent video data of human operations as "fuel" for training machine models.

[0026] A common approach is to build a professional data acquisition environment, such as a data acquisition room equipped with multiple high-definition, high-speed synchronous cameras and motion capture systems. Specialized workers in the field are invited to strictly follow standard operating procedures to perform the corresponding operations, recording the entire process from multiple viewpoints. After obtaining the raw operation videos, the most crucial step is to conduct quality control and data annotation on these videos. The quality of these operation videos varies greatly; some substandard videos cannot be used as "feed" for machine training and may even worsen the training effect. Therefore, it is necessary to refine the videos to obtain accurate and flawless ones that can serve as high-quality demonstration data. The quality of these videos directly determines the performance of the machine learning model.

[0027] Currently, the review process largely relies on manual review. Reviewers need to watch the operation video frame by frame and from multiple angles, rigorously comparing every action of the professional worker—from posture and force to timing—with the pre-set standard operating procedure to determine if there are any operational deviations, process errors, redundant actions, or potential risks that may lead to operational failure. For fully qualified data segments, precise action segmentation and semantic annotation are then performed. Only after this process can the annotated data be input into the machine learning model for learning and training.

[0028] However, manual review has many drawbacks, such as inconsistent review standards, insufficient review depth, and lack of feedback loop. The results of manual review are highly dependent on the personal factors of the reviewers, exhibiting significant subjectivity. Different reviewers may have different evaluations of the same concept or data, leading to chaotic training data received by the machine. Manual review only focuses on a superficial assessment of right and wrong, without analyzing the underlying causes. The one-way review process cannot quickly and systematically update the problems and information discovered during the review process into the knowledge base, thus failing to form a closed loop. The improvement of data quality can only rely on human-recognized experience.

[0029] Based on this, this application proposes a video review method that solves a series of problems caused by manual review, such as low review efficiency, inconsistent review standards, insufficient review depth, high cost, and lack of effective feedback loop, through machine review training and learning.

[0030] In some embodiments of this application, such as Figure 1 As shown, the method includes, but is not limited to, the following steps: 101: Acquire multi-view operation video data; the multi-view operation video data includes video data acquired simultaneously from multiple different shooting angles for recording the same operation process.

[0031] It is understood that the multi-view operation video is an operation video recorded from multiple perspectives of the professional worker's operation. Different perspectives are used to record the same operation subject, the same operation process, and the same moment in time. This avoids invalid video data due to data loss caused by misaligned perspectives or process deviations. The goal is to ensure that the captured operation video details are accurately captured, the perspective is unobstructed, and the time is synchronized.

[0032] Specifically, a multi-view data access and synchronization module can be set up to capture the actions of professional workers through multiple acquisition devices and receive the raw video stream. The core of this module is to perform precise time synchronization. A precise time protocol can be used to ensure that the time error of all video frames is controlled at the millisecond level. Then, the timestamp-aligned data is standardized in format and finally integrated into a multimodal data packet under a unified time as the multi-view operation video data to provide high-quality input.

[0033] The data acquisition devices can be multiple cameras or sensors. Cameras can record the movements and postures of professional workers, such as the bending angle of their arms; acquire three-dimensional spatial position information, such as the spatial relative position of their hands and tools; and capture high-speed, minute movements, such as the action of using a small tool to pick up small objects. Sensors can collect precise information about the professional worker's movements; for example, motion capture sensors can help collect the motion trajectories of various body parts and information on coordinated limb movements. Acquiring accurate and effective multi-view operation videos is fundamental for video review and ensuring the smooth operation of machine learning.

[0034] 102: The multi-view operation video data is parsed using a preset video language big model to obtain an action unit sequence.

[0035] It should be noted that the Vision-Language Model (VLM) is a pre-defined, powerful language model that serves as the foundational model for video review. It can simultaneously process both video visual information and textual language information. The core capabilities of this foundational model are: in-depth understanding of video content from a single perspective, and several powerful basic capabilities. These basic capabilities include, but are not limited to: high-precision human pose estimation (especially tracking fine hand movements), object detection and tracking, temporal action segmentation, and video question answering and description generation. For example, the multimodal model Qwen3-VL-8B-Thinking can be selected. This foundational model is only used as a primary model for iterations in the review process; the specific model chosen is not limited.

[0036] In some embodiments, before the video language model is formally implemented in the video review method provided in this application, it undergoes supervised fine-tuning using existing historical datasets that have been reviewed by human experts. Supervised fine-tuning involves training the video language model using specific, pre-labeled data to better perform specific tasks; the historical dataset includes operation video clips, manually labeled tags, and natural language review comments written by experts. Specifically, data fine-tuning is performed using labeled operation video clips, expert review comments, and other historical data to enable the video language model to learn the specific semantics of operation actions, such as understanding that touching point A with a finger constitutes an operation error.

[0037] The goal of supervised fine-tuning is to align the judgment criteria and language style of the video language model with those of human experts. By injecting domain-specific content moderation knowledge into the video language model, it addresses some of the cold-start issues and ensures that the model possesses high accuracy and professionalism from the outset. Cold start essentially occurs when the model lacks prior knowledge and must perform tasks from a near-blank state. Therefore, supervised fine-tuning uses high-quality, manually labeled data to fine-tune the basic video language model, providing an effective initial state and preventing unstable training or chaotic output.

[0038] Understandably, expert reviews are not simply written descriptions, but rather opinions from human experts that include specific experience and reasons. For example, an evaluation of an error in a particular operation must include: why it was wrong, and the specific reasons for the error. Its core value lies in solving the pain point of reviewing systems that only provide results without explanation. It enables machines to learn not just by judging right or wrong, but by analyzing the specific reasons, resulting in more accurate conclusions.

[0039] In some embodiments, the step of parsing the multi-view operation video data using a preset video language big data model to obtain an action unit sequence includes: analyzing the multi-view operation video data using the video language big data model to identify human posture, hand movements, object interactions, and scene information in the multi-view operation video data; summarizing the identified human posture, hand movements, object interactions, and scene information to form content information of the multi-view operation video data; and segmenting the content information to generate independent action unit sequences.

[0040] It's important to note that the video language big data model transforms continuous, multi-view, and unstructured operational video data into discrete, structured, and analyzable sequences of action units through multi-dimensional recognition, structured content aggregation, and temporal segmentation. Essentially, it converts visual information into semantic operational steps. The recognition of human posture, hand movements, object interactions, and scene information aims to address the incompleteness of single-dimensional recognition; that is, it requires integrating and aggregating information from different dimensions and contents to collectively reflect the content information of multi-view operational video data. This content information is a description of the overall operational panorama based on timeline alignment and multi-view, multi-dimensional fusion. For example, after aggregating the recognition information, we obtain content information such as "at a certain moment, in what scene, what action did the operator perform, and what object did they interact with?" This content information is then segmented to obtain independent sequences of action units. A sequence of action units represents the indivisible minimum functional operational steps in an operation; it consists of multiple action units ordered by time, corresponding to the decomposition of a complete operation.

[0041] Understandably, segmenting the operation content information into independent action unit sequences enables precise matching of operation steps. Without specific action unit segments, continuous actions are prone to missing steps during review. Segmenting them into independent action unit sequences allows for better extraction of operation steps for review and learning purposes.

[0042] 103: Extract key features from the action unit sequence to obtain the feature vector corresponding to the action unit sequence, and retrieve target knowledge data matching the feature vector from a preset video data knowledge base based on the feature vector.

[0043] By extracting key features, structured action units can be transformed into machine-retrievalable numerical vectors. Key features are characteristic information extracted from the action unit sequence that characterizes the core attributes of the operation (such as action type, interaction object, execution speed, etc.). This effectively filters out a large amount of irrelevant and redundant information from the action unit sequence, retaining the core information for review, enabling more efficient and accurate matching between action units and the knowledge base. The feature vectors transformed from key features are generated by converting multi-dimensional key features into numerical feature vectors according to a unified standard. In this way, different key features can be transformed into a computable and comparable set of values, which can then be retrieved in the video data knowledge base to obtain matching target knowledge data. The target knowledge data is the set of knowledge units in the video data knowledge base retrieval results that have a similarity greater than a preset threshold to the feature vector. It represents the data retrieved in the video data knowledge base that is similar to the feature vector and can be used as reference content for review.

[0044] In some embodiments, the method further includes constructing the video data knowledge base; constructing the video data knowledge base includes: acquiring objective knowledge data, digitizing and structuring the objective knowledge data to generate minimum knowledge units, and summarizing the minimum knowledge units to obtain an explicit rule base; acquiring subjective experience data from human experts, vectorizing the subjective experience data to generate soft knowledge units, and summarizing the soft knowledge units to obtain an implicit experience base; and integrating the explicit rule base and the implicit experience base to generate the video data knowledge base.

[0045] It should be noted that the video data knowledge base differs from traditional knowledge bases, comprising an explicit rule base and an implicit experience base. The explicit rule base generates the smallest existing knowledge units by digitizing and structuring all written, objective knowledge. Specifically, standard operating procedures, tool manuals, safety specifications, quality checklists, and other documents in PDF or Word format are parsed into the smallest knowledge units, such as each step in a standard operating procedure. Furthermore, these smallest knowledge units are vectorized and embedded into the video data knowledge base, ensuring that every judgment in the large video language model is based on evidence, guaranteeing the objectivity and rigidity of the review. The implicit experience base transforms expert soft knowledge into machine-understandable soft knowledge units. Soft knowledge specifically refers to some unspoken, implicit knowledge of experts, such as operational intuition, muscle memory, weighing judgments in specific situations, and the smoothness of the next step anticipated by the expert's gaze while completing one step. This is soft knowledge that is overlooked and lost in ordinary reviews. Establishing such an implicit experience base is the core of the video data knowledge base proposed in this application embodiment. Specifically, through interviews and other methods, expert experience and knowledge are acquired. Their professional experience, judgment logic, and intuition about cases are refined and textualized. The resulting data is then vectorized and stored in an implicit experience base. This allows the video language model to make more insightful and human-expert-like judgments during video review. Examples of soft knowledge units include: when an operator picks up object A, a slight hesitation in their hand usually indicates uncertainty about the next instruction and should be marked as flawed data; or, when completing step N, the operator's gaze anticipates the material's location in step N+1 0.5 seconds in advance, a characteristic of high-quality operation and should be marked as valid data.

[0046] Understandably, building an explicit rule base provides the most direct, core, and clear objective basis and support for video language model review. It contains a large amount of reliable objective data, ensuring that the review meets all requirements and is more comprehensive. Building an implicit experience base, on the other hand, addresses the issue of tacit knowledge loss based on the explicit rule base. During manual review, expert intuition and experience cannot be transmitted in a standardized way; this information does not belong to the objective explicit rule base. However, in many cases, such tacit knowledge is particularly important. Therefore, textualizing and automating this tacit experience improves the depth and accuracy of the review, supporting machine-based judgment. The integration of the explicit rule base and the implicit experience base includes both rigid rules and soft experience; a single query can simultaneously retrieve both explicit rules and implicit experience—the target knowledge data.

[0047] 104: Using the video language big model, the multi-view operation video data and the target knowledge data are fused and analyzed to generate a first review report.

[0048] After completing all the preliminary preparations, the video language big data model officially started the review process. It compared and analyzed the acquired multi-view operation video data with the target knowledge data in the video data knowledge base to complete the initial review judgment.

[0049] In some embodiments, the step of fusing and analyzing the multi-view operation video data and the target knowledge data through the video language big model to generate a first review report includes: matching the action unit sequence of the multi-view operation video data with the knowledge units of the target knowledge data to generate contextual prompts; the contextual prompts include operation actions, object interactions, scene information, and preset rule matching results; inputting the contextual prompts into the video language big model, and the video language big model performing semantic analysis, anomaly detection, and compliance judgment on the action unit sequence based on the contextual information to generate a first review report.

[0050] It should be noted that the preset rule matching result is essentially a comparison result obtained by accurately comparing the sequence of action units parsed from multi-view video data with preset explicit rules and implicit experiences in the video data knowledge base. Specifically, the preset rules include standard operating procedures and safety regulations in the explicit rule base, and expert experience in the implicit experience base. Based on these, the key features of the action units are compared for similarity and logically verified. For example, if the action unit is fingers touching the bottom contact point when picking up or placing an item, the preset rule in the explicit rule base is: hands are strictly prohibited from touching the bottom contact point when picking up or placing an item; a tool must be used. Therefore, the matching result is a mismatch. Similarly, if the action unit is no "click" sound when placing an item, the preset rule in the implicit experience base is: hearing a "click" sound when placing an item indicates proper installation; not hearing it may indicate improper installation and a risk of poor contact. Therefore, the matching result is a risk present.

[0051] Understandably, matching using pre-defined rules provides clear criteria for the video language model's judgment. It directly determines whether a rule is violated or satisfied based on the matched content, reducing errors and improving the model's efficiency. Furthermore, it ensures the traceability and interpretability of the review process, clearly showing which specific rule the judgment is based on, making it verifiable.

[0052] By matching action unit sequences with knowledge units in the target knowledge data, a structured context is constructed. Based on this, deep reasoning analysis is performed to obtain a preliminary first review report. Contextual prompts are structured, semantic prompt texts that serve as input material for the video language model to reason about multi-view operation video data. Through the construction of contextual prompts, the nature of the operation, the knowledge basis, and the preliminary matching result are clarified, i.e., semantic analysis, anomaly detection, and compliance judgment. Step 104 is the main step in machine video review. Overall, the multi-view operation video data received upstream and the information in the established video data knowledge base are compared for machine-based review and judgment.

[0053] Understandably, step 104 addresses the issues of low efficiency in manual reasoning, inconsistent review standards, and insufficient review depth. The machine can quickly complete the reasoning of action units using pre-set information. Furthermore, because it is based on the same video data knowledge base, the review standards are completely consistent, all based on standard logic. In addition, the implicit experience base within the video data knowledge base allows the review in step 104 to incorporate in-depth analysis, enhancing the reliability of the review results.

[0054] 105: Send the first audit report to the human-machine collaborative audit module; the human-machine collaborative audit module is used to perform in-depth audit operations on the first audit report and output the audit results.

[0055] 106: Receive the audit result returned by the human-machine collaborative audit module, and generate a target audit report based on the audit result.

[0056] In some embodiments, the human-machine collaborative review module is specifically used to: receive the review results of human experts judging each item of the review content in the first review report, confirm whether each review result matches the preset conditions, and if not, modify and improve the first review report to generate a second review report.

[0057] Specifically, the human-machine collaborative review module provides a visual interactive interface as a bridge between human experts and machines. This interface has, but is not limited to, the following functions: synchronously playing multi-view operation videos; overlaying the analysis content and results of the video language model onto the video screen; clearly displaying the first review report given by the video language model and providing the content from the video data knowledge base; and providing feedback tools to allow human experts to confirm, correct, or supplement the judgments of the video language model.

[0058] In some embodiments, such as Figure 2 The method described above involves receiving the audit result returned by the human-machine collaborative audit module and generating a target audit report based on the audit result, including: 1061: Compare the audit results with the preset conditions. If they match, proceed to step 1062; if they do not match, proceed to steps 1063 and 1064.

[0059] 1062: Set the first audit report as the target audit report and increase the confidence of the multi-view operation video data in knowledge base matching and deep reasoning analysis.

[0060] 1063: Obtain data on the correction operations performed by human experts on the first audit report, generate a second audit report based on the correction operation data and the first audit report, and set the second audit report as the target audit report.

[0061] 1064: Record the data of the correction operation into the video data knowledge base to update the content of the video data knowledge base.

[0062] Understandably, the human-machine collaborative review module first completes the expert verification and optimization of the first review report. Then, the review results are compared with preset conditions to confirm the final target review report, and the content of the video data knowledge base is updated. The review results include confirmation (no objection), correction (errors found in the content), or supplementation (omissions). The preset conditions are the criteria for determining that the first review report can be used as the final target review report without modification. The core is that the machine review results are completely consistent with the human expert judgment, with no omissions, no errors, and no ambiguities. Specifically, this includes: accurate expert conclusions and matching knowledge basis.

[0063] The confidence level of multi-view operation video data in knowledge base matching and deep reasoning analysis is equivalent to the reliability of the matching between the multi-view operation video data and the knowledge units of the video data knowledge base after rule matching, reasoning analysis of the video language big model, and collaborative review by the human-machine collaborative review module, as well as the credibility of the resulting review result. Its core is a quantified value of reliability, combining the quality of the multi-view operation video, the accuracy of key feature extraction, the reliability of knowledge base matching, and the rigor of the reasoning logic of the video language big model. When the confidence level of the multi-view operation video data in knowledge base matching and deep reasoning analysis increases, it means that the review judgment result and matching process have increased weight in subsequent machine reviews. This allows the machine to prioritize the review judgment logic of this data when processing similar operation reviews, leading to improved efficiency and a positive cycle in subsequent reviews.

[0064] Specifically, by leveraging the complementary strengths of humans and machines, and using a human-machine collaborative review module as a vehicle, machines undertake large-scale and standardized basic review work, while human experts focus on the precise optimization of cases. Then, by generating target review reports through multi-path approaches, confidence levels are enhanced or the knowledge base is updated, thus achieving a positive cycle of more accurate review results, system capability evolution, and more efficient subsequent reviews.

[0065] In some embodiments, the method further includes: inputting the target review report into the video data knowledge base to update the video data knowledge base; wherein, updating the video data knowledge base includes: updating the smallest knowledge unit in the explicit rule base and updating the soft knowledge unit in the implicit experience base; inputting the target review report and the updated video data knowledge base into the video language big model, and iteratively learning the video language big model based on the review results in the target review report and the updated video data knowledge base to adjust the parameters of the video language big model to obtain an adjusted video language big model; the adjusted video language big model is used to parse subsequently acquired multi-view operation video data.

[0066] Understandably, inputting the target review report into the video data knowledge base for updates, and further iteratively evolving the video language model, is the core of the review, training, and learning closed loop in this application's embodiment. Its essence is to maximize the value of the target review report: both updating the video data knowledge base with the report to fill knowledge gaps, and continuously fine-tuning the video language model based on the updated knowledge and the report, achieving a positive cycle of knowledge accumulation, model upgrades, and more accurate review.

[0067] The updates to the minimum knowledge units address specific explicit rules not covered in the target audit report, such as omitted tool usage requirements, by adding new minimum knowledge units; or by correcting vague descriptions or errors in existing units. The updates to the soft knowledge units address new implicit experiences supplemented by experts in the target audit report, such as unrecorded risk characteristics or details related to high-quality operations, by adding new soft knowledge units; or by optimizing the audit judgment logic of existing experiences. Through these updates to both minimum and soft knowledge units, the video data knowledge base can be expanded and corrected during audits, continuously improving knowledge coverage and further refining standards, thus becoming more beneficial for subsequent audits.

[0068] After updating the video data knowledge base, the target review report and the updated video data knowledge base are input into the video language model to perform iterative learning and parameter adjustment. Specifically, using the data in the target review report and the new knowledge units in the updated video data knowledge base as training data, the parameters of the video language model are fine-tuned, allowing the model to learn new scenarios, rules, and experiences. The parameter adjustments of the video language model retain its original capabilities while adding or enhancing its adaptability to new knowledge. Through iterative updates of the video language model, its recognition capabilities are continuously upgraded, and its adaptability is enhanced.

[0069] By combining video data knowledge base updates with the iteration of video language models, a feedback loop for overall video review can be achieved, allowing the machine to continuously update and upgrade the knowledge base during video review, and become more adaptable to more scenarios after each review.

[0070] In some embodiments, the method can output structured data of the obtained content information through multiple application output interfaces, including: timestamps, action unit segmentation, judgment results, confidence scores, referenced original video data knowledge base text, and machine-generated natural language review comments. This structured report serves a wider range of scenarios through different interfaces, including but not limited to: automatically filtering and labeling qualified data to build high-quality imitation learning datasets; converting it into dense reward signals to guide and train machine models; and serving as teaching cases with expert analysis to train new human reviewers.

[0071] Specifically, in training new human auditors, a standardized and replicable training system is established: the data output by machine audits, with its absolute knowledge completeness and judgment consistency, provides a unified training standard. This ensures that every new auditor can quickly master the auditing criteria completely consistent with those of senior experts, fundamentally solving the problems of inconsistent standards and varying quality inherent in manual training models. Furthermore, through an instant feedback loop, the learning curve is accelerated exponentially: new auditors can receive expert-level feedback instantly during practice; the system immediately points out errors and omissions in their judgments and displays the correct basis for those judgments. This instant closed loop of practice, feedback, and correction, compared to traditional mentoring models, can significantly shorten the growth cycle of new auditors, enabling them to become competent in their positions more quickly. Explicitizing tacit knowledge bridges the cognitive gap between experts and novices: The greatest training value of this method lies in its ability to present some implicit experiences of senior experts—experiences that are "only understood intuitively"—to new auditors in clear text and case studies. This allows new auditors not only to learn the rules but also to understand the deep logic and expert intuition behind them, thus truly achieving rapid alignment at the cognitive level.

[0072] In summary, unlike related technologies, this application provides a video review method that proposes automated machine review to replace frame-by-frame comparison in manual review, improving review efficiency, reducing review costs, and freeing human experts from repetitive tasks. By utilizing knowledge retrieval and matching from a video data knowledge base, machine judgment is grounded in evidence and depth, accurately matching review content with knowledge data to ensure the review process is based on evidence and the results are verifiable, thus enhancing data credibility. Furthermore, a closed-loop architecture integrating review, training, and learning is constructed. Review feedback updates the video data knowledge base and is used to optimize the video language model, enabling the system to have self-learning capabilities. Review accuracy and depth continuously improve with use, allowing issues discovered during review to be fed back to upstream processes, driving iterative optimization of operational standards. Moreover, the real-time closed-loop feedback from machine review and the explicit display reports of deep knowledge provide more standardized review knowledge.

[0073] For example, in some embodiments, the video review method flow is as follows: Figure 3 As shown, a video data knowledge base is first established, including an explicit rule base and an implicit experience base. Then, multi-view operation videos are received and processed using timestamp alignment and format unification for structured operation. Next, a large-scale video language model is used to parse the processed multi-view operation videos, summarizing the information contained within and segmenting them to obtain action unit sequences. Further, key features of the action unit sequences are extracted, and their feature vectors are obtained. These feature vectors are then queried in the video data knowledge base to obtain target knowledge data, and a first review report is generated based on the large-scale video language model. The first review report is sent to the human-machine collaborative review module, which compares it with preset review conditions. If they match, the first review report is directly output; otherwise, a second review report modified by the human-machine collaborative review module is output. The final output review report is further used to update the video data knowledge base and optimize the large-scale video language model.

[0074] For example, taking "reviewing and applying a teaching video on assembling a mobile phone motherboard" as an example, the embodiments of this application will be further described in detail: In building the video data knowledge base, PDF documents such as "Mobile Phone Motherboard Assembly SOP Manual" and "Component Safety Operation Specifications" were uploaded to the system in batches. The system automatically segmented the documents, extracted key information, and used a pre-trained text embedding model to vectorize them before storing them in the explicit rule base. Next, through knowledge engineering interviews with several senior auditors, several implicit experiences were extracted, such as: "When a skilled worker aligns the slot, their gaze will move to the next material to be picked up 0.5 seconds in advance, which is a high-quality action," and "When placing a CPU, if a slight 'click' sound is heard, it means that the installation is in place; the absence of this sound may indicate a risk." These experiences were textualized and stored in the implicit experience base.

[0075] Then, the model is fine-tuned. Using historical data accumulated internally and already manually reviewed, each video clip and expert commentary is used as training data. Through supervised fine-tuning, the basic video language model learns the experts' language style and judgment logic, completing the cold start of the video language model.

[0076] Example 1: A video clip of a worker's operation is captured. After VLM analysis, a structured audit report is generated. For example, for a specific time segment of the video, the report content would be: { "timestamp": "[10.5, 12.3]"; "action": "Place CPU chip", "judgement": "unacceptable" "confidence": 0.98, "reason": "The operator's fingers directly touched the gold contacts on the bottom of the CPU chip, violating cleanroom operation procedures." "knowledge_source": [ {"type": "explicit", "id": "SOP-2.1.3", "text": "It is strictly forbidden to touch any chip-based electronic components with bare hands..."}, {"type": "implicit", "id": "EXP-012", "text": "One common mistake beginners make is using their fingers to help with positioning when gripping something difficult, just for convenience..."} ] } This report is pushed to the human-machine collaborative review module. After the human reviewer reviews the report and confirms that the operation error identified in the report is correct, the operation video clip is automatically marked as "unqualified" and archived, and will not enter the final training and learning iteration.

[0077] Example 2: Scoring Robot Actions. A robot model learning to assemble a motherboard is fed in video clips of its attempted actions. The video language model analyzes this data and outputs a dense reward signal. For example, for the action of placing a CPU, the video language model's output is parsed by a module: Qualitative feedback: "Two minor jitters occurred when aligning with the slot; the placement speed was too fast, posing a risk of impact." Quantitative reward: Based on the video language model's internal confidence evaluation, a comprehensive reward value is calculated; in this example, it is negative. This negative value serves as a penalty signal, updating the robot model's policy network through backpropagation, causing it to reduce jitter and slow down on the next attempt.

[0078] Example 3: Training a newly hired human auditor. A new auditor enters practice mode and reviews the same video of a violation. He only marks "finger touching the chip" as an error, citing "non-standard operation." After submission, the system interface displays his answer alongside the video language model's answer, highlighting the differences: Omission Analysis. The system will then prompt: "You have missed another error: at time t, the tweezers used by the worker are model B, which does not match the model A required by the standard operating procedure. Click here to view the original standard operating procedure." Furthermore, such as Figure 4 As shown, Figure 4 This is a schematic diagram of the structure of a video review device 10 provided in an embodiment of this application. The device 10 includes: a data acquisition module 11, a data parsing module 12, a data retrieval module 13, a first review module 14, a second review module 15, and a report output module 16.

[0079] The system includes the following modules: a data acquisition module 11, which acquires multi-view operation video data, including video data simultaneously captured from multiple different shooting angles to record the same operation process; a data parsing module 12, which parses the multi-view operation video data using a preset video language model to obtain action unit sequences; a data retrieval module 13, which extracts key features from the action unit sequences to obtain feature vectors corresponding to the action unit sequences, and retrieves target knowledge data matching the feature vectors from a preset video data knowledge base; a first review module 14, which performs a fusion analysis of the multi-view operation video data and the target knowledge data using the video language model to generate a first review report; a second review module 15, which sends the first review report to a human-machine collaborative review module, which performs a deep review of the first review report and outputs the review results; and a report output module 16, which receives the review results returned by the human-machine collaborative review module and generates a target review report based on the review results.

[0080] The video review device 10 can be a software module, which includes several instructions stored in a memory. The processor can access the memory and execute the instructions to complete the video review method described in the above embodiments.

[0081] In some embodiments, the video review device 10 can also be constructed from hardware devices. For example, the video review device 10 can be constructed from one or more chips, which can work together to complete the immersive conferencing implementation method described in the various embodiments. As another example, the video review device 10 can also be constructed from various logic devices, such as general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), microcontrollers, ARM (Acorn RISC Machine) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components.

[0082] It should be noted that the video review device 10 described above can execute the video review method for electronic devices provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the embodiments of the video review device 10 can be found in the video review method for electronic devices provided in the embodiments of this application.

[0083] Please see Figure 5 , Figure 5This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 20 includes one or more processors 21 and a memory 22. The memory 22 is connected to one or more processors 21, for example, via a bus.

[0084] Processor 21 is configured to support the electronic device 20 in performing the corresponding functions in the methods described in the above method embodiments. Processor 21 may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0085] Memory 22 is used to store program code, etc. Memory 22 may include volatile memory (VM), such as random access memory (RAM); memory 22 may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory 22 may also include combinations of the above types of memory.

[0086] The memory 22 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the video review method in the embodiments of this application. The processor 21 executes various functional applications and data processing of the video review method and video review device by running the non-volatile software programs, instructions, and modules stored in the memory 22, that is, it realizes the functions of each module or unit of the video review method and video review device provided in the above method embodiments.

[0087] The memory 22 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function. The data storage area may store data created based on the use of the immersive conferencing implementation device. In some embodiments, the memory 22 may include remotely located memories 22 relative to the processor 21, which can be connected to the video auditing device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0088] The one or more modules are stored in the memory 22. When executed by the one or more processors 21, they perform the video review method in any of the above method embodiments. For example, they perform the method steps described in the above method embodiments to realize the functions of the modules described in the above device embodiments.

[0089] The electronic device in this application embodiment may specifically be an embedded intelligent device, an engineering control computer, an MCU or DSP chip, etc.

[0090] Furthermore, embodiments of this application provide a robot, which includes an electronic device 20 and is capable of performing any of the methods described above. Figure 1 Steps 101 to 106 in the method are as follows. Figure 2 Steps 1061 to 1064 in the method are implemented. Figure 4 Functions of modules 11-16 in the document.

[0091] This application provides a non-volatile computer-readable storage medium storing computer-executable instructions that are executed by one or more processors 21, for example... Figure 5 One of the processors 21 can be configured to execute the video review method in any of the above method embodiments, for example, to perform the above-described... Figure 1 Steps 101 to 106 in the method are as follows. Figure 2 Steps 1061 to 1064 in the method are implemented. Figure 4 Functions of modules 11-16 in the document.

[0092] This application provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions, which, when executed by the electronic device, enable the electronic device to perform any of the above-described method embodiments. Figure 1 Steps 101 to 106 in the method are as follows. Figure 2 Steps 1061 to 1064 in the method are implemented. Figure 4Functions of modules 11-16 in the document.

[0093] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0094] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A video review method, characterized in that, The method includes: Acquire multi-view operation video data; the multi-view operation video data includes video data acquired simultaneously from multiple different shooting angles for recording the same operation process; The multi-view operation video data is parsed using a pre-defined video language model to obtain an action unit sequence; Key features are extracted from the action unit sequence to obtain the feature vector corresponding to the action unit sequence. Based on the feature vector, target knowledge data matching the feature vector is retrieved from a preset video data knowledge base. The video language big model is used to fuse and analyze the multi-view operation video data and the target knowledge data to generate a first review report. The first audit report is sent to the human-machine collaborative audit module; the human-machine collaborative audit module is used to perform in-depth audit operations on the first audit report and output the audit results; Receive the audit result returned by the human-machine collaborative audit module, and generate a target audit report based on the audit result.

2. The method according to claim 1, characterized in that, The step involves parsing the multi-view operation video data using a preset video language model to obtain an action unit sequence, including: The video language big model is used to analyze the multi-view operation video data to identify human posture, hand movements, object interactions and scene information in the multi-view operation video data. The identified human posture, hand movements, object interactions, and scene information are summarized to form the content information of the multi-view operation video data; The content information is segmented to generate independent action unit sequences.

3. The method according to claim 1, characterized in that, The method further includes: constructing the video data knowledge base; The construction of the video data knowledge base includes: Obtain objective knowledge data, digitize and structure the objective knowledge data to generate the smallest knowledge unit, and summarize the smallest knowledge unit to obtain an explicit rule base; Subjective experience data from human experts is acquired, the subjective experience data is vectorized to generate soft knowledge units, and the soft knowledge units are summarized to obtain an implicit experience base. The explicit rule base and the implicit experience base are integrated to generate the video data knowledge base.

4. The method according to claim 1, characterized in that, The process of fusing and analyzing the multi-view operation video data and the target knowledge data using the video language big model to generate a first review report includes: The action unit sequence of the multi-view operation video data is matched with the knowledge units of the target knowledge data to generate contextual prompt information; the contextual prompt information includes operation actions, object interactions, scene information, and preset rule matching results; The contextual prompts are input into the video language model, which performs semantic analysis, anomaly detection, and compliance assessment on the action unit sequence based on the contextual information, and generates a first review report.

5. The method according to claim 1, characterized in that, The human-machine collaborative review module is specifically used to: receive the review results of human experts judging each item in the first review report, confirm whether each review result matches the preset conditions, and if not, modify and improve the first review report to generate a second review report. The step of receiving the audit result returned by the human-machine collaborative audit module and generating a target audit report based on the audit result includes: If the audit result matches the preset conditions, the first audit report is set as the target audit report, and the confidence of the multi-view operation video data in knowledge base matching and deep reasoning analysis is improved. If the audit result does not match the preset conditions, the data of the correction operation performed by human experts on the first audit report is obtained, a second audit report is generated based on the data of the correction operation and the first audit report, the second audit report is set as the target audit report, and the data of the correction operation is recorded in the video data knowledge base to realize the content update of the video data knowledge base.

6. The method according to claim 3, characterized in that, The method further includes: The target review report is input into the video data knowledge base to update the video data knowledge base; wherein, updating the video data knowledge base includes: updating the smallest knowledge unit in the explicit rule base and updating the soft knowledge unit in the implicit experience base; The target review report and the updated video data knowledge base are input into the video language big model. Based on the review results in the target review report and the updated video data knowledge base, the video language big model is iteratively learned to adjust the parameters of the video language big model, resulting in an adjusted video language big model. The adjusted video language big model is used to parse the subsequently acquired multi-view operation video data.

7. A video review device, characterized in that, The device includes: The data acquisition module acquires multi-view operation video data; the multi-view operation video data includes video data acquired simultaneously from multiple different shooting angles for recording the same operation process; The data parsing module parses the multi-view operation video data using a preset video language model to obtain an action unit sequence; The data retrieval module extracts key features from the action unit sequence to obtain the feature vector corresponding to the action unit sequence, and retrieves target knowledge data matching the feature vector from a preset video data knowledge base based on the feature vector. The first review module uses the video language big model to fuse and analyze the multi-view operation video data and the target knowledge data to generate a first review report. The second review module sends the first review report to the human-machine collaborative review module; the human-machine collaborative review module is used to perform in-depth review operations on the first review report and output the review results. The report output module receives the audit results returned by the human-machine collaborative audit module and generates a target audit report based on the audit results.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 6.

9. A robot, characterized in that, The robot includes the electronic equipment described in claim 8.

10. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-executable instructions that, when executed by an electronic device, cause the electronic device to perform the method of any one of claims 1 to 6.