Multi-trigger-mode video auditing problem intelligent marking method and device and medium
By employing a multi-trigger intelligent marking method for video review issues, this system automatically acquires and processes video pause information, determines the issue type, and marks it on the progress bar. This solves the problems of cumbersomeness and inaccuracy in existing video marking methods, achieving efficient and intelligent video issue marking and management.
Patent Information
- Application Number
- CN202511736306.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-02-03
AI Technical Summary
Existing video tagging methods suffer from problems such as cumbersome tag creation, poor time accuracy, non-standard information, weak display, lack of management, and limited interaction, making it difficult to meet the high efficiency and standardization requirements of modern video processing.
The marking interface can be configured with multiple trigger methods, supporting pause by clicking, progress bar click, shortcut key, right-click menu and double-click trigger. It automatically obtains pause information and brings up the marking interface, determines the problem type based on image information, uses preset review script content templates, and accurately marks the progress bar.
It has automated and intelligentized video review, improved the efficiency and accuracy of tag creation, enhanced visualization capabilities, and ensured the consistency and accuracy of review results.
Smart Images

Figure CN121462818A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of video content review and marking, and particularly relates to a multi-trigger video review problem intelligent marking method and device and medium. BACKGROUND
[0002] With the rapid development of Internet technology, video content is showing an explosive growth trend. In many scenarios such as video content review, quality control, and teaching evaluation, users need to accurately mark and record the problems found in the video. This is a key link to ensure the quality and compliance of video content and to improve teaching effectiveness. However, the traditional video marking method has been difficult to adapt to the high efficiency and standardization requirements of modern video processing. Therefore, how to conveniently and accurately mark problem points during video playback and provide effective marking management functions has become an important issue to be solved in the field of video review technology.
[0003] Currently, video content review mainly relies on manual review methods, and the reviewer records video problems synchronously while viewing the video. In this way, the reviewer needs to manually pause the video, manually input the time point, and input the problem description to complete the creation of the mark. The mark information is generally displayed in a list form, and lacks perfect mark editing, deletion, recovery and other management functions, and the interactive way only provides a single mark creation approach.
[0004] The existing manual review and marking method has many drawbacks. First, the marking creation process is complex, and the tedious operation steps seriously affect the review efficiency. Second, the time recording precision is insufficient, and the error of the manually recorded time point makes it difficult to achieve accurate problem positioning, and it is not possible to accurately find the problem location when reviewing. Third, the marking information lacks standardization, and the marking format and description method of different reviewers differ greatly, resulting in uneven marking quality. Fourth, the visualization display capability is weak, and the mark list and the video playback interface lack intuitive association, making it difficult for users to clearly understand the mark distribution on the video progress bar and quickly locate the problem point. Fifth, the marking management function is missing, making it difficult to effectively maintain the marks. Sixth, the interactive way is single and rigid, and cannot meet the diversified needs of different use scenarios. SUMMARY
[0005] Therefore, the embodiments of the present disclosure provide a multi-trigger video review problem intelligent marking method, device and medium, which can solve the problems of complex marking creation, poor time precision, non-standard information, weak display, missing management, and single interaction in the prior art.
[0006] In a first aspect, the embodiments of the present disclosure provide a multi-trigger video review problem intelligent marking method, comprising: configuring a marking interface and configuring different types of review speech content templates in the marking interface; mapping relationship between the trigger types of different trigger video pause and the mark interface is configured, and is stored to a preset database; The trigger types are click pause video trigger, progress bar click trigger, shortcut key trigger, right-click menu trigger, or double-click trigger. In response to a video pause trigger signal, corresponding pause information and a current trigger type of trigger video pause are acquired, and the mark interface is automatically called out; The pause information includes a current timestamp, a formatted label, a unique identifier, and image information corresponding to a current pause time. A question type is determined based on the image information in the pause information. A corresponding review script content template is determined in the mark interface according to the question type, and is recorded as target content. The pause information and the target content are marked in a preset manner at a corresponding time point of a progress bar according to the current timestamp in the pause information.
[0007] In a second aspect, the embodiments of the present disclosure further provide a computer device, which adopts the following technical solution: The computer device comprises: at least one processor; and a memory in communication connection with the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the multi-trigger mode video review question intelligent marking method described in any one of the above.
[0008] In a third aspect, the embodiments of the present disclosure further provide a computer readable storage medium, which stores computer instructions for causing a computer to execute the multi-trigger mode video review question intelligent marking method described in any one of the above.
[0009] In a fourth aspect, the embodiments of the present disclosure further provide a computer program product, which comprises computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the steps of the method described in any one of the above.
[0010] The multi-trigger mode video review problem intelligent marking method disclosed in the application can quickly obtain corresponding pause information and the current trigger type of triggering video pause when the video pause trigger signal occurs, and automatically call out the marking interface; then determine the problem type based on the image information in the pause information, determine the corresponding review speech content template in the marking interface according to the problem type, and mark it as the target content; finally, mark the pause information and the target content according to the preset mode at the corresponding time point of the progress bar according to the current timestamp in the pause information, realize accurate, efficient and intelligent automatic marking, and can be effectively displayed.
[0011] The above description is only a summary of the technical solutions of the present disclosure. In order to more clearly understand the technical means of the present disclosure, the content of the specification can be implemented, and in order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the following preferred embodiments are described in detail below, and the accompanying drawings are described as follows. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0013] Figure 1 The flowchart of the multi-trigger mode video review problem intelligent marking method provided by the embodiments of the present disclosure.
[0014] Figure 2 The flowchart of the method for determining the problem type based on the image information in the pause information provided by the embodiments of the present disclosure.
[0015] Figure 3 The flowchart of the method for obtaining the target content provided by the embodiments of the present disclosure.
[0016] Figure 4 The structural schematic diagram of a computer device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION
[0017] The embodiments of the present disclosure will be described in detail below with reference to the drawings.
[0018] It should be apparent that the following description illustrates by way of example only a number of possible embodiments of the present disclosure. Those skilled in the art will readily understand other advantages and benefits of the present disclosure from the description that follows, without departing from the scope of the present disclosure. Obviously, the described embodiments are merely some, but not all, of the embodiments of the present disclosure. The present disclosure can also be implemented or applied in other different specific embodiments, and the details in the description can be modified or changed based on different views and applications, without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present disclosure.
[0019] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It should be apparent to those skilled in the art that the aspects described herein can be embodied in a wide variety of forms, and that any specific structure and / or function described herein is merely illustrative. Based on the teachings herein one skilled in the art should appreciate that an aspect described herein can be implemented independently of any other aspects and that an aspect can be implemented both as any number of software, firmware, and / or hardware structures.
[0020] It should also be noted that the figures provided in the following embodiments are only schematically illustrating the basic concepts of the present disclosure, and only the components related to the present disclosure are shown in the figures, not drawn according to the number, shape and size of the components in actual implementation, and the actual implementation of each component can be a random change in shape, number and proportion, and the layout of the components can be more complex.
[0021] In addition, in the following description, specific details are provided in order to facilitate a thorough understanding of the examples. However, one skilled in the art will understand that the aspects described can be practiced without these specific details.
[0022] Referring to Figure 1 The present application discloses a multi-trigger video review problem intelligent marking method, which can realize intelligent creation of video problem marking based on multiple user interaction modes. The method comprises: S100, configuring a marking interface and configuring different types of review script content templates in the marking interface.
[0023] Through this step, different marking needs can be effectively met.
[0024] S200, configure a mapping relationship between trigger types of different video pause triggers and a marking interface, and store to a preset database.
[0025] The trigger type is a click pause video trigger, a progress bar click trigger, a shortcut key trigger, a right-click menu trigger, or a double-click trigger.
[0026] The method configures a mapping relationship between trigger types of different video pause triggers and a marking interface. This multi-trigger mode effectively improves the flexibility and convenience of user operations. Reviewers can freely choose the most suitable trigger mode to start the marking function according to their habits and actual operation scenarios, without having to manually pause the video, manually input the time point, and other series of cumbersome operations as in the traditional mode.
[0027] S300, in response to a video pause trigger signal, obtaining corresponding pause information and the current trigger type of the video pause trigger, and automatically calling out a marking interface.
[0028] The pause information includes a current timestamp, a formatted label, a unique identifier, and image information corresponding to the current pause time.
[0029] When the video pause trigger is triggered, the system will automatically execute this step to automatically call out the marking interface, greatly simplifying the marking creation process and improving the review efficiency.
[0030] S400, determining a problem type based on the image information in the pause information.
[0031] S500, determining a corresponding review script content template in the marking interface according to the problem type, denoted as a target content.
[0032] Through the above configuration of different types of review script content templates, in this step, the corresponding review script content template can be directly determined as the target content from the marking interface according to the problem type, without the need for the user to manually input the problem description, further reducing the operation steps and improving the marking creation speed.
[0033] S600, marking the pause information and the target content at the corresponding time point of the progress bar according to the current timestamp in the pause information in a preset manner.
[0034] By automatically obtaining the current timestamp as part of the pause information when the video is paused, the current timestamp in the pause information is used to mark the pause information and the target content at the corresponding time point on the progress bar in a preset manner, ensuring that the mark corresponds to the time point on the video progress bar, and further improving the accuracy of problem positioning. Further, the pause information and the target content are marked at the corresponding time point on the progress bar in a preset manner. The user can clearly see the distribution of the marks on the video progress bar, quickly locate the problem point, enhance the intuitive association between the mark information and the video playback interface, and improve the visual display capability.
[0035] The multi-trigger video review problem intelligent marking method disclosed in the present application forms a complete video review problem intelligent marking process from configuring the marking interface, triggering the video pause, obtaining the pause information, determining the problem type, matching the review script to marking on the progress bar, which can realize the automation and intelligentization of the video review process, greatly improve the review efficiency and quality consistency, ensure the consistency and accuracy of the review results, and improve the review quality.
[0036] The method has multiple triggering methods, automatic information acquisition, automatic callout of the marking interface, and preset review script content templates, effectively reducing the operation steps of the reviewer and the time of writing the review script, making the review process more efficient; the problem type is determined by image recognition technology, the standard review script content template is used, and the mark is accurately marked on the progress bar, effectively improving the accuracy of problem identification and review opinion expression; the mapping relationship, pause information and review opinion are stored in the database, which is convenient for the system to manage and maintain, and also facilitates the reviewer to trace and query the review history.
[0037] For S100 "configure the marking interface and configure different types of review script content templates in the marking interface", assuming that the video to be reviewed is mainly news video, the review script content templates can be configured according to common problem types; for example, for the problem of false information, the configured review script content template can be "there is false information in the [specific description event] part of the video, which needs to be verified and modified"; for the problem of sensitive content, the template can be "sensitive content appears in [specific location of sensitive content in the video], which is suggested to be deleted". When designing the marking interface, a simple and intuitive layout can be used to display different types of review script content templates in the form of a list, which facilitates the reviewer to quickly find and select, and the configured marking interface provides a unified operation platform for the reviewer, making the review process more standardized and orderly. Different types of review script content templates can help the reviewer to quickly and accurately express review opinions and improve review efficiency.
[0038] For S200, "configure the mapping relationship between different trigger types of video pause and the marking interface, and store it in the preset database", specifically, for the click pause trigger, when the reviewer clicks the pause button of the video player, the system associates the trigger type with the marking interface, that is, the marking interface is automatically called out after clicking pause; for the progress bar click trigger, the reviewer clicks a certain time point on the progress bar, the video is paused and the marking interface is called out at the same time; for the shortcut key trigger, for example, set "Ctrl + P" as the shortcut key for pausing and calling out the marking interface, when the reviewer presses the combination key, the corresponding function is realized; for the right-click menu trigger, click the right mouse button on the video player, set the "pause and mark" option in the pop-up menu, and after clicking the option, the video is paused and the marking interface is called out; for the double-click trigger, the reviewer double-clicks on the video screen, the video is paused and the marking interface is called out. Store these mapping relationships in the preset database for subsequent system query and call. Multiple trigger methods provide more operation options for reviewers, meet the use habits of different reviewers, improve the convenience of operation, store the mapping relationship in the database, facilitate system management and maintenance, and also ensure the consistency and stability of data.
[0039] For the method of "obtaining corresponding pause information in response to a video pause trigger signal" in S300, specifically comprising: A100, listen to multiple pause events of the video player, when detecting a pause operation, immediately capture the current playing state and generate an initial pause information package containing the original timestamp, player state code and trigger source identification.
[0040] Specifically, the pause and play state changes of the video player are captured through an event listening mechanism, which is a common programming technique for monitoring the occurrence of specific events. In the video player, the pause event indicates that the video is paused, and the play event indicates that the video starts playing. By listening to these two events, the change of the video playing state can be perceived in real time. For example, when the user clicks the pause button of the video player, the pause event is triggered; clicking the play button triggers the play event.
[0041] The setPauseInfo() method is called when a pause event is detected to accurately capture the play timestamp at the time of pausing. When the video player triggers a pause event, the module calls the setPauseInfo() method, which is primarily responsible for obtaining the precise play timestamp at the moment the video is paused. The timestamp is a numerical value representing a specific time point, usually measured in milliseconds, which accurately records the moment of video pausing. After obtaining the timestamp, it is automatically processed to calculate the corresponding minute (M) and second (S) parts, and formatted into the standard time label MM:SS. For example, if the timestamp corresponds to 1 minute and 30 seconds, the formatted time label is 01:30, which is more in line with people's reading habits, making it easier for auditors to intuitively understand the time point of video pausing.
[0042] A200, analyze the initial pause information package through the event source analyzer, identify the trigger type, and generate a unique type code.
[0043] Specifically, a mapping table of trigger types and codes is established, such as "button_pause" corresponding to code "T001", "seek_pause" (seek bar drag causing pause) corresponding to code "T002". The event source analyzer matches the trigger source identifier in the initial information package to generate the corresponding type code. Various trigger types are represented by unique codes, facilitating subsequent data storage, query, and analysis, making data processing more standardized. When performing large-scale data statistics, using codes can quickly classify and summarize different trigger types, improving analysis efficiency.
[0044] A300, feature extraction of images at the time of video pausing, generating image feature package.
[0045] Specifically, the SIFT (Scale-Invariant Feature Transform) algorithm is used to detect key points in the image and calculate feature descriptors around the key points, combining these descriptors into an image feature package.
[0046] A400, fusion processing of image feature package and type code to generate formatted label and globally unique identifier, combining current timestamp and image information corresponding to the current pause moment to construct structured pause information.
[0047] Specifically, the image feature bag and type code are spliced, and then dimension reduction processing is performed through a fully connected layer to obtain a fused feature vector; according to the fused feature vector, a formatted label such as "user actively pauses at highlight frame" is generated through a preset rule or classifier; a globally unique identifier is generated using a UUID algorithm to ensure that each pause information has a unique identifier; the formatted label, globally unique identifier, current timestamp, and image information at the pause moment are combined together to form a structured pause information object; the formatted label converts complex feature information into an easily understandable text description, facilitating manual viewing and analysis.
[0048] In this embodiment, the algorithm responding to the pause event mainly includes: 1) register a video player pause event listener: in the program, a pause event listener is registered first, so that when the video player triggers a pause event, the program can receive the notification and perform corresponding operations. 2) execute setPauseInfo() when the pause event is triggered: when the pause event is triggered, the program will call the setPauseInfo() method to start capturing and processing the relevant information at the pause moment. 3) get accurate currentTime timestamp (millisecond level): the setPauseInfo() method will get the currentTime attribute of the video player, which represents the current playing time of the video in milliseconds, thereby obtaining the accurate pause timestamp. 4) generate a formatted time label MM:SS by calculating the minute and second part of the time: according to the obtained timestamp, a specific calculation method is used to separate the minute and second part and format it into a MM:SS time label. 5) create pause information object pauseInfo: according to the structure mentioned above, create pauseInfo object containing accurate timestamp, formatted time label and unique identifier field. 6) trigger the display of the mark creation interface (provide standardized review scripts options for reviewers to select whether the current video point has a problem): after completing the capture and processing of the pause information, the module will trigger the display of the mark creation interface, which will provide standardized review script options, and the reviewer can quickly judge whether the current video pause point has a problem according to these options and perform corresponding marking operations.
[0049] For "automatic callout interface" in S300, specifically includes: 1) when triggered by clicking pause video: listen to the video `pause` event, automatically pop up the mark creation interface; 2) when triggered by clicking the progress bar: listen to the progress bar `click` event, create a mark at the click position; 3) when triggered by shortcut key: listen to the `keydown` event, support custom shortcut key combination to create a mark; 4) when triggered by right-click menu: implement custom `contextmenu` in the video area, provide the "add mark" option; 5) when triggered by double-click: listen to the video area `dblclick` event, quickly create a mark.
[0050] Referring to Figure 2 For the method of S400 "determining the question type based on the image information in the pause information", it includes: S410, based on the image information of the current pause time, extract the image feature vector and perform normalization processing.
[0051] Specifically, the image information of the current pause time is preprocessed: first, the original image is standardized to 512x512 pixels to ensure consistency in subsequent processing, then the image is smoothed using Gaussian filtering to remove random noise in the image, making the image clearer, and finally, the contrast of the image is enhanced through histogram equalization to highlight the details of the image. The objects and scenes in the image are more clear; use multi-scale convolution kernel to convolve the preprocessed image, extract the edge, texture and color distribution of the image, etc. The bottom features usually have high dimensions, in order to reduce the amount of calculation and storage cost, use PCA (Principal Component Analysis) dimensionality reduction technology to compress high-dimensional features into 256-dimensional feature vectors; use Z-score standardization method to normalize the 256-dimensional feature vector, so that each element in the feature vector has zero mean and unit variance, eliminating the dimensional differences between different features.
[0052] S420, input the preprocessed image feature vector into the pre-trained multi-classification neural network model, and output the confidence distribution of each question type.
[0053] Specifically, a pre-trained ResNet-50 improved multi-classification neural network model is used, which adds a feature enhancement layer, an attention mechanism layer and a multi-head classifier to the traditional ResNet-50. The feature enhancement layer can further extract and enhance the feature information of the image, the attention mechanism layer can make the model pay more attention to the important areas in the image, and the multi-head classifier can classify multiple question types at the same time. A 256-dimensional standardized image feature vector is input into the model for forward propagation. After a series of calculations by the model, a 12-dimensional confidence probability distribution vector is output by the Softmax activation function, each dimension corresponding to a question type, and the value representing the confidence of the question type, ranging from 0 to 1. The pre-trained ResNet-50 model has been trained on a large-scale image dataset and has strong feature extraction and classification capabilities. The improved model architecture (feature enhancement layer, attention mechanism layer and multi-head classifier) can further improve the classification accuracy of the model for different question types. The multi-head classifier can classify multiple question types at the same time, output the confidence distribution of all question types at once, and improve the classification efficiency.
[0054] In S430, the corresponding question type is determined according to the confidence distribution and a preset threshold strategy.
[0055] Specifically, the main threshold 0.7 is set for high-confidence determination, and the secondary threshold 0.4 is set for suspicious question marking. The 12-dimensional confidence distribution vector output in the previous step is received, and the question type with the highest confidence is found. If the highest confidence is higher than 0.7, the question type is directly determined. If the highest confidence is between 0.4 and 0.7, the question is marked as “to be reviewed”. If all confidences are lower than 0.4, it is marked as “normal”. Finally, a structured question type label object containing question type label, confidence value and determination basis is output, which is convenient for subsequent processing and analysis. The dynamic threshold strategy can make different determinations according to the confidence level, avoiding the misjudgment problem caused by a single threshold. For high-confidence question types, the question type can be directly determined. For suspicious questions, further review is needed, which improves the accuracy of determination. The structured question type label object contains information such as question type, confidence and determination basis, which is convenient for subsequent manual review, data analysis and processing.
[0056] Reference Figure 3 For S500, “determine the corresponding review speech content template in the marking interface according to the question type, denoted as target content”, which is the method for obtaining the target content, specifically including: In S510, the question type label is input to retrieve the associated basic speech template from the pre-constructed speech knowledge graph.
[0057] Specifically, assume that the problem type label is "picture quality abnormality". The pre-constructed dialogue knowledge graph is a knowledge base containing various problem types and corresponding basic dialogue templates. In this knowledge graph, "picture quality abnormality" is associated with basic dialogue templates such as "the video has picture quality problems, which may affect the viewing experience" and "the video picture quality is abnormal, please check and handle". The system will search for these basic dialogue templates associated with the "picture quality abnormality" label in the knowledge graph according to the input. Through the pre-constructed knowledge graph, the basic dialogue template related to the problem type can be quickly located, avoiding the need to rewrite the dialogue every time and saving time and labor costs. The basic dialogue templates in the knowledge graph are carefully designed and audited, and the use of these templates can ensure that the audit dialogue is consistent in different audit personnel and audit scenarios.
[0058] S520, based on the image feature vector, obtaining a relevance score of the basic dialogue template and the current image content through a semantic matching algorithm.
[0059] Specifically, for the previously retrieved basic dialogue template "the video has picture quality problems, which may affect the viewing experience", semantic matching is performed between it and the feature vector of the current "picture quality abnormality" image. For example, the image feature vector reflects that the image has features such as blurring and color distortion, and the semantic matching algorithm analyzes the relevance of "picture quality problems" in the dialogue template to these specific image features. If the dialogue template can well cover the picture quality problem features in the image, a higher relevance score will be given; otherwise, if the dialogue template is too broad and not closely related to the specific image features, the score will be lower. This step can measure the fit of the basic dialogue template and the current image content, ensure that the selected dialogue is more in line with the actual problem, improve the relevance of the audit dialogue, and through the relevance score, the basic dialogue template most relevant to the image content can be selected to provide a better basis for generating customized dialogue.
[0060] S530, according to the relevance score, individualizing the basic dialogue template to generate a customized audit dialogue containing specific descriptions.
[0061] Specifically, assume that the relevance score of the previously retrieved basic dialogue template "the video has picture quality problems, which may affect the viewing experience" is high, and according to the specific situation reflected by the image feature vector, individualize it. If the image features show that the picture is blurred and the color is dark, the customized audit dialogue can be adjusted to "the video has picture quality problems, the picture is blurred and the color is dark, which seriously affects the viewing experience". In this step, the customized audit dialogue contains specific problem descriptions, which can help the audited party better understand the problem and improve the effectiveness of communication.
[0062] S540, the customized review script is prioritized according to the severity and urgency, and a structured script content list is formed.
[0063] Specifically, assuming that there are multiple customized review scripts, each of which is for different problem types and image situations. For the customized script of the "content violation" type, due to its high severity and urgency, it will be placed at the front of the list; while for the script of the "subtitle error" type, the severity and urgency are relatively low, and it will be placed at the back of the list, finally forming a structured script content list according to the priority, which is convenient for review personnel to view and use. Review personnel can process review tasks in order according to the priority order, improve the review efficiency, ensure that important problems are handled in time, and the structured script content list enables review personnel to quickly understand the priority and general situation of different problems, facilitating the overall grasp of the review work.
[0064] In another embodiment, the method for obtaining target content can also include: taking the problem type label as input, retrieving the basic script template with the highest content similarity from the pre-built script knowledge graph as the target content, without modification when tagging, directly calling and automatically annotating, that is, after determining the problem type, automatically matching the most suitable review script, ensuring that review personnel use uniform script when handling the same type of problem, thereby improving the standardization of the review and avoiding the problem of inconsistent review results due to subjective factors of the review personnel.
[0065] For the method of S600 "tagging the pause information and the target content according to the corresponding time point of the progress bar at the current timestamp in the pause information in a preset manner", it includes: S610, calculating the accurate pixel position on the video progress bar based on the current timestamp, and selecting the corresponding visual marker style according to the problem severity.
[0066] Assuming that the total length of the current video is 60 minutes, and the length of the video progress bar is 600 pixels. The current timestamp is 15 minutes, so through simple proportional calculation (15 / 60 600 = 150 pixels), the accurate pixel position on the progress bar is 150 pixels; for the problem severity, if the problem type is "content violation", it is a serious problem, and a prominent red triangle can be selected as the visual marker style; if the problem type is "subtitle error", the severity is relatively low, and a yellow circle can be selected as the visual marker style; by calculating the accurate pixel position, the time point of the problem can be accurately marked on the progress bar, which is convenient for review personnel to quickly locate the problem; according to the problem severity, different visual marker styles can be selected, which can enable the review personnel to see the severity of the problem at a glance, improving the review efficiency.
[0067] S620, create a multi-level mark point including the basic time mark layer, the problem type identification layer and the dialogue preview layer at the corresponding position of the progress bar, form an interactive mark node, and encapsulate the pause information, the problem type, the target content and the visual mark style into a standardized mark data package, and write into the video metadata storage area.
[0068] A multi-level mark point is created at the 150 pixel of the progress bar. The basic time mark layer displays the current timestamp "15 minutes"; the problem type identification layer displays the problem type with an icon or text, such as "content violation"; and the dialogue preview layer displays the previously generated target content, such as "the video contains serious content violation, please handle immediately". The pause information (such as image features at the pause moment), the problem type ("content violation"), the target content and the visual mark style (red triangle) are encapsulated into a standardized mark data package, which is then written into the video metadata storage area for subsequent query and use. The multi-level mark point contains time, problem type and review dialogue, etc. information, which can help the reviewer understand the problem comprehensively. The interactive mark node formed by the multi-level mark point is convenient for the reviewer to click and view detailed information, improving the convenience of operation. The related information is encapsulated into a standardized mark data package and stored in the video metadata storage area, which is convenient for unified management and maintenance of the mark information.
[0069] Further, the global mark index table is updated to establish a bidirectional mapping relationship between the timestamp and the mark data package, and a interface refresh event is triggered to output complete mark state information. The global mark index table is a table recording all mark information. After the mark is completed, a record is added in the index table to establish a bidirectional mapping relationship between the timestamp (15 minutes) and the mark data package, that is, the corresponding mark data package can be quickly found through the timestamp, and vice versa. The interface refresh event is triggered to update the progress bar and mark information on the interface in time, output complete mark state information, and let the reviewer see the latest mark situation. The bidirectional mapping relationship facilitates quick search and positioning of mark information at a specific time point, improves the data query efficiency, and the interface refresh event ensures that the reviewer can see the latest mark state in real time, avoiding misoperation caused by lagging information.
[0070] The multi-trigger video review problem intelligent marking method disclosed in the application further comprises: acquiring all the marking information, and judging whether the adjacent marking is within a preset period; If the adjacent marking is within the preset period, the marking within the same preset period is merged from front to back according to the video playing order to generate a comprehensive marking and display.
[0071] Further, the marking in the embodiment can be marked with a green bubble on the progress bar at the current time point, and when the auditor does not select the review script content or cancels the selected review script content, the marking data and the bubble are lost or cleared.
[0072] Further, the multi-trigger video review question intelligent marking method disclosed in the application further comprises: the reviewer judges whether the marked problem pause picture matches the marked corresponding script content, that is, checks whether there is a marked error, if not, reacquires the corresponding review script content and corrects the marking, or cancels the marking when the marked problem pause picture does not belong to the problem picture. Further, the application supports the functions of floating prompt, highlighting, and click interaction of the marked.
[0073] The application further comprises: verifying that the video is in a pause state, checking whether the user has selected a review script, obtaining time-based information from pauseInfo, recording the current marking time point, checking whether pauseInfo has a unique identifier identity, calling guid() to generate a new identifier if not, and then checking whether the identifier record exists in the custom marking record list customTagRecord. If found, only modify the problem description of the existing record, if not found, it represents a new record point data, then add pauseInfo information to the custom marking record list customTagRecord storage array, then clear the pauseInfo object, then sort customTagRecord in time sequence, and then call the mergePreCheckByTime method as a parameter to perform intelligent merging of time ranges. After merging, trigger the visual update process.
[0074] The application further comprises real-time marking information display, which can listen to the current time to display related marking information during video playback, check whether there is a mark at the current time point by using the checkPreCheck method, dynamically display and hide the marking detail panel, and support real-time editing and management of marking information. Specifically, the video timeupdate event is listened to, the current playback time currentTime is obtained, customCheckList is traversed to find a matching mark: the time difference is calculated: currentTime - mark time, it is judged whether it is within the display window (usually within 3 seconds), if there is a matching mark, the curCustomCheckIndex points to the current mark, the marking detail panel is displayed (all marking record data under the current mark point are displayed, which can be modified, deleted, closed, etc.), and the corresponding bubble mark on the highlight progress bar is displayed; if there is no matching mark, the curCustomCheckIndex is reset to -1, the marking detail panel is hidden, and the marking highlight state is canceled.
[0075] Further, the application also includes: based on all valid markers to automatically generate a structured rejection reason document: according to all valid marker information, a structured rejection reason document is automatically generated to provide clear and standardized rejection basis for the reviewer; the marker information is arranged in chronological order to make the content in the rejection reason document present a clear time context, facilitating viewing and understanding; a unified standardized template is used to format the rejection content to ensure that the generated document format is standardized and unified; the user can preview the generated rejection reason document and edit it to meet different needs.
[0076] The computer device according to the embodiments of the present disclosure includes a memory and a processor. The memory is configured to store non-transitory computer readable instructions. Specifically, the memory can include one or more computer program products, which can include various forms of computer readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory, etc. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.
[0077] The processor can be a central processing unit (CPU) or other forms of processing units with data processing and / or instruction execution capabilities, and can control other components in the computer device to perform desired functions. In one embodiment of the present disclosure, the processor is configured to execute the computer readable instructions stored in the memory, so that the computer device performs all or part of the steps of the multi-trigger mode video review problem intelligent marking method according to the embodiments of the present disclosure.
[0078] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain a good user experience effect, the present embodiment can also include well-known structures such as communication bus, interface, etc., which should also be included in the protection scope of the present disclosure.
[0079] As Figure 4 A structural schematic diagram of a computer device according to an embodiment of the present disclosure is shown. It shows a structural schematic diagram suitable for implementing the computer device in the embodiments of the present disclosure. Figure 4 The computer device shown is only an example and should not limit the functions and use range of the embodiments of the present disclosure.
[0080] As Figure 4As shown, the computer device can include a processor (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) or loaded from a storage device into a random access memory (RAM). Various programs and data required for the operation of the computer device are also stored in the RAM. The processor, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0081] Generally, the following devices can be connected to the I / O interface: input devices including, for example, sensors or visual information collection devices; output devices including, for example, display screens; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. The communication devices can allow the computer device to communicate wirelessly or wired with other devices (such as edge computing devices) to exchange data. Although Figure 4 The computer device is shown with various devices, but it should be understood that not all of the shown devices are required to be implemented or possessed. More or fewer devices can alternatively be implemented or possessed.
[0082] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device, or installed from the storage device, or installed from the ROM. When the computer program is executed by the processor, all or part of the steps of the multi-trigger mode video review problem intelligent labeling method of embodiments of the present disclosure are performed.
[0083] Detailed descriptions of the present embodiments can refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.
[0084] The computer-readable storage medium according to the embodiments of the present disclosure has non-transitory computer-readable instructions stored thereon. When the non-transitory computer-readable instructions are run by a processor, all or part of the steps of the multi-trigger mode video review problem intelligent labeling method of the embodiments of the present disclosure described above are performed.
[0085] The computer-readable storage medium described above includes, but is not limited to, optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or mobile hard disk), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).
[0086] For details of the present embodiment, reference can be made to the corresponding description in the foregoing embodiments, which will not be repeated here.
[0087] The above has described the basic principles of the present disclosure in combination with specific embodiments, but it needs to be pointed out that the advantages, benefits, effects and the like mentioned in the present disclosure are only examples and not limitations, and these advantages, benefits, effects and the like cannot be considered as necessary for each embodiment of the present disclosure. In addition, the above specific details of the disclosure are only for the purpose of example and for the purpose of understanding, and not for limitation, and the above details do not limit the present disclosure to be necessarily implemented with the above specific details.
[0088] In the present disclosure, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations, and the block diagrams of the devices, apparatuses, equipment, systems involved in the present disclosure are only illustrative examples and are not intended to require or imply the connection, arrangement, configuration as shown in the block diagram. As those skilled in the art will recognize, these devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner. Words such as "include", "contain", "have" and the like are open-ended words, which mean "including but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.
[0089] In addition, as used herein, "or" used in a list of items, starting with "at least one of", indicates a disjunctive list such that, for example, "at least one of A, B, or C" means A or B or C or AB or AC or BC or ABC (i.e., A and B and C). In addition, the phrase "exemplary" does not mean that a described example is preferred or better than other examples.
[0090] It also needs to be pointed out that in the systems and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions of the present disclosure.
[0091] Various changes, modifications, and alterations to the techniques described herein can be made without departing from the teachings of the attached claims. Moreover, the scope of the claims should not be limited to the particular aspects described herein, but should be given the broadest interpretation available to them under the law. All patents, patent applications, and publications identified are expressly incorporated herein by reference for the purpose of describing and disclosing, for example, the methodologies described in such publications that might be used in connection with the technology described herein. These publications are provided solely for their disclosure prior to the filing date of the present application. Nothing in this regard should be construed as a representation by the inventor and / or the assignee that the inventors and / or the assignee has made or maintains any dedication to the public of the patentable matter in the publications other than the inventor and / or assignee's own intellectual property. No admission is made that any portion of the patent literature can be prior art. The claims should not be limited to the specific aspects and embodiments described herein but should be given the broadest interpretation available to them under the law.
[0092] The above description of disclosed aspects is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0093] The above description has been presented for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the disclosure to the forms disclosed herein. Although various example aspects and embodiments have been discussed above, those of skill in the art will recognize certain modifications, permutations, additions, and sub-combinations thereof.
Claims
1. A method for intelligently marking video review issues using multiple triggering methods, characterized in that, include: Configure the tagging interface and configure different types of audit script content templates in the tagging interface; Configure the mapping relationship between different trigger types for pausing video and the marking interface, and store it in a preset database; The trigger types are: click to pause video, click on progress bar, shortcut key, right-click menu, or double-click. In response to a video pause trigger signal, the corresponding pause information and the current trigger type that triggered the video pause are obtained, and the marking interface is automatically brought up; The pause information includes the current timestamp, format label, unique identifier, and image information corresponding to the current pause time; The problem type is determined based on the image information in the pause information; Based on the question type, determine the corresponding review script template in the marking interface and record it as the target content; Based on the current timestamp in the pause information, the pause information and the target content are marked at the corresponding time point on the progress bar according to a preset method.
2. The intelligent tagging method for video review issues with multiple triggering methods according to claim 1, characterized in that, The step of responding to a video pause trigger signal and obtaining the corresponding pause information includes: Monitor various pause events of the video player. When a pause operation is detected, immediately capture the current playback state and generate an initial pause information packet containing the original timestamp, player status code, and trigger source identifier. The initial pause information packet is analyzed by the event source analyzer to identify the trigger type and generate a unique type code; Extract features from images when the video is paused to generate an image feature package; The image feature package is fused with the type encoding to generate a formatted label and a globally unique identifier. Combined with the current timestamp and the image information corresponding to the current pause time, structured pause information is constructed.
3. The intelligent marking method for video review issues with multiple triggering methods according to claim 1, characterized in that, The process of determining the problem type based on the image information in the pause information includes: Image feature vectors are extracted based on the image information at the current pause time and then normalized. The preprocessed image feature vector is input into a pre-trained multi-class neural network model, which outputs the confidence distribution of each problem type. Based on the confidence distribution and the preset threshold strategy, the corresponding problem type is determined.
4. The intelligent marking method for video review issues with multiple triggering methods according to claim 3, characterized in that, The step of determining the corresponding review script content template in the marking interface according to the question type, denoted as the target content, includes: Using the question type label as input, retrieve the associated basic dialogue templates from the pre-built dialogue knowledge graph; Based on the image feature vector, a semantic matching algorithm is used to obtain the relevance score between the basic dialogue template and the current image content; The basic script template is personalized based on the relevance score to generate a customized audit script containing specific descriptions; The customized review scripts are prioritized according to their severity and urgency to form a structured list of script content.
5. The intelligent marking method for video review issues with multiple triggering methods according to claim 3, characterized in that, The step of determining the corresponding review script content template in the marking interface according to the question type, denoted as the target content, includes: Using the question type label as input, the basic dialogue template with the highest content similarity is retrieved from the pre-built dialogue knowledge graph and used as the target content.
6. The intelligent tagging method for video review issues with multiple triggering methods according to claim 1, characterized in that, The step of marking the pause information and the target content according to a preset method at the corresponding time point of the progress bar based on the current timestamp in the pause information includes: Calculate the precise pixel position on the video progress bar based on the current timestamp, and select the corresponding visual marker style according to the severity of the problem; Multi-level marker points, including a basic time marker layer, a question type identifier layer, and a script preview layer, are created at the corresponding positions of the progress bar to form interactive marker nodes. The pause information, the question type, the target content, and the visual marker style are encapsulated into standardized marker data packages and written to the video metadata storage area.
7. The intelligent marking method for video review issues with multiple triggering methods according to claim 6, characterized in that, Also includes: Obtain all marking information and determine whether several adjacent markings are within a preset period; If several adjacent tags are within a preset period, they will be merged according to the video playback order from front to back to generate a single composite tag and display it.
8. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the multi-trigger video review issue intelligent tagging method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to execute the multi-trigger video review issue smart tagging method as described in any one of claims 1-7.
10. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the method according to any one of claims 1-7.