Method and apparatus for analyzing and controlling output video content
By using a video input unit to transmit frames to an analysis unit with a frame buffer for real-time AI-based content recognition, the method addresses latency issues in existing systems, enabling efficient and timely management of video content.
Patent Information
- Application Number
- CN202510615302.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-14
AI Technical Summary
In the prior art, the video content monitoring system has problems such as high transmission delay and long processing delay, which is difficult to meet the real-time monitoring needs of live broadcast programs, and relies on manual review to be inefficient and prone to errors.
Video frames are received through the video input unit and transmitted to the analysis unit and control unit at the same time. The artificial intelligence model is used to extract features and identify contents, generate analysis results, and set the picture status flags of the video frame based on the analysis results to realize timely output and flexible control of the video frame.
It realizes timely monitoring and management of video content, avoids missed inspection, improves identification accuracy and response efficiency, and supports dynamic adjustment of parameters to adapt to monitoring needs of different scenarios.
Smart Images

Figure CN120166199B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method and device for analyzing and controlling output video content. Background Art
[0002] With the increase in the sources and scale of video content, such as short videos, long videos, and various programs with live or quasi-live nature, it has brought challenges to the intelligent analysis and supervision of video content. If relying on manual methods to review and filter sensitive or abnormal content in video content, it will result in low efficiency and easy errors. In the prior art, a centralized video monitoring system is generally adopted, where a large amount of video content is centralized on the background server, and then software programs are used to monitor and filter the video content. It can also be combined with manual supervision. However, such a centralized video monitoring system means increased transmission delay and processing delay. Therefore, it is difficult to meet the requirements of live programs such as large online conferences and live broadcasts of sports and cultural festivals. Another solution in the prior art is a distributed video monitoring system, which deploys a video monitoring system at the receiving end, such as the user side. However, the distributed video monitoring system in the prior art needs to perform content analysis after video capture, and needs to transmit the signal to the memory for storage and then scan the stored video content, resulting in high processing delay and unable to achieve timely monitoring and timely response.
[0003] Therefore, this application provides a method and device for analyzing and controlling output video content to address the technical problems in the prior art. Summary of the Invention
[0004] In a first aspect, this application provides a method for analyzing and controlling output video content. The method includes: through a video input unit, simultaneously transmitting a plurality of video frames received by the video input unit within a first time period to an analysis unit and a control unit; through the control unit, storing the plurality of video frames into a frame buffer area in the control unit, and then, based on the analysis result, uniformly setting the picture state flag bits of the plurality of video frames respectively. Then, based on the picture state flag bit of the video frame located at the output position in the frame buffer area and the picture state flag bits of at least the previous one video frame located at the output position in the frame buffer area, output the video frame located at the output position, or output a preset picture, where the analysis result is generated by the analysis unit extracting features and identifying the content of the plurality of video frames, the length of the first time period is a preset video output delay time, and the plurality of video frames respectively, starting from being received by the video input unit, after the preset video output delay time, serve as the video frame located at the output position.
[0005] Through the first aspect of the present application, a plurality of video frames received by the video input unit within a first time period are simultaneously transmitted to the analysis unit and the control unit. In this way, the video frames move along the data frame movement direction in the frame buffer towards the output bit of the frame buffer, that is, the video frame becomes the video frame at the output bit after a preset video output delay time. And synchronously in the analysis unit, feature extraction and content recognition are performed on the plurality of video frames to generate an analysis result. In this way, the artificial intelligence model efficiently utilizes the preset video output delay time, so that an analysis result can be generated based on the detection frames of the artificial intelligence model, and then the picture state flag bits of these video frames can be uniformly set. By using the picture state flag bit of the video frame at the output bit in the frame buffer and the picture state flag bits of the video frames at at least the previous bit before the output bit in the frame buffer, content recognition with better timeliness is achieved, which is beneficial to avoiding missed detection situations. A flexible control method capable of dynamically adjusting various parameters is realized, and more effective and timely responsive content management and monitoring are achieved.
[0006] In a possible implementation manner of the first aspect of the present application, through the control unit, based on the analysis result, the picture state flag bits of the plurality of video frames are uniformly set, including: through the control unit, based on the analysis result, the picture state flag bit of each video frame in the plurality of video frames is set to indicate the existence of abnormal content or indicate the non-existence of abnormal content.
[0007] In a possible implementation manner of the first aspect of the present application, the video frames at at least the previous bit before the output bit are the video frames at the previous N bits before the output bit, where N is a positive integer greater than or equal to 1 and N is determined by the configuration file of the control unit.
[0008] In a possible implementation manner of the first aspect of the present application, when the picture state flag bits of the video frame at the output bit and the video frames at the previous N bits before the output bit both indicate the non-existence of abnormal content, the video frame at the output bit is output, otherwise, the preset picture is output.
[0009] In a possible implementation manner of the first aspect of the present application, the analysis result is generated by the analysis unit performing feature extraction and content recognition on the plurality of video frames, including: through the analysis unit, performing fuzzy judgment on the overall abnormality of the plurality of video frames, and then respectively performing precise judgment on the individual abnormality of one or more video frames in the plurality of video frames, so as to obtain the analysis result, where the number and distribution of the one or more video frames are determined by the fuzzy judgment result of the overall abnormality.
[0010] In a possible implementation of the first aspect of the present application, through the analysis unit, a fuzzy determination of overall anomalies of the multiple video frames is performed, including: determining the degree of change of the reference indicators of the multiple video frames, and increasing or decreasing the number of the one or more video frames according to the level of the degree of change of the reference indicators of the multiple video frames, where the reference indicators are brightness, pixels, grayscale, or specific contour recognition.
[0011] In a possible implementation of the first aspect of the present application, through the analysis unit, an accurate determination of individual anomalies of the one or more video frames among the multiple video frames is performed respectively, including: through the analysis unit, for the one or more video frames respectively, performing text feature extraction and text content recognition to generate a first confidence level, and performing image feature extraction and image content recognition to generate a second confidence level, so as to obtain an accurate determination result of individual anomalies based on the first confidence level and the second confidence level.
[0012] In a possible implementation of the first aspect of the present application, the analysis result is generated by the analysis unit performing feature extraction and content recognition on the multiple video frames, including: through the analysis unit, for one or more video frames among the multiple video frames respectively, performing text feature extraction and text content recognition to generate a first confidence level, and performing image feature extraction and image content recognition to generate a second confidence level, so as to generate the analysis result based on the first confidence level and the second confidence level.
[0013] In a possible implementation of the first aspect of the present application, the control unit includes a decision logic and a rule library, and the rule library is used to provide decision rules. Among them, through the control unit, based on the analysis result, the frame state flag bits of the multiple video frames are uniformly set, including: through the decision logic, based on the decision rules and the analysis result, determining a decision output, and the decision output is used to set the frame state flag bit of each video frame among the multiple video frames to indicate the existence of abnormal content or indicate the non-existence of abnormal content.
[0014] In a possible implementation of the first aspect of the present application, the analysis unit is based on an artificial intelligence model, and the configuration file of the analysis unit includes an image model path, an image anomaly label, an image anomaly threshold, an image scaling ratio, a text box detection model path, a text box score threshold, a text box ratio threshold, a text recognition model path, a sensitive word file path, a text classification enable flag, a language detection enable flag, a language type detection model path, a language type detection threshold, a text classification model path, a text classification anomaly label, and a text classification threshold.
[0015] In a second aspect, an embodiment of the present application further provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method according to any implementation manner of any aspect above is implemented.
[0016] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium storing computer instructions. When the computer instructions run on a computer device, the computer device is caused to execute the method according to any implementation manner of any aspect above.
[0017] In a fourth aspect, an embodiment of the present application further provides a computer program product, which includes instructions stored on a computer-readable storage medium. When the instructions run on a computer device, the computer device is caused to execute the method according to any implementation manner of any aspect above.
[0018] In a fifth aspect, the present application provides a device for analyzing and controlling output video content. The device includes: a video input unit configured to simultaneously transmit a plurality of video frames received by the video input unit within a first time period to an analysis unit and a control unit; the analysis unit configured to perform feature extraction and content recognition on the plurality of video frames to generate an analysis result; the control unit configured to store the plurality of video frames in a frame buffer area in the control unit, and then, based on the analysis result, uniformly set the picture state flag bits of the plurality of video frames respectively, and then, based on the picture state flag bit of the video frame located at the output position in the frame buffer area and the picture state flag bits of at least one previous video frame located at the output position in the frame buffer area, output the video frame located at the output position, or output a preset picture, where the length of the first time period is a preset video output delay time, and the plurality of video frames respectively, starting from when they are received by the video input unit, after the preset video output delay time, serve as the video frame located at the output position.
[0019] Through the fifth aspect of the present application, by simultaneously transmitting a plurality of video frames received by the video input unit within a first time period to the analysis unit and the control unit, in this way, the video frames move towards the output bit of the frame buffer along the data frame movement direction in the frame buffer area, that is, after a preset video output delay time, the video frames serve as the video frames located at the output bit, and, synchronously in the analysis unit, feature extraction and content recognition are performed on the plurality of video frames to generate an analysis result; in this way, the artificial intelligence model efficiently utilizes the preset video output delay time, so that an analysis result can be generated based on the detection frames of the artificial intelligence model, and then the picture state flag bits of these video frames can be uniformly set; by using the picture state flag bits of the video frames located at the output bit in the frame buffer area and the picture state flag bits of the video frames located at at least the previous one bit of the output bit in the frame buffer area, content recognition with better timeliness is realized, which is beneficial to avoiding missed detection situations; a flexible control method capable of dynamically adjusting various parameters is realized, and more effective and timely response content management and monitoring are realized.
[0020] In a possible implementation manner of the fifth aspect of the present application, the control unit is configured to, based on the analysis result, set the picture state flag bit of each of the plurality of video frames to indicate the presence of abnormal content or indicate the absence of abnormal content.
[0021] In a possible implementation manner of the fifth aspect of the present application, the analysis unit is configured to perform a fuzzy judgment on the overall abnormality of the plurality of video frames, and then respectively perform an accurate judgment on the individual abnormality of one or more of the plurality of video frames, so as to obtain the analysis result, where the number and distribution of the one or more video frames are determined by the fuzzy judgment result of the overall abnormality.
[0022] In a possible implementation manner of the fifth aspect of the present application, the analysis unit is configured to determine the degree of change of the reference index of the plurality of video frames, and increase or decrease the number of the one or more video frames according to the level of the degree of change of the reference index of the plurality of video frames, where the reference index is brightness, pixels, grayscale or specific contour recognition.
[0023] In a possible implementation manner of the fifth aspect of the present application, the analysis unit is configured to respectively perform text feature extraction and text content recognition on the one or more video frames to generate a first confidence level, and perform image feature extraction and image content recognition to generate a second confidence level, so as to obtain an accurate judgment result of individual abnormality based on the first confidence level and the second confidence level. Description of the Drawings
[0024] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0025] Figure 1 It is a schematic flowchart of a method for analyzing and controlling the output video content provided by an embodiment of the present application;
[0026] Figure 2 It is a schematic diagram of the internal process of an analysis module provided by an embodiment of the present application;
[0027] Figure 3 It is a schematic diagram of the internal logic of a control unit provided by an embodiment of the present application;
[0028] Figure 4 It is a schematic diagram of a device for analyzing and controlling the output video content provided by an embodiment of the present application;
[0029] Figure 5 It is a schematic diagram of the structure of a computing device provided by an embodiment of the present application. Detailed implementation manners
[0030] The following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0031] It should be understood that in the description of the present application, "at least one" means one or more, and "a plurality" means two or more. In addition, words such as "first" and "second" are only used for the purpose of distinguishing descriptions unless otherwise specified, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.
[0032] Figure 1 It is a schematic flowchart of a method for analyzing and controlling the output video content provided by an embodiment of the present application. As Figure 1 shown, the method for analyzing and controlling the output video content includes the following steps.
[0033] Step S101: Through the video input unit, simultaneously transmit a plurality of video frames received by the video input unit within the first time period to the analysis unit and the control unit.
[0034] Step S103: Through the control unit, store the multiple video frames into the frame buffer in the control unit. Then, based on the analysis result, uniformly set the picture status flag bits of the multiple video frames respectively. Then, based on the picture status flag bit of the video frame located at the output position in the frame buffer and the picture status flag bits of at least the previous one video frame located at the output position in the frame buffer, output the video frame located at the output position, or output a preset picture.
[0035] Figure 1 The method for analyzing and controlling the output video content as shown, the analysis result is generated by the analysis unit extracting features and identifying the content of the multiple video frames, the length of the first time period is a preset video output delay time, and the multiple video frames respectively, starting from when they are received by the video input unit, after the preset video output delay time, serve as the video frame located at the output position. Figure 1 The method for analyzing and controlling the output video content as shown can be applied to the monitoring and filtering of video content in various application scenarios, especially in application scenarios that require low latency, timely monitoring, and timely response, such as various programs with a live or quasi-live nature, such as large-scale online meetings, live broadcasts of sports and cultural programs, etc. With the increasing richness of video content, the increase in video sources, and the expansion of the video data scale, if relying on predefined rules or pre-written image recognition algorithms to identify the information to be filtered in the video content (such as sensitive information, inappropriate information, abnormal information, etc.), it may not be possible to achieve effective identification and timely processing. For this reason, Figure 1 The method for analyzing and controlling the output video content as shown introduces an artificial intelligence algorithm based on deep learning, and uses feature extraction and content recognition to generate an analysis result. In this way, the accuracy and real-time performance of video content recognition are improved. The analysis unit, such as an analysis module based on an artificial intelligence model, can analyze the video stream in real time and identify more complex and variable sensitive content. In addition, the control unit can, according to the analysis result obtained by the artificial intelligence model, timely decide whether to cut off the video output, that is, output a preset picture (such as a still picture indicating that the current video output is interrupted), so as to achieve more effective content management and monitoring.
[0036] Refer to Figure 1, in step S101, through the video input unit, a plurality of video frames received by the video input unit within the first time period are simultaneously transmitted to the analysis unit and the control unit. In this way, compared with first storing the video content and then analyzing the stored video content, by simultaneously transmitting a plurality of video frames received by the video input unit within the first time period to the analysis unit and the control unit, the storage step is skipped, and the analysis unit directly performs real-time analysis. Then, the control unit decides how to output externally according to the real-time analysis result of the analysis unit, for example, decides whether to transmit the video content or broadcast a preset picture. In step S103, through the control unit, the plurality of video frames are stored in the frame buffer area in the control unit. Then, based on the analysis result, the picture state flag bits of the plurality of video frames are uniformly set. Then, based on the picture state flag bit of the video frame located at the output position in the frame buffer area and the picture state flag bits of at least the previous one video frame located at the output position in the frame buffer area, the video frame located at the output position is output, or a preset picture is output. And, the analysis result is generated by the analysis unit performing feature extraction and content recognition on the plurality of video frames. The length of the first time period is a preset video output delay time. The plurality of video frames, respectively, starting from when they are received by the video input unit, after the preset video output delay time, serve as the video frame located at the output position. In this way, by using the frame buffer in the control unit, it is realized that the received video content, such as the data received by the receiving end of the High Definition Multimedia Interface (HDMI), that is, a plurality of video frames received by the video input unit within the first time period, is first stored in the frame buffer area, and simultaneously detected based on an artificial intelligence model through the analysis unit. Based on the analysis result generated by the analysis unit, the picture state flag bits of the plurality of video frames are uniformly set. In this way, each video frame can be understood as including a data frame and a picture state flag bit, where the data frame corresponds to the video information when the video frame is played. Here, the analysis result generated by the analysis unit is generated by the analysis unit performing feature extraction and content recognition on the plurality of video frames, so it is for the plurality of video frames received by the video input unit within the first time period. In other words, the video input unit receives multiple batches of video frames at intervals according to the length of the first time period. For each batch of video frames, feature extraction and content generation are performed based on an artificial intelligence model to generate an analysis result, and then based on this analysis result, the picture state flag bits of these video frames in the same batch are uniformly set, for example, set to have abnormal content or no abnormal content. In this way, it is possible to effectively avoid leaking abnormal content and realize the monitoring and control of video content.In addition, considering that the input video frames are continuously stored in the frame buffer and then, according to the moving direction of the data frames, serve as the output video frames, the video frames at the output position in the frame buffer are continuously changing. The output position in the frame buffer is a fixed position in the frame buffer, and each video frame at this fixed position is the current output data frame. The video frames at at least one position before the output position in the frame buffer mean one or more positions before the output position, that is, the video frames at at least one position before the output position in the frame buffer are output after the video frames at the output position in the frame buffer are output.
[0037] Continue to refer to Figure 1, through the control unit, based on the analysis result, uniformly set the picture status flag bits of the multiple video frames. Thus, taking the length of the first time period as the time interval length, receive multiple batches of video frames at intervals and perform feature extraction and content generation on each batch of video frames based on the artificial intelligence model to generate corresponding analysis results, and then uniformly set the picture status flag bits of these video frames in the corresponding batch of video frames based on the analysis results. Therefore, as long as one video frame is identified as having abnormal content, such as sensitive words or inappropriate picture information, as an artificial intelligence detection frame, all video frames in the same batch as this video frame will be set with the picture status flag bit indicating abnormal content. Further, through the control unit, based on the picture status flag bit of the video frame located at the output position in the frame buffer and the picture status flag bits of at least the previous one video frame located at the output position in the frame buffer, output the video frame located at the output position, or output a preset picture. Thus, when the picture status flag bit of the video frame located at the output position in the frame buffer indicates the existence of abnormal content, or the picture status flag bits of at least the previous one video frame located at the output position in the frame buffer indicate the existence of abnormal content, output a preset picture. In other words, only when the picture status flag bit of the video frame located at the output position and the picture status flag bits of at least the previous one video frame located at the output position both indicate the non-existence of abnormal content, will the video frame located at the output position be output. Here, after the video frame is stored in the frame buffer, it moves towards the output position of the frame buffer according to the data frame movement direction until it is located at the output position of the frame buffer. The at least previous one video frame located at the output position refers to one or more video frames in front of this output position along the data frame movement direction, that is, the at least previous one video frame located at the output position in the frame buffer is output after the video frame located at the output position in the frame buffer is output. Such a design is considered because after multiple video frames in a batch are output one by one, for the last few or the last video frame to be output, the timeliness of its own picture status flag bit may be insufficient. For this reason, by setting the picture status flag bit of the video frame located at the output position in the frame buffer and the picture status flag bits of at least the previous one video frame located at the output position in the frame buffer, in this way, for the video frames received earlier and later in a batch, good timeliness content recognition can be achieved. In some embodiments, the video frame located at the output position in the frame buffer is a video frame of the previous batch, and the at least previous one video frame located at the output position in the frame buffer is a video frame of the next batch relative to the previous batch. Such a design is beneficial to continuously and intermittently obtain multiple batches of video frames and perform feature extraction and content recognition, thereby avoiding missed detection situations.Furthermore, the length of the first time period is a preset video output delay time. The multiple video frames, respectively, starting from when they are received by the video input unit, after passing through the preset video output delay time, serve as the video frames at the output position. This means that after each video frame is received, it is delayed to a certain extent, that is, after passing through the preset video output delay time, and then serves as the video frame at the output position in the frame buffer. For example, the preset video output delay time can be set to 2 seconds, that is, the length of the first time period is 2 seconds. In this way, the video input unit transmits multiple video frames received within a time period of 2 seconds to the analysis unit and the control unit at intervals. Here, the number of multiple video frames received within a time period of 2 seconds may vary. For example, it may be 120 video frames or 100 video frames, which may depend on factors such as the video data stream. Additionally, the delay playback time can be dynamically adjusted, the size of the frame buffer can be dynamically adjusted, the occlusion screen, that is, the preset screen broadcast when abnormal content is detected, can be dynamically adjusted, and the control strategy for judging normal and abnormal can also be dynamically adjusted, etc.
[0038] In summary, Figure 1 The method for analyzing and controlling the output video content shown, by simultaneously transmitting multiple video frames received by the video input unit within the first time period to the analysis unit and the control unit, thus enabling the video frames to move towards the output position of the frame buffer along the data frame movement direction in the frame buffer, that is, the video frame serves as the video frame at the output position after passing through the preset video output delay time, and, synchronously in the analysis unit, performing feature extraction and content recognition on the multiple video frames to generate an analysis result; thus, the preset video output delay time is efficiently utilized by the artificial intelligence model, so that an analysis result can be generated based on the detection frames of the artificial intelligence model, and then the picture state flag bits of these video frames can be uniformly set; by using the picture state flag bits of the video frames at the output position in the frame buffer and the picture state flag bits of at least the previous one video frame at the output position in the frame buffer, content recognition with better timeliness is achieved, which is beneficial to avoiding missed detection situations; a flexible control method that can dynamically adjust multiple parameters is realized, and more effective and timely response content management and monitoring are achieved.
[0039] Figure 2 This is a schematic diagram of the internal process of an analysis module provided by an embodiment of the present application. As Figure 2As shown, the video signal input 205 undergoes feature extraction 207 and then content recognition 208, and finally the analysis result output 209 is obtained. Moreover, the video signal input 205 is transmitted to the deep learning model 206 to improve the effect of content recognition 208 through the deep learning model 206. Here, the video signal input 205 refers to receiving a video signal from an input end, such as a High-Definition Multimedia Interface (HDMI). The deep learning model 206 is used to apply the trained deep learning model 206 to process video frames. Feature extraction 207 extracts key features from video frames for subsequent content recognition 208. Content recognition 208 identifies whether the video content contains inappropriate or sensitive information based on the extracted features. The analysis result output 209 sends the recognition result to the control unit for decision-making processing.
[0040] Figure 3 This is a schematic diagram of the internal logic of a control unit provided by an embodiment of the present application. As Figure 3 shown, the analysis result input 310 is transmitted to the decision logic 311 and the rule library 312, and then the decision output 313 is obtained through the rule library 312. The analysis result input 310 refers to the control unit receiving the recognition result from the artificial intelligence analysis module. The decision logic 311 means that according to preset rules and the artificial intelligence analysis result, the decision logic 311 is executed. The rule library 312 stores rules for decision-making, and these rules can be configured according to different application scenarios. The decision output 313 means that according to the result of the decision logic 311, a control signal is output to the output unit to determine whether to output a video signal.
[0041] Refer to Figure 1 、 Figure 2 and Figure 3, in a possible implementation, through the control unit, based on the analysis result, the picture state flag bits of the multiple video frames are uniformly set, including: through the control unit, based on the analysis result, the picture state flag bit of each video frame in the multiple video frames is set to indicate the existence of abnormal content or indicate the non-existence of abnormal content. In this way, when the picture state flag bit of the video frame located at the output position in the frame buffer indicates the existence of abnormal content, or the picture state flag bit of the video frame located at at least the previous position of the output position in the frame buffer indicates the existence of abnormal content, a preset picture is output. In other words, only when the picture state flag bit of the video frame located at the output position and the picture state flag bits of the video frames located at at least the previous position of the output position both indicate the non-existence of abnormal content, the video frame located at the output position will be output. Taking the length of the first time period as the time interval length, multiple batches of video frames are received at intervals and feature extraction and content generation are performed on each batch of video frames based on an artificial intelligence model to generate corresponding analysis results, and then the picture state flag bits of the corresponding batch of video frames are uniformly set based on the analysis results. Therefore, as long as one video frame is identified as having abnormal content, such as sensitive words or inappropriate picture information, as an artificial intelligence detection frame, all the video frames in the same batch as this video frame will be set with the picture state flag bit indicating the existence of abnormal content. In this way, content recognition with better timeliness is achieved, which is beneficial to avoiding missed detection situations.
[0042] In a possible implementation, the video frame located at least one bit before the output bit is the video frame located at the first N bits before the output bit, where N is a positive integer greater than or equal to 1 and N is determined by the configuration file of the control unit. After the video frame is stored in the frame buffer, it moves towards the output bit of the frame buffer in the data frame moving direction until it is located at the output bit of the frame buffer. The video frame located at least one bit before the output bit refers to the video frame located one or more bits before the output bit along the data frame moving direction. That is, the video frame located at least one bit before the output bit in the frame buffer is output after the video frame located at the output bit in the frame buffer is output. Such a design is considered because after multiple video frames in a batch are output one by one, the picture status flag bits of the last few or the last video frame output may lack timeliness. Therefore, by setting the picture status flag bit of the video frame located at the output bit in the frame buffer and the picture status flag bit of the video frame located at least one bit before the output bit in the frame buffer, both the video frames received earlier and the video frames received later in a batch can achieve content recognition with good timeliness. In some embodiments, the video frame located at the output bit in the frame buffer is the video frame of the previous batch, and the video frame located at least one bit before the output bit in the frame buffer is the video frame of the next batch relative to the previous batch. Such a design is beneficial for continuously and intermittently obtaining multiple batches of video frames and performing feature extraction and content recognition, thereby avoiding missed detection. Further, the length of the first time period is the preset video output delay time. Each of the multiple video frames, starting from when it is received by the video input unit, passes through the preset video output delay time and serves as the video frame located at the output bit. This means that after each video frame is received, it undergoes a certain degree of delay, that is, passes through the preset video output delay time, and then serves as the video frame located at the output bit of the frame buffer. For example, the preset video output delay time can be set to 2 seconds, that is, the length of the first time period is 2 seconds. In this way, the video input unit simultaneously transmits multiple video frames received within a time period of 2 seconds to the analysis unit and the control unit at intervals. Here, the number of multiple video frames received within a time period of 2 seconds may vary, such as 120 video frames or 100 video frames, which may depend on factors such as the video data stream. In addition, the delay playback time can be dynamically adjusted, the size of the frame buffer can be dynamically adjusted, the occlusion screen, that is, the preset screen played when abnormal content is detected, can be dynamically adjusted, and the control strategy for judging normal and abnormal can also be dynamically adjusted.
[0043] In a possible implementation, when the picture status flag bit of the video frame at the output position and the picture status flag bits of the video frames at the first N positions before the output position both indicate the absence of abnormal content, the video frame at the output position is output; otherwise, the preset picture is output. In this way, when the picture status flag bit of the video frame at the output position in the frame buffer indicates the presence of abnormal content, or the picture status flag bit of at least one of the video frames before the output position in the frame buffer indicates the presence of abnormal content, the preset picture is output. In other words, only when the picture status flag bit of the video frame at the output position and the picture status flag bits of at least one of the video frames before the output position both indicate the absence of abnormal content, the video frame at the output position will be output. The artificial intelligence model efficiently utilizes the preset video output delay time, so that analysis results can be generated based on the detection frames of the artificial intelligence model, and then the picture status flag bits of these video frames can be uniformly set; by using the picture status flag bit of the video frame at the output position in the frame buffer and the picture status flag bits of at least one of the video frames before the output position in the frame buffer, content recognition with better timeliness is achieved, which is beneficial to avoiding missed detection; a flexible control method that can dynamically adjust multiple parameters is realized, and more effective and timely content management and monitoring are achieved.
[0044] In a possible implementation, the analysis result is generated by the analysis unit extracting features and identifying the content of the multiple video frames, including: through the analysis unit, making a fuzzy judgment on the overall abnormality of the multiple video frames, and then respectively making an accurate judgment on the individual abnormality of one or more of the multiple video frames, so as to obtain the analysis result. Among them, the number and distribution of the one or more video frames are determined by the fuzzy judgment result of the overall abnormality. As described above, by setting the length of the first time period as the preset video output delay time. And setting each of the multiple video frames as the video frame at the output position after passing through the preset video output delay time since it is received by the video input unit. In this way, a lightweight layout is achieved, which is beneficial to quickly and effectively analyzing and judging whether there is abnormal content, and helps to be deployed in application scenarios with high requirements for quick response and system efficiency, such as edge devices or input / output ports. Further, considering that abnormal content may appear in the form of malicious algorithms, such as using a splash screen or inserting frames into the video data stream, attempting to achieve the purpose of evading the detection algorithm. For this reason, a fuzzy judgment on the overall abnormality of the multiple video frames can be made first, so as to determine the number and distribution of the one or more video frames. In some embodiments, the fuzzy judgment of the overall abnormality can be made by using the change in brightness or the change in the reference object. For a situation with a drastic change, a larger number of video frames are extracted for subsequent accurate judgment of individual abnormality, while for a situation with a gentle change, a smaller number of video frames are extracted for subsequent accurate judgment of individual abnormality. For example, assuming that the length of the first time period is 2 seconds and corresponds to 120 video frames, 3 video frames can be extracted for subsequent analysis in the case of a drastic change, while 2 or 1 video frame can be extracted for subsequent analysis in the case of a gentle change. This means that the video frames extracted within a certain time period are first analyzed preliminarily, the degree of change (such as the reference object) is judged, and a fuzzy judgment is made. Then, based on the fuzzy judgment result of the overall abnormality, a certain number of video frames are extracted for accurate judgment, which helps to avoid the phenomenon of missed detection. Among them, the reference object is an object that can be preferentially labeled during training. For example, sensitive content to be reviewed may be more common in billboards, flags, hats (for example, based on the shape outline of the object, objects with certain specific outlines can be used as key reference objects, which can be used for fuzzy judgment and for blocking the output of the screen). In this way, by combining fuzzy judgment and accurate judgment, the accuracy of identifying abnormal content is improved, resource waste is avoided, and a high system efficiency is maintained.
[0045] In a possible implementation, through the analysis unit, a fuzzy determination of the overall abnormality of the multiple video frames is performed, including: determining the degree of change of the reference indicators of the multiple video frames, and increasing or decreasing the number of the one or more video frames according to the level of the degree of change of the reference indicators of the multiple video frames, where the reference indicators are brightness, pixels, grayscale, or specific contour recognition. In this way, the video frames extracted within a certain period of time are first analyzed preliminarily to judge the degree of change (such as the reference object), and a fuzzy determination is made. Then, based on the fuzzy determination result of the overall abnormality, a certain number of video frames are extracted for accurate determination, which helps to avoid the phenomenon of missed detection. The fuzzy determination can be based on the reference indicators, and the reference indicators can be brightness, pixels, grayscale, or specific contour recognition. For example, the fuzzy determination can be based on the brightness analysis result or the pixel analysis result or the grayscale analysis result. The reference object is an object that can be preferentially marked during training. For example, sensitive content to be reviewed may be mostly seen on billboards, flags, hats (for example, based on the shape and contour of the object for classification, certain objects with specific contours can be used as key reference objects for fuzzy determination and for outputting the occluded picture). The fuzzy determination algorithm can be the brightness value, grayscale analysis, or specific contour object recognition. For example, a square frame with sensitive content pops up. The sensitive content can be played in a flashing manner, appearing and then disappearing in some video frames, taking advantage of the phenomenon of human eye retention to bypass the algorithm obstacles of video frame sampling inspection. By performing the fuzzy determination algorithm, abnormal phenomena in the video frame data stream can be identified, such as abnormal changes in brightness and pixels. In this way, by combining fuzzy determination and accurate determination, the accuracy of identifying abnormal content is improved, resource waste is avoided, and a high system efficiency is maintained.
[0046] In a possible implementation, through the analysis unit, precise judgment of individual anomalies is performed on the one or more video frames among the multiple video frames respectively, including: through the analysis unit, for the one or more video frames respectively, text feature extraction and text content recognition are performed to generate a first confidence level, and image feature extraction and image content recognition are performed to generate a second confidence level, so as to obtain a precise judgment result of individual anomalies based on the first confidence level and the second confidence level. The built-in artificial intelligence analysis module of the analysis unit itself has two branches, one is the text feature extraction branch, and the other is the image feature extraction branch. Here, by combining the confidence levels generated by the two branches respectively, the recognition efficiency can be improved. For example, a probability of 70% can be set as the threshold. Then, assuming that the text feature extraction branch gives a first confidence level of 65% and the image feature extraction branch gives a second confidence level of 65%, both confidence levels are high but neither exceeds the threshold. In the case where the confidence levels of the two branches are both relatively high but do not reach the threshold, the sensitivity requirements of the application scenario can be combined to obtain a precise judgment result of individual anomalies, which in turn helps to make a video processing result of occlusion or blurring.
[0047] In a possible implementation, the analysis result is generated by the analysis unit performing feature extraction and content recognition on the multiple video frames, including: through the analysis unit, for one or more video frames among the multiple video frames respectively, text feature extraction and text content recognition are performed to generate a first confidence level, and image feature extraction and image content recognition are performed to generate a second confidence level, so as to generate the analysis result based on the first confidence level and the second confidence level. In this way, by using the text feature extraction branch to perform text feature extraction and text content recognition to generate a first confidence level, and using the image feature extraction branch to perform image feature extraction and image content recognition to generate a second confidence level, the sensitivity requirements of the application scenario can be combined to obtain a precise judgment result of individual anomalies, which in turn helps to make a video processing result of occlusion or blurring. In this way, the deep learning algorithm is used to perform real-time analysis on the video content. This technical means significantly improves the recognition accuracy of inappropriate or sensitive information in the video content. Compared with the traditional image recognition algorithm, it has higher recognition accuracy and adaptability.
[0048] In a possible implementation, the control unit includes a decision logic and a rule library, and the rule library is used to provide decision rules. Wherein, through the control unit, based on the analysis result, the picture status flag bits of the multiple video frames are uniformly set, including: through the decision logic, based on the decision rules and the analysis result, determining a decision output, and the decision output is used to set the picture status flag bit of each video frame in the multiple video frames to indicate the existence of abnormal content or indicate the non-existence of abnormal content. In this way, an intelligent decision-making and control mechanism is realized. The control unit intelligently decides whether to cut off the video output according to the output result of the AI analysis module and the preset rules, realizing the automation and intelligence of content management, and improving the practicability and flexibility of the system.
[0049] In a possible implementation, the analysis unit is based on an artificial intelligence model, and the configuration file of the analysis unit includes an image model path, an image anomaly label, an image anomaly threshold, an image scaling ratio, a text box detection model path, a text box score threshold, a text box ratio threshold, a text recognition model path, a sensitive word file path, a text classification enable flag, a language detection enable flag, a language type detection model path, a language type detection threshold, a text classification model path, a text classification anomaly label, and a text classification threshold. In this way, by using the configuration file of the analysis unit, parameters such as the delayed playback time, buffer size, occluded picture, normal and abnormal control strategies can be dynamically adjusted, realizing flexible sensitivity configuration. For example, the image anomaly threshold is used to indicate the score threshold calculated by the artificial intelligence model, and only when it is greater than this threshold is it valid. The image scaling ratio is used to indicate the width ratio threshold of the image box to the original picture, and those less than this threshold are filtered out. The language type detection threshold is used to indicate the detection threshold of the language type output by the artificial intelligence model, and only when it is greater than this threshold is it credible. The text classification threshold is used to indicate the detection threshold of the text classification type output by the artificial intelligence model, and only when it is greater than this threshold is it credible.
[0050] Figure 4 This is a schematic diagram of a device for analyzing and controlling the output video content provided by an embodiment of the present application. As Figure 4As shown, the device for analyzing and controlling the output video content includes: a video input unit 401, configured to simultaneously transmit multiple video frames received by the video input unit 401 within a first time period to an analysis unit 403 and a control unit 405; the analysis unit 403, configured to perform feature extraction and content recognition on the multiple video frames to generate an analysis result; the control unit 405, configured to store the multiple video frames in a frame buffer area in the control unit 405, and then, based on the analysis result, uniformly set the picture status flag bits of the multiple video frames respectively, and then, based on the picture status flag bit of the video frame located at the output position in the frame buffer area and the picture status flag bits of at least the previous one video frame located at the output position in the frame buffer area, output the video frame located at the output position, or output a preset picture. Wherein, the length of the first time period is a preset video output delay time, and each of the multiple video frames, starting from the time when it is received by the video input unit 401, after the preset video output delay time, serves as the video frame located at the output position. Figure 4 An output unit 410 is also shown. The output unit 410 is used to provide an external output, and the external output of the output unit 410 is determined by the control unit 405, that is, it is determined by the control unit 405 to output the video frame located at the output position, or output a preset picture.
[0051] Figure 4 The shown device for analyzing and controlling the output video content, by simultaneously transmitting multiple video frames received by the video input unit 401 within a first time period to the analysis unit 403 and the control unit 405, thus enables the video frames to move along the data frame movement direction in the frame buffer area towards the output position of the frame buffer area, that is, the video frame serves as the video frame located at the output position after the preset video output delay time, and, synchronously in the analysis unit 403, performs feature extraction and content recognition on the multiple video frames to generate an analysis result; thus, the artificial intelligence model efficiently utilizes the preset video output delay time, so that an analysis result can be generated based on the detection frames of the artificial intelligence model, and then the picture status flag bits of these video frames are uniformly set; by using the picture status flag bit of the video frame located at the output position in the frame buffer area and the picture status flag bits of at least the previous one video frame located at the output position in the frame buffer area, content recognition with better timeliness is achieved, which is beneficial to avoiding missed detection situations; a flexible control means capable of dynamically adjusting multiple parameters is realized, and more effective and timely response content management and monitoring are achieved.
[0052] In a possible implementation, the control unit 405 is configured to set the picture status flag bit of each of the multiple video frames to indicate the presence of abnormal content or the absence of abnormal content based on the analysis result. In this way, when the picture status flag bit of the video frame at the output position in the frame buffer indicates the presence of abnormal content, or the picture status flag bit of the video frame at least one position before the output position in the frame buffer indicates the presence of abnormal content, a preset picture is output. In other words, only when the picture status flag bit of the video frame at the output position and the picture status flag bit of the video frame at least one position before the output position both indicate the absence of abnormal content, the video frame at the output position will be output. Taking the length of the first time period as the time interval length, multiple batches of video frames are received at intervals, and feature extraction and content generation are performed on each batch of video frames based on an artificial intelligence model to generate corresponding analysis results, and then the picture status flag bits of these video frames in the corresponding batch are uniformly set based on the analysis results. Therefore, as long as one video frame is identified as having abnormal content, such as sensitive words or inappropriate picture information, as an artificial intelligence detection frame, all the video frames in the same batch as this video frame will be set with the picture status flag bit indicating the presence of abnormal content. In this way, content recognition with better timeliness is achieved, which is beneficial to avoiding missed detection situations.
[0053] In a possible implementation, the analysis unit 403 is configured to perform a fuzzy judgment on the overall abnormality of the multiple video frames, and then perform an accurate judgment on the individual abnormality of one or more of the multiple video frames respectively, so as to obtain the analysis result, where the number and distribution of the one or more video frames are determined by the result of the fuzzy judgment on the overall abnormality. In this way, by combining fuzzy judgment and accurate judgment, the accuracy of identifying abnormal content is improved, resource waste is avoided, and a relatively high system efficiency is maintained.
[0054] In a possible implementation, the analysis unit 403 is configured to determine the degree of change of the reference metrics of the multiple video frames, and increase or decrease the number of the one or more video frames according to the level of the degree of change of the reference metrics of the multiple video frames, where the reference metrics are brightness, pixels, grayscale, or specific contour recognition. In this way, the video frames extracted within a certain period of time are initially analyzed to judge the degree of change (such as the reference object), and a fuzzy judgment is made. Then, based on the fuzzy judgment result of overall abnormality, a certain number of video frames are extracted for accurate judgment, which helps to avoid missed detection. The fuzzy judgment can be based on the reference metrics, and the reference metrics can be brightness, pixels, grayscale, or specific contour recognition. For example, the fuzzy judgment can be based on the brightness analysis result, pixel analysis result, or grayscale analysis result. The reference object is an object that can be preferentially labeled during training. For example, sensitive content to be reviewed may be more common in billboards, flags, hats (for example, based on the shape contour of the object, objects with certain specific contours can be used as key reference objects for fuzzy judgment and for outputting the obscured picture). The fuzzy judgment algorithm can be the brightness value, grayscale analysis, or specific contour object recognition. For example, a square frame with sensitive content pops up. The sensitive content can be played in a flashing manner, appearing and then disappearing in some video frames, taking advantage of the phenomenon of human eye persistence to bypass the algorithm obstacle of video frame sampling inspection. By performing the fuzzy judgment algorithm, abnormal phenomena in the video frame data stream, such as abnormal changes in brightness and pixels, can be identified. In this way, by combining fuzzy judgment and accurate judgment, the accuracy of identifying abnormal content is improved, resource waste is avoided, and a relatively high system efficiency is maintained.
[0055] In a possible implementation, the analysis unit 403 is configured to separately perform text feature extraction and text content recognition on the one or more video frames to generate a first confidence level, and perform image feature extraction and image content recognition on the one or more video frames to generate a second confidence level, so as to obtain an accurate judgment result of individual abnormality based on the first confidence level and the second confidence level. In this way, the sensitivity requirements of the application scenario can be combined to obtain an accurate judgment result of individual abnormality, which further helps to make a video processing result of occlusion or blurring.
[0056] It should be understood that a method and apparatus for analyzing and controlling output video content provided by specific embodiments of the present application can be applied to public security monitoring, such as crowded areas like stations and shopping malls, and can also be applied to content review, such as content management on video platforms, and can also be applied to home entertainment systems to provide parental control functions, as well as other extended applications. By using the method and apparatus for analyzing and controlling output video content, the intelligent level of video content management is improved, and the spread of inappropriate content is effectively prevented; the system has a fast response speed, can process video streams in real time, and meets the requirements of real-time monitoring; it can be flexibly configured to adapt to different monitoring environments and requirements. Moreover, a method and apparatus for analyzing and controlling output video content provided by specific embodiments of the present application have the following technical improvement points:
[0057] (1) Application of deep learning algorithms: Deep learning algorithms are used to perform real-time analysis on video content. This technical means significantly improves the recognition accuracy of inappropriate or sensitive information in video content. Compared with traditional image recognition algorithms, it has higher recognition accuracy and adaptability.
[0058] (2) Real-time video stream processing ability: It can achieve real-time processing of HDMI input video streams, which is realized by a high-performance Artificial Intelligence (AI) analysis module. This real-time nature ensures that the system can respond and process video content in a timely manner, meeting the requirements of real-time monitoring.
[0059] (3) Intelligent decision-making and control mechanism: The control unit intelligently decides whether to cut off the video output according to the output results of the AI analysis module and preset rules. It realizes the automation and intelligence of content management, and improves the practicality and flexibility of the system.
[0060] (4) Flexible sensitivity configuration: Allows users to flexibly configure the sensitivity level of the system according to different application scenarios and requirements. This design enables the system to adapt to different monitoring environments, enhancing its versatility and applicability.
[0061] (5) Modular system design: The system adopts a modular design, including HDMI input, AI analysis, control, and output units. This design makes the system easy to expand and maintain, and is also convenient for customization according to different requirements.
[0062] (6) High-precision feature extraction technology: In the AI analysis module, the feature extraction technology is the key to achieving high-precision content recognition. Key features in video frames are extracted through deep learning models, providing a reliable basis for content recognition.
[0063] (7) Adaptive learning and model optimization: The AI analysis module has the ability of adaptive learning, which can continuously optimize the model according to new data, improving the recognition accuracy and the robustness of the system.
[0064] (8) User-friendly interface and operation process: The system provides a user-friendly operation interface, enabling users to easily configure system parameters, monitor and manage video content. This design improves the usability of the system and reduces the operation difficulty.
[0065] Figure 5 FIG. is a schematic structural diagram of a computing device provided by an embodiment of the present application. The computing device 500 includes: one or more processors 510, a communication interface 520, and a memory 530. The processor 510, the communication interface 520, and the memory 530 are interconnected through a bus 540. Optionally, the computing device 500 may further include an input / output interface 550, and the input / output interface 550 is connected to input / output devices for receiving parameters set by users, etc. The computing device 500 can be used to implement some or all of the functions of the device embodiment or the system embodiment in the above embodiment of the present application; the processor 510 can also be used to implement some or all of the operation steps of the method embodiment in the above embodiment of the present application. For example, the specific implementation of various operations performed by the computing device 500 can refer to the specific details in the above embodiments, such as the processor 510 is used to execute some or all of the steps in the above method embodiment or some or all of the operations in the above method embodiment. Again, for example, in the embodiment of the present application, the computing device 500 can be used to implement some or all of the functions of one or more components in the above device embodiment. In addition, the communication interface 520 can specifically be used for communication functions necessary for implementing the functions of these devices and components, etc., and the processor 510 can specifically be used for processing functions necessary for implementing the functions of these devices and components, etc.
[0066] It should be understood that Figure 5 the computing device 500 may include one or more processors 510, and multiple processors 510 can cooperate to provide processing capabilities in a parallel connection mode, a serial connection mode, a serial-parallel connection mode, or any connection mode, or multiple processors 510 can form a processor sequence or a processor array, or the multiple processors 510 can be divided into a main processor and an auxiliary processor, or the multiple processors 510 can have different architectures such as using a heterogeneous computing architecture. In addition, Figure 5 for the computing device 500 shown, the related structural descriptions and functional descriptions are exemplary and non-limiting. In some exemplary embodiments, the computing device 500 may include more than Figure 5More or fewer components as shown, or combining certain components, or splitting certain components, or having different component arrangements.
[0067] The processor 510 can have various specific implementation forms. For example, the processor 510 can include one or more combinations of a central processing unit (CPU), a graphic processing unit (GPU), a neural-network processing unit (NPU), a tensor processing unit (TPU), or a data processing unit (DPU), etc. The embodiments of the present application do not make specific limitations. The processor 510 can also be a single-core processor or a multi-core processor. The processor 510 can be a combination of a CPU and a hardware chip. The above hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. The processor 510 can also be implemented separately using a logic device with built-in processing logic, such as an FPGA or a digital signal processor (DSP), etc. The communication interface 520 can be a wired interface or a wireless interface for communicating with other modules or devices. The wired interface can be an Ethernet interface, a local interconnect network (LIN), etc. The wireless interface can be a cellular network interface or use a wireless local area network interface, etc.
[0068] The memory 530 can be a non-volatile memory, for example, a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The memory 530 can also be a volatile memory, and the volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). The memory 530 can also be used to store program code and data, so as to facilitate the processor 510 to call the program code stored in the memory 530 to execute some or all of the operation steps in the above method embodiments, or to execute the corresponding functions in the above device embodiments. In addition, the computing device 500 may include more or fewer components than Figure 5 shown, or have a different component configuration.
[0069] The bus 540 can be a Peripheral Component Interconnect Express (PCIe) bus, or an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. The bus 540 can be divided into an address bus, a data bus, a control bus, etc. In addition to including a data bus, the bus 540 can also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity,Figure 5 It is represented by only one thick line in the figure, but it does not mean that there is only one bus or one type of bus.
[0070] The method and device provided by the embodiments of the present application are based on the same inventive concept. Since the principles of the method and the device for solving problems are similar, the embodiments, implementation manners, examples or implementation modes of the method and the device can be referred to each other, and the repeated parts will not be described again. The embodiments of the present application further provide a system, which includes a plurality of computing devices, and the structure of each computing device can refer to the structure of the computing device described above. The functions or operations that the system can implement can refer to the specific implementation steps in the above method embodiments and / or the specific functions described in the above device embodiments, and will not be described again here.
[0071] The embodiments of the present application further provide a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on a computer device (such as one or more processors), the method steps in the above method embodiments can be implemented. The specific implementation of the processor of the computer-readable storage medium in executing the above method steps can refer to the specific operations described in the above method embodiments and / or the specific functions described in the above device embodiments, and will not be described again here.
[0072] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. The present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. The embodiments of the present application can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The present application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center containing one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium, or a semiconductor medium. The semiconductor medium can be a solid-state drive, a random access memory, a flash memory, a read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, a register, or any other suitable form of storage medium.
[0073] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. Each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means for implementing the functions specified in Figure 1One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the processes Figure 1 One or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes.
[0074] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present application without departing from the spirit and scope of the embodiments of the present application. The steps in the method of the embodiments of the present application can be adjusted, combined or deleted according to actual needs; the modules in the system of the embodiments of the present application can be divided, combined or deleted according to actual needs. If these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these changes and variations.
Claims
1. A method for analyzing and controlling output video content, characterized in that, The method includes: Through a video input unit, transmitting multiple video frames received by the video input unit within a first time period to an analysis unit and a control unit simultaneously; Through the control unit, storing the multiple video frames in a frame buffer area in the control unit, then, based on an analysis result, uniformly setting the picture status flag bits of the multiple video frames respectively, and then, based on the picture status flag bit of the video frame at the output position in the frame buffer area and the picture status flag bits of at least the previous one video frame at the output position in the frame buffer area, outputting the video frame at the output position, or outputting a preset picture, Wherein, the analysis result is generated by the analysis unit performing feature extraction and content recognition on the multiple video frames, the length of the first time period is a preset video output delay time, and the multiple video frames respectively, from the time of being received by the video input unit, after the preset video output delay time, serve as the video frame at the output position, Through the control unit, uniformly setting the picture status flag bits of the multiple video frames respectively based on the analysis result includes: through the control unit, setting the picture status flag bit of each video frame in the multiple video frames to indicate the existence of abnormal content or indicate the non-existence of abnormal content based on the analysis result, When the picture status flag bit of the video frame at the output position and the picture status flag bits of the video frames at the previous N positions at the output position both indicate the non-existence of abnormal content, output the video frame at the output position, otherwise, output the preset picture.
2. The method according to claim 1, wherein The video frame at at least the previous one position at the output position is the video frame at the previous N positions at the output position, where N is a positive integer greater than or equal to 1 and N is determined by a configuration file of the control unit.
3. The method according to claim 1, wherein The analysis result is generated by the analysis unit performing feature extraction and content recognition on the multiple video frames, including: through the analysis unit, performing a fuzzy judgment on the overall abnormality of the multiple video frames, and then, respectively performing an accurate judgment on the individual abnormality of one or more of the multiple video frames, so as to obtain the analysis result, wherein the number and distribution of the one or more video frames are determined by the fuzzy judgment result of the overall abnormality.
4. The method according to claim 3, wherein Through the analysis unit, performing a fuzzy judgment on the overall abnormality of the multiple video frames includes: determining the degree of change of a reference index of the multiple video frames, and increasing or decreasing the number of the one or more video frames according to the level of the degree of change of the reference index of the multiple video frames, wherein the reference index is brightness, pixels, grayscale or specific contour recognition.
5. The method according to claim 3, characterized in that, Through the analysis unit, precise judgment of individual anomalies is respectively performed on the one or more video frames among the multiple video frames, including: through the analysis unit, respectively performing text feature extraction and text content recognition on the one or more video frames to generate a first confidence level, and performing image feature extraction and image content recognition to generate a second confidence level, so as to obtain a precise judgment result of individual anomalies based on the first confidence level and the second confidence level.
6. The method according to claim 1, characterized in that, The analysis result is generated by the analysis unit performing feature extraction and content recognition on the multiple video frames, including: through the analysis unit, respectively performing text feature extraction and text content recognition on one or more video frames among the multiple video frames to generate a first confidence level, and performing image feature extraction and image content recognition to generate a second confidence level, so as to generate the analysis result based on the first confidence level and the second confidence level.
7. The method according to claim 1, wherein The control unit includes a decision logic and a rule library, and the rule library is used to provide decision rules. Among them, through the control unit, based on the analysis result, the picture state flag bits of the multiple video frames are uniformly set, including: through the decision logic, based on the decision rules and the analysis result, determining a decision output, and the decision output is used to set the picture state flag bit of each video frame among the multiple video frames to indicate the existence of abnormal content or indicate the non-existence of abnormal content.
8. The method according to claim 1, characterized in that, The analysis unit is based on an artificial intelligence model, and the configuration file of the analysis unit includes an image model path, an image anomaly label, an image anomaly threshold, an image scaling ratio, a text box detection model path, a text box score threshold, a text box ratio threshold, a text recognition model path, a sensitive word file path, a text classification enable flag, a language detection enable flag, a language type detection model path, a language type detection threshold, a text classification model path, a text classification anomaly label, and a text classification threshold.
9. An apparatus for analyzing and controlling output video content, characterized in that, The device includes: A video input unit for simultaneously transmitting the multiple video frames received by the video input unit within a first time period to the analysis unit and the control unit; The analysis unit for performing feature extraction and content recognition on the multiple video frames to generate an analysis result; The control unit for storing the multiple video frames in the frame buffer area in the control unit, and then, based on the analysis result, uniformly setting the picture state flag bits of the multiple video frames respectively, and then, based on the picture state flag bit of the video frame located at the output position in the frame buffer area and the picture state flag bits of at least the previous one video frame located at the output position in the frame buffer area, outputting the video frame located at the output position, or outputting a preset picture, wherein the length of the first time period is a preset video output delay time, and the multiple video frames respectively, starting from when they are received by the video input unit, after the preset video output delay time, serve as the video frame located at the output position. The control unit is configured to set the picture status flag bit of each of the multiple video frames to indicate the presence of abnormal content or the absence of abnormal content based on the analysis result. When the picture status flag bits of the video frame at the output position and the video frames at the first N positions before the output position both indicate the absence of abnormal content, the video frame at the output position is output; otherwise, the preset picture is output.
10. The device according to claim 9, characterized in that, The analysis unit is configured to perform a fuzzy determination of overall abnormality on the multiple video frames, and then perform an accurate determination of individual abnormality on one or more of the multiple video frames respectively, so as to obtain the analysis result, wherein the number and distribution of the one or more video frames are determined by the result of the fuzzy determination of overall abnormality.
11. The device according to claim 10, characterized in that, The analysis unit is configured to determine the degree of change of the reference index of the multiple video frames, and increase or decrease the number of the one or more video frames according to the level of the degree of change of the reference index of the multiple video frames, wherein the reference index is brightness, pixel, grayscale or specific contour recognition.
12. The device according to claim 10, characterized in that, The analysis unit is configured to perform text feature extraction and text content recognition on the one or more video frames respectively to generate a first confidence level, and perform image feature extraction and image content recognition to generate a second confidence level, so as to obtain an accurate determination result of individual abnormality based on the first confidence level and the second confidence level.
Citation Information
Patent Citations
Method for switching photographing mode, and electronic device
CN103716535A
Method for precisely and remotely monitoring broadcast content through frames and controlling player
CN116017025A