A video detection method, device, electronic device and storage medium

By identifying and text alignment of the audio and subtitles of the video clips, the integrity of the video clips is automatically detected, and the problems of high and low-efficiency manual review costs in the prior art are solved, and efficient and low-cost video content integrity detection is achieved.

CN113591530BActive Publication Date: 2025-06-13SHENZHEN YAYUE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110207090.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-24
Publication Date
2025-06-13
Estimated Expiration
2041-02-24

AI Technical Summary

Technical Problem

In the prior art, incomplete review of video content mainly relies on manual review, resulting in high costs and low detection efficiency, which cannot meet the needs of a large number of video content to be launched.

Method used

By obtaining the audio information and subtitle text of the target video clip in the video to be detected, voice recognition and subtitle recognition are performed, text alignment is performed, and the integrity of the video clip is detected based on the alignment results.

Benefits of technology

It improves the efficiency of video integrity detection, reduces detection costs, and realizes automated video content integrity detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113591530B_ABST
    Figure CN113591530B_ABST
Patent Text Reader

Abstract

The present application discloses a video detection method, apparatus, electronic device, and storage medium. The present application can obtain audio information and subtitle text corresponding to a target video segment in a video to be detected, where the subtitle text includes at least one text unit. Perform speech recognition on the audio information to obtain target text corresponding to the audio information, where the target text includes at least one text unit. According to the text units of the target text and the text units of the subtitle text, perform text alignment on the target text and the subtitle text to obtain a text alignment result. According to the text alignment result, detect the integrity of the target video segment. The embodiments of the present application can perform video content integrity detection based on the audio and subtitles of a target video segment, improve the efficiency of video integrity detection, and reduce the detection cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a video detection method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of computer technology, the application of multimedia has become increasingly widespread, and various videos have emerged on the Internet. Among them, the quality of videos varies. In the highly competitive video content market, improving the user experience is very important, and ensuring the quality of video content is an important means to improve the user experience. Since most video content is spontaneously produced and uploaded by users, the production quality of videos usually requires the review of a review team. Video incompleteness is a very important item in quality review, which usually includes situations such as abrupt or incomplete beginnings or endings of videos, missing words, subtitles, and unexpressed content.

[0003] In the current related technologies, for the review of video content incompleteness, manual review is mainly adopted. The review team will formulate very detailed review criteria. After training the reviewers, the review background will transmit all the content in real time to the reviewers for the review of content incompleteness. With the continuous growth of the number of video content, the cost of traditional manual review also shows a linear growth trend, and the detection efficiency is relatively low, which cannot meet the needs of a large number of video content going online. Summary of the Invention

[0004] Embodiments of this application provide a video detection method, apparatus, electronic device, and storage medium, which can improve the efficiency of video integrity detection and reduce the detection cost.

[0005] Embodiments of this application provide a video detection method, including:

[0006] Obtain the audio information and subtitle text corresponding to the target video segment in the video to be detected, where the subtitle text includes at least one text unit;

[0007] Perform speech recognition on the audio information to obtain the target text corresponding to the audio information, where the target text includes at least one text unit;

[0008] Align the target text and the subtitle text according to the text units of the target text and the text units of the subtitle text to obtain a text alignment result;

[0009] Detect the integrity of the target video segment according to the text alignment result.

[0010] Correspondingly, embodiments of this application provide a video detection apparatus, including:

[0011] An acquisition unit, configured to acquire audio information and subtitle text corresponding to a target video segment in a video to be detected, where the subtitle text includes at least one text unit;

[0012] An identification unit, configured to perform speech recognition on the audio information to obtain target text corresponding to the audio information, where the target text includes at least one text unit;

[0013] An alignment unit, configured to align the target text and the subtitle text according to the text units of the target text and the text units of the subtitle text, to obtain a text alignment result;

[0014] A detection unit, configured to detect the integrity of the target video segment according to the text alignment result.

[0015] Optionally, in some embodiments of the present application, the acquisition unit may include an acquisition subunit, an extraction subunit, an extraction subunit, and an identification subunit, as follows:

[0016] The acquisition subunit is configured to acquire audio information corresponding to a target video segment in a video to be detected;

[0017] The extraction subunit is configured to extract video frames of the target video segment in the video to be detected to obtain at least one video frame image of the target video segment;

[0018] The extraction subunit is configured to determine a subtitle area of the video frame image by performing feature extraction on the video frame image;

[0019] The identification subunit is configured to identify subtitles in the subtitle area to obtain subtitle text corresponding to the target video segment.

[0020] Optionally, in some embodiments of the present application, the extraction subunit may specifically be configured to perform downsampling and upsampling processing on the video frame image at multiple scales to obtain a target feature map of the video frame image; perform a convolution operation on the target feature map to obtain a text unit heat map of the video frame image; and determine the subtitle area of the video frame image based on the text unit heat map.

[0021] Optionally, in some embodiments of the present application, the step of "performing downsampling and upsampling processing on the video frame image at multiple scales to obtain a target feature map of the video frame image" may include:

[0022] Performing multiple downsampling processes on the video frame image to obtain downsampled feature maps of the video frame image at multiple scales;

[0023] The downsampled feature map of the target scale is subjected to multiple upsampling processes to obtain the upsampled fusion feature maps of the video frame image at multiple scales, where the input for upsampling at each scale is the fusion feature obtained by fusing the upsampled feature map and the downsampled feature map of the adjacent scale;

[0024] Determine the target feature map of the video frame image from the upsampled fusion feature maps of each scale.

[0025] Optionally, in some embodiments of the present application, the text unit heat map includes a character heat map and an inter-character heat map; the step of "determining the subtitle area of the video frame image based on the text unit heat map" may include:

[0026] Select character areas from the character heat map according to the heat values of the heat points in the character heat map;

[0027] Select inter-character areas from the inter-character heat map according to the heat values of the heat points in the inter-character heat map;

[0028] Based on the character areas and the inter-character areas, determine the subtitle area of the video frame image.

[0029] Optionally, in some embodiments of the present application, the recognition sub-unit may specifically be used to extract features from the subtitle area to obtain a feature sequence of the subtitle area, where the feature sequence includes at least one feature information; according to the front and back feature information in the feature sequence, predict each feature information in the feature sequence to obtain the subtitle text corresponding to the target video segment.

[0030] Optionally, in some embodiments of the present application, the recognition unit may include an audio extraction sub-unit, a first determination sub-unit, and a second determination sub-unit, as follows:

[0031] The audio extraction sub-unit is used to perform semantic extraction on the audio information to obtain the audio semantic feature information of the audio information;

[0032] The first determination sub-unit is used to determine the prediction probability of translating the audio information into each candidate text based on the audio semantic feature information;

[0033] The second determination sub-unit is used to determine the target text corresponding to the audio information from the candidate texts based on the prediction probability.

[0034] Optionally, in some embodiments of the present application, the step of "performing downsampling and upsampling processes on the video frame image at multiple scales to obtain the target feature map of the video frame image" may include:

[0035] Through a subtitle area recognition model, perform downsampling and upsampling processing on the video frame image at multiple scales to obtain the target feature map of the video frame image;

[0036] The step of "performing a convolution operation on the target feature map to obtain the text unit heat map of the video frame image" may include:

[0037] Through a subtitle area recognition model, perform a convolution operation on the target feature map to obtain the text unit heat map of the video frame image.

[0038] Optionally, in some embodiments of the present application, the video detection device may further include a training unit; the training unit is used to train the subtitle area recognition model. Specifically, the training unit may be used to:

[0039] Obtain training data, where the training data includes sample images and the target subtitle areas corresponding to the sample images;

[0040] Through a preset subtitle area recognition model, perform downsampling and upsampling processing on the sample image at multiple scales to obtain the target feature map of the sample image;

[0041] Perform a convolution operation on the target feature map of the sample image to obtain the text unit heat map of the sample image;

[0042] Based on the text unit heat map, determine the reference subtitle area of the sample image;

[0043] Based on the reference subtitle area and the target subtitle area, adjust the parameters of the preset subtitle area recognition model to obtain the subtitle area recognition model.

[0044] Optionally, in some embodiments of the present application, the alignment unit may include a third determination subunit, a matching subunit, an update subunit, and a return subunit, as follows:

[0045] The third determination subunit is used to determine the target text unit at the target position in the target text;

[0046] The matching subunit is used to match the target text unit with the text units starting from the starting alignment position in the subtitle text;

[0047] The update subunit is used to, when there is a text unit in the text units starting from the starting alignment position in the subtitle text that matches the target text unit, update the target position in the target text and the starting alignment position in the subtitle text, and update the target text unit to the text unit at the updated target position;

[0048] A return subunit, configured to return the step of matching the target text unit with the text unit starting from the starting alignment position in the subtitle text until all text units in the target text are matched, so as to obtain a text alignment result.

[0049] An electronic device provided in an embodiment of the present application includes a processor and a memory. The memory stores multiple instructions, and the processor loads the instructions to execute the steps in the video detection method provided in the embodiment of the present application.

[0050] In addition, an embodiment of the present application further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the video detection method provided in the embodiment of the present application are implemented.

[0051] An embodiment of the present application provides a video detection method, apparatus, electronic device, and storage medium. Audio information and subtitle text corresponding to a target video segment in a video to be detected can be obtained. The subtitle text includes at least one text unit. Speech recognition is performed on the audio information to obtain target text corresponding to the audio information. The target text includes at least one text unit. According to the text units of the target text and the text units of the subtitle text, text alignment is performed on the target text and the subtitle text to obtain a text alignment result. According to the text alignment result, the integrity of the target video segment is detected. The embodiment of the present application can perform video content integrity detection based on the audio and subtitles of the target video segment, improve the efficiency of video integrity detection, and reduce the detection cost. Description of the Drawings

[0052] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1a is a schematic diagram of the scenario of the video detection method provided in the embodiment of the present application;

[0054] Figure 1b is a flowchart of the video detection method provided in the embodiment of the present application;

[0055] Figure 1c is a model structure diagram of the video detection method provided in the embodiment of the present application;

[0056] Figure 1d is an explanatory diagram of the video detection method provided in the embodiment of the present application;

[0057] Figure 1eIt is another explanatory diagram of the video detection method provided by the embodiments of the present application;

[0058] Figure 1f It is another model structure diagram of the video detection method provided by the embodiments of the present application;

[0059] Figure 1g It is another model structure diagram of the video detection method provided by the embodiments of the present application;

[0060] Figure 1h It is another explanatory diagram of the video detection method provided by the embodiments of the present application;

[0061] Figure 2a It is another flowchart of the video detection method provided by the embodiments of the present application;

[0062] Figure 2b It is another flowchart of the video detection method provided by the embodiments of the present application;

[0063] Figure 2c It is another explanatory diagram of the video detection method provided by the embodiments of the present application;

[0064] Figure 2d It is another explanatory diagram of the video detection method provided by the embodiments of the present application;

[0065] Figure 3a It is a schematic structural diagram of the video detection device provided by the embodiments of the present application;

[0066] Figure 3b It is another schematic structural diagram of the video detection device provided by the embodiments of the present application;

[0067] Figure 3c It is another schematic structural diagram of the video detection device provided by the embodiments of the present application;

[0068] Figure 3d It is another schematic structural diagram of the video detection device provided by the embodiments of the present application;

[0069] Figure 3e It is another schematic structural diagram of the video detection device provided by the embodiments of the present application;

[0070] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0071] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0072] An embodiment of the present application provides a video detection method, device, electronic device, and storage medium. The video detection device can be specifically integrated in an electronic device, and the electronic device can be a device such as a terminal or a server.

[0073] It can be understood that the video detection method in this embodiment can be executed on a terminal, on a server, or jointly executed by a terminal and a server. The above examples should not be construed as a limitation to the present application.

[0074] As Figure 1a shown, taking the case where a terminal and a server jointly execute the video detection method as an example. The video detection system provided by the embodiment of the present application includes a terminal 10 and a server 11, etc.; the terminal 10 and the server 11 are connected through a network, for example, through a wired or wireless network connection, etc., where the video detection device can be integrated in the server.

[0075] Among them, the server 11 can be used to: obtain the audio information and subtitle text corresponding to the target video segment in the video to be detected, where the subtitle text includes at least one text unit; perform speech recognition on the audio information to obtain the target text corresponding to the audio information, where the target text includes at least one text unit; perform text alignment on the target text and the subtitle text according to the text units of the target text and the text units of the subtitle text to obtain a text alignment result; detect the integrity of the target video segment according to the text alignment result, and send the detection result to the terminal 10. Among them, the server 11 can be a single server, or a server cluster or cloud server composed of multiple servers.

[0076] Among them, the terminal 10 can send the video information of the video to be detected to the server 11, so that the server 11 can detect the integrity of the video content of the video to be detected. The video information may include the video information in multiple modalities of the video to be detected, such as audio, images, text, etc. The terminal 10 can also receive the detection result sent by the server 11. If the detection result is that the video content of the target video segment is incomplete, the terminal 10 can edit and modify the target video segment of the video to be detected, and then send the modified video to be detected to the server 11 for video content integrity detection again; if the detection result is that the video content of the target video segment is complete, the video to be detected can be published on the video platform. Among them, the terminal 10 may include a mobile phone, a smart TV, a tablet computer, a notebook computer, a personal computer (PC, Personal Computer), a wearable device, or an in-vehicle computer, etc. A client can also be set on the terminal 10, and the client can be an application client or a browser client, etc.

[0077] The steps for the above-mentioned server 11 to perform video detection can also be executed by the terminal 10.

[0078] The video detection method provided by the embodiments of the present application relates to computer vision technology and speech technology in the field of artificial intelligence. The embodiments of the present application can perform video content integrity detection based on the audio and subtitles of the target video segment, improve the efficiency of video integrity detection, and reduce the detection cost.

[0079] Among them, artificial intelligence (AI, Artificial Intelligence) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Among them, artificial intelligence software technology mainly includes directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0080] Among them, computer vision technology (CV) is a science that studies how to enable machines to "see". Further, it refers to using cameras and computers to replace human eyes for machine vision such as target recognition, tracking, and measurement, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0081] Among them, the key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.

[0082] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0083] This embodiment will be described from the perspective of a video detection device, which can be specifically integrated in an electronic device, and the electronic device can be a device such as a server or a terminal.

[0084] The video detection method of the embodiments of this application can be applied to various scenarios that require video integrity detection of videos, where the video duration and video type are not limited. For example, a certain video platform needs to perform content integrity detection on millions of videos to conduct quality audits on the videos, remove the videos with incomplete content from the video platform, and improve the quality of the videos on the video platform. Specifically, the video detection method provided in this embodiment can be used to perform video content integrity detection based on the audio and subtitles of the target video segment, and can quickly detect a large number of videos.

[0085] As Figure 1b shown, the specific process of this video detection method can be as follows:

[0086] 101. Obtain the audio information and subtitle text corresponding to the target video segment in the video to be detected, and the subtitle text includes at least one text unit.

[0087] Among them, the video to be detected is a video for which the integrity of the video content needs to be detected. In a specific scenario, videos on a video playback platform need to undergo quality review, and video integrity is an important item in the quality review. Video incompleteness usually includes situations such as abrupt and incomplete beginnings or endings of the video, missing words, subtitles, and unfinished content expression.

[0088] Among them, the target video segment is a video segment in the video to be detected, which can specifically be the beginning segment, ending segment, etc. of the video to be detected. In one embodiment, the beginning segment and ending segment of the video to be detected can be subjected to video content integrity detection, and the integrity of the video to be detected can be determined based on the integrity of the beginning segment and ending segment. For example, after performing video content integrity detection on the beginning segment based on the audio information and subtitle text corresponding to the beginning segment, if it is found that there is unfinished content in the subtitles of the beginning segment, it can be determined that the beginning segment is incomplete, that is, the video content of the video to be detected is incomplete.

[0089] Specifically, the video to be detected can be divided into several video segments at regular time intervals; alternatively, the video to be detected can be divided into a certain number of video segments. The longer the duration of the video to be detected, the longer the time period length of the divided video segments, that is, the longer the lengths of the divided beginning segment and ending segment; it is also possible to use each sentence in the video to be detected as an interval to divide the video to be detected, and each video segment includes one sentence. The division method of the video segments in this embodiment is not limited and can be set according to actual situations.

[0090] Among them, the subtitle text includes at least one text unit. Specifically, when the subtitle text is in Chinese, one text unit can be one Chinese character; when the subtitle text is in English, one text unit can be one word. When it is other languages, the text unit is adjusted according to the corresponding language. When the subtitle is displayed vertically or in other directions, the text unit is adjusted according to the size of the text in the vertical or other corresponding directions.

[0091] The subtitles of the video can be divided into soft subtitles and hard subtitles. Soft subtitles are separately saved subtitle files, which can be directly extracted from the video data stream; hard subtitles are subtitles embedded in the video, and the subtitle file and the video stream are compressed together, and the subtitle file cannot be separated.

[0092] In some embodiments, if the subtitle corresponding to the target video segment of the video to be detected is a hard subtitle, then text recognition can be performed on the image sequence obtained by processing the target video segment to extract the subtitles in the image sequence.

[0093] Optionally, in this embodiment, the step of "obtaining the audio information and subtitle text corresponding to the target video segment in the video to be detected" may include:

[0094] Obtain the audio information corresponding to the target video segment in the video to be detected;

[0095] Extract video frames from the target video segment in the video to be detected to obtain at least one video frame image of the target video segment;

[0096] Determine the subtitle area of the video frame image by performing feature extraction on the video frame image;

[0097] Recognize the subtitles in the subtitle area to obtain the subtitle text corresponding to the target video segment.

[0098] Among them, in some embodiments, uniform frame extraction can be performed on the target video segment. For example, the target video segment can be divided into video sub-segments of equal duration. For example, taking 0.3 seconds as the division unit, the target video segment is divided into at least one 0.3-second video sub-segment, and one frame of picture, that is, a video frame image, is extracted from each video sub-segment. In other embodiments, only the video frame images containing subtitle text in the target video segment can be extracted. This embodiment does not limit the video frame extraction method.

[0099] Optionally, in this embodiment, the step of "determining the subtitle area of the video frame image by performing feature extraction on the video frame image" may include:

[0100] Perform downsampling and upsampling processing on the video frame image at multiple scales to obtain the target feature map of the video frame image;

[0101] Perform a convolution operation on the target feature map to obtain the text unit heat map of the video frame image;

[0102] Determine the subtitle area of the video frame image based on the text unit heat map.

[0103] Among them, a neural network can be used to perform feature extraction on the video frame image. This neural network can be a convolutional neural network (CNN, Convolutional Neural Networks), a visual geometry group network (VGGNet, Visual Geometry Group Network), a residual network (ResNet, Residual Network), a dense connection convolutional network (DenseNet, Dense Convolutional Network), etc. However, it should be understood that the neural network in this embodiment is not limited to the above-listed types.

[0104] Optionally, in this embodiment, the step of "performing downsampling and upsampling processing on the video frame image at multiple scales to obtain the target feature map of the video frame image" may include:

[0105] Performing multiple downsampling processes on the video frame image to obtain downsampled feature maps of the video frame image at multiple scales;

[0106] Performing multiple upsampling processes on the downsampled feature map of the target scale to obtain upsampled fusion feature maps of the video frame image at multiple scales, where the upsampling input for each scale is the fusion feature obtained by fusing the upsampled feature map and the downsampled feature map of the adjacent scale;

[0107] Determining the target feature map of the video frame image from the upsampled fusion feature maps of each scale.

[0108] During the sampling process, it is divided into upsampling and downsampling. For a low-resolution feature map, it can be restored to a high resolution by upsampling. Upsampling can upsample the finally obtained output to the size of the original image. The essence of upsampling is to enlarge the image and image interpolation, and the interpolation method can be the nearest neighbor method, bilinear interpolation method, cubic convolution interpolation method, etc. Downsampling is to shrink the image, which can make the image conform to the size of the display area and can generate a thumbnail of the corresponding image.

[0109] Among them, the target scale is the lowest scale among multiple scales. The step of "the upsampling input for each scale is the fusion feature obtained by fusing the upsampled feature map and the downsampled feature map of the adjacent scale" specifically includes: the upsampling input for each scale is the upsampled fusion feature map of the adjacent scale, that is, the fusion feature obtained by fusing the upsampled feature map and the downsampled feature map of the adjacent scale is the upsampled fusion feature map of the adjacent scale, where the upsampled feature map for each scale is obtained by performing an upsampling process on the upsampled fusion feature map of the adjacent scale. Reference Figure 1c, if the size of the video frame image is h*w*3, after multiple downsamplings of the video frame image, downsampled feature maps with sizes of h / 2*w / 2*64, h / 4*w / 4*128, h / 8*w / 8*256, h / 16*w / 16*512, and h / 32*w / 32*512 are obtained. Then, multiple upsamplings are performed on the downsampled feature map of h / 32*w / 32*512 to obtain upsampled fusion feature maps of the video frame image at multiple scales. For the upsampled feature map with a scale of h / 8*w / 8*128, its upsampling input is the upsampled fusion feature map with a scale of h / 16*w / 16*256, because upsampling the upsampled fusion feature map with a scale of h / 16*w / 16*256 can obtain the upsampled feature map with a scale of h / 8*w / 8*128. And the upsampled fusion feature map with a scale of h / 16*w / 16*256 is obtained by fusing the upsampled feature map with a scale of h / 16*w / 16*256 and the downsampled feature map with a scale of h / 16*w / 16*512, and so on for other scales. The adjacent scale of a certain scale can refer to the largest scale among the scales smaller than this scale. Specifically, it can also refer to the scale that is half of this scale. For example, there are scales h / 2*w / 2, h / 4*w / 4, h / 8*w / 8, h / 16*w / 16, and h / 32*w / 32. Among them, the adjacent scale of h / 8*w / 8 is h / 16*w / 16.

[0110] Among them, fusion refers to feature fusion. Fusing features of different scales can improve the feature representation ability. The resolution of low-level features is relatively high, containing more detailed information, but due to fewer convolutions, there is more noise and lower semanticity; high-level features have stronger semantic information, but their resolution is low and more details are lost. There are multiple ways of fusion. For example, the upsampled feature map and the downsampled feature map at the same scale can be concatenated; or the corresponding pixels of the upsampled feature map and the downsampled feature map at the same scale can be added. It can be understood that the ways of fusion are not limited to the above examples, and this embodiment does not limit this.

[0111] Specifically, in some embodiments, the step of "performing multiple upsampling processes on the downsampled feature map of the target scale to obtain upsampled fusion feature maps of the video frame image at multiple scales, where the upsampling input of each scale is the fusion feature obtained by fusing the upsampled feature map and the downsampled feature map of the adjacent scale" may include:

[0112] Based on the processing of the downsampled feature map with the lowest scale among multiple scales, a reference feature map with the same scale as the downsampled feature map with the lowest scale is obtained;

[0113] Upsample the reference feature map to obtain an upsampled feature map, and fuse the upsampled feature map with the downsampled feature map of the same scale as it to obtain the upsampled fused feature map of the video frame image;

[0114] Use the upsampled fused feature map as the new reference feature map, and return to execute the step of upsampling the reference feature map to obtain an upsampled feature map, and fusing the upsampled feature map with the downsampled feature map of the same scale as it to obtain the upsampled fused feature map of the video frame image, so as to obtain the upsampled fused feature maps of each scale of the video frame image.

[0115] Among them, the processing of the downsampled feature map with the lowest scale can specifically be to fuse the downsampled feature map of h / 32*w / 32*512 obtained in convolutional stage 5 and the downsampled feature map of h / 32*w / 32*512 obtained in convolutional stage 6 to obtain a reference feature map of h / 32*w / 32*512 with the same scale as it, as Figure 1c shown.

[0116] Among them, the step of "determining the target feature map of the video frame image from the upsampled fused feature maps of each scale" can specifically be: select the upsampled fused feature map with the largest scale from the upsampled fused feature maps of each scale, and perform upsampling on the upsampled fused feature map with the largest scale, and the obtained upsampled feature map is the target feature map of the video frame image. For example, referring to Figure 1c the fusion stage 4 in, fuse the upsampled feature map of scale h / 4*w / 4*64 with the downsampled feature map of h / 4*w / 4*128 to obtain an upsampled fused feature map of scale h / 4*w / 4*64, and then perform upsampling processing on it to obtain a target feature map of h / 2*w / 2*32.

[0117] Optionally, in this embodiment, the text unit heat map includes a character heat map and an inter-character heat map; the step of "determining the subtitle area of the video frame image based on the text unit heat map" may include:

[0118] Select character regions from the character heat map according to the heat values of the heat points in the character heat map;

[0119] Select inter-character regions from the inter-character heat map according to the heat values of the heat points in the inter-character heat map;

[0120] Determine the subtitle area of the video frame image based on the character regions and the inter-character regions.

[0121] Among them, the text unit heat map contains the distribution information of text units, that is, the distribution information of characters. The heat value of each heat point in the character heat map represents the probability that a character exists at that heat point. The greater the heat value, the greater the probability. The heat value of each heat point in the inter-character heat map represents the probability that the heat point is between two or more characters. The greater the heat value, the greater the probability.

[0122] Specifically, the step of "selecting a character region from the character heat map according to the heat value of the heat point in the character heat map" may include: taking the heat points in the character heat map with heat values greater than a preset first threshold as candidate first heat points, and determining the character region according to the positions of the candidate first heat points. Among them, the preset first threshold can be set according to the actual situation. In some embodiments, the character heat map can be divided into multiple regions, and the regions where the percentage of candidate first heat points in the heat points exceeds a preset percentage are used as candidate character regions, and then the obtained candidate character regions are fused to obtain at least one character region. The fusion method can specifically be splicing, etc., and this embodiment does not limit this.

[0123] Specifically, the step of "selecting an inter-character region from the inter-character heat map according to the heat value of the heat point in the inter-character heat map" may include: taking the heat points in the inter-character heat map with heat values greater than a preset second threshold as candidate second heat points, and determining the inter-character region according to the positions of the candidate second heat points. Among them, the preset second threshold can be set according to the actual situation. In some embodiments, the inter-character heat map can be divided into multiple regions, and the regions where the percentage of candidate second heat points in the heat points exceeds a preset percentage are used as candidate inter-character regions, and then the obtained candidate inter-character regions are fused to obtain at least one inter-character region. The fusion method can specifically be splicing, etc., and this embodiment does not limit this.

[0124] As Figure 1c shown, after downsampling, upsampling, and convolutional processing of the video frame image, a h / 2*w / 2*1 region heat map and a h / 2*w / 2*1 affinity heat map of the video frame image can be obtained. The region heat map (Regionheatmap) is the character heat map (character Gaussian heat map) in the above embodiment. The heat value of each heat point in the region heat map represents the region score. The heat (i.e., the region score) at the center of the character is the highest, and the heat at the character edge and the background (non-character region) is lower (specifically, it can be 0). The affinity heat map (Affinityheatmap) is the inter-character heat map (inter-character Gaussian heat map) in the above embodiment. The heat value of each heat point in the affinity heat map represents the affinity score. The heat (i.e., the affinity score) at the interval between characters is the highest, and the heat at the interval between non-characters and characters is lower (specifically, it can be 0). As Figure 1dAs shown, it is a schematic diagram of the character Gaussian heatmap and the inter-character Gaussian heatmap. It can be seen from the figure that the character Gaussian heatmap has a higher brightness at the position corresponding to the character, that is, a higher regional score; the inter-character Gaussian heatmap has a higher brightness in the area between characters, that is, a higher affinity score.

[0125] Optionally, in some embodiments, the step of "determining the subtitle area of the video frame image based on the character area and the inter-character area" can specifically be: based on the character area and the inter-character area, obtaining a text box, that is, the subtitle area, through a quadrilateral synthesis algorithm. As Figure 1e shown, the ellipse represents the character area. The center points of each character area can be taken, and the center points are connected to obtain the center line. At the center point of each character area, a perpendicular line perpendicular to the center line is drawn respectively to obtain the support points of the upper and lower edges. Finally, all the support points are connected to obtain a quadrilateral text box (subtitle area). A similar method can obtain a polygon box.

[0126] As Figure 1c shown, it is the network structure diagram of the CRAFT (Character Region Awareness for Text Detection) model. The CRAFT model can be used to identify the subtitle area. Specifically, the video frame image can be first downsampled and upsampled to obtain a target feature map, and then the target feature map is convolved through four convolutional layers to output the regional heatmap and the affinity heatmap of the video frame image. Based on the regional score in the regional heatmap and the affinity score in the affinity heatmap, the subtitle area of the video frame image can be obtained. Among them, when upsampling, the feature map can be regularized.

[0127] Optionally, in this embodiment, the step of "identifying the subtitles in the subtitle area to obtain the subtitle text corresponding to the target video segment" may include:

[0128] Performing feature extraction on the subtitle area to obtain a feature sequence of the subtitle area, where the feature sequence includes at least one feature information;

[0129] Predicting each feature information in the feature sequence according to the front and back feature information in the feature sequence to obtain the subtitle text corresponding to the target video segment.

[0130] Among them, the subtitles in the subtitle area can be identified through a neural network. The neural network can be a convolutional neural network, a recurrent neural network, a densely connected convolutional network, etc. However, it should be understood that the neural network in this embodiment is not limited to the above-listed types.

[0131] Specifically, the caption text in the caption area can be recognized by a CRNN (Convolutional Recurrent Neural Network), as Figure 1f shown, which is the network structure of CRNN. The CRNN network structure can be divided into three parts, namely CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), and CTC (Connectionist Temporal Classification).

[0132] Among them, as the convolutional layer, CNN can extract features from the input image (i.e., the image corresponding to the caption area) to obtain the feature sequence of the caption area; as the recurrent layer, RNN can specifically use a bidirectional RNN (such as BLSTM) to predict the feature sequence to obtain the candidate characters corresponding to the caption area; as the transcription layer, CTC can determine the target characters from the candidate characters obtained from the recurrent layer to obtain the caption text.

[0133] Among them, BLSTM (Bi-directional Long Short-Term Memory) is a bidirectional long short-term memory network, and BiLSTM is composed of a forward long short-term memory network (LSTM, Long Short-Term Memory) and a backward long short-term memory network. LSTM is a time recurrent neural network, that is, a kind of recurrent neural network (RNN, Recurrent Neural Network). LSTM is more suitable for extracting semantic features from time series data and is often used to extract semantic features from context information in natural language processing tasks. LSTM can selectively forget some historical data, add some current input data through three gate structures (input gate, forget gate, output gate), and finally integrate it into the current state and generate an output state. However, LSTM advances from left to right, making the subsequent data more important than the previous data. And through BiLSTM, bidirectional semantic information can be better captured.

[0134] In this embodiment, the caption text can be recognized by the CRNN algorithm, which combines CNN for image feature engineering and LSTM for sequential recognition, not only extracting robust features but also avoiding the extremely difficult single-character segmentation and single-character recognition in traditional algorithms through sequential recognition. At the same time, sequential recognition also embeds temporal dependencies (implicitly using the corpus).

[0135] Optionally, in some embodiments, before performing text recognition on the subtitle area in the video frame image, the recognized subtitle area in the video frame image can also be filtered to select a target subtitle area, and only the subtitles in the target subtitle area are subjected to text recognition.

[0136] Specifically, filtering the subtitle area in the video frame image can be to select a valid subtitle area from the recognized subtitle areas. Specifically, the subtitle area that meets the preset valid conditions can be determined as the valid subtitle area, and the valid subtitle area can be used as the target subtitle area to be text recognized. Among them, the preset valid conditions can be set according to the actual situation, and this embodiment does not limit this.

[0137] For example, the valid subtitle area can be comprehensively judged according to the rotation angle, aspect ratio, and position of the subtitle area. For example, the preset valid conditions can include: a) the height of the subtitle area is less than 28 pixel points and greater than 5 pixel points, b) the aspect ratio of the subtitle area is not less than 2, c) the rotation angle of the subtitle area does not exceed 15 degrees, d) the position of the subtitle area in the video frame image is not more than 5% of the height of the video frame image from the edge of the video frame image. The subtitle area that meets these 4 conditions at the same time is the valid subtitle area.

[0138] In some embodiments, for a video frame image, if multiple valid subtitle areas are recognized, the valid subtitle areas can be further filtered. Specifically, the valid subtitle area with the widest width can be selected as the target subtitle area. Then, the subtitles in the target subtitle area are recognized to obtain the subtitle text of the video frame image.

[0139] Through the video detection method provided in this embodiment, video frames of the target video segment of the video to be detected can be extracted and de-duplicated to obtain the video frame images of the target video segment; for each video frame image, the subtitle area of the video frame image is detected; then it is judged whether the detected subtitle area is a valid subtitle area to obtain the target subtitle area corresponding to the valid subtitle, and then the subtitles in the target subtitle area are recognized. Finally, the subtitle texts of each recognized video frame image are de-duplicated to obtain the subtitle text corresponding to the target video segment.

[0140] Optionally, in some embodiments, the subtitle text of the target video segment can also be extracted by an end-to-end OCR (Optical Character Recognition) method. For example, FOTS (Fast Oriented Text Spotting) can be used to directly extract the subtitle text from the video frame image without first detecting the subtitle area in the video frame image and then recognizing the subtitles in the subtitle area. FOTS is a fast end-to-end integrated detection and recognition framework, through which subtitle text can be extracted faster.

[0141] Optionally, in this embodiment, the step of "performing downsampling and upsampling processing on the video frame image at multiple scales to obtain the target feature map of the video frame image" may include:

[0142] Performing downsampling and upsampling processing on the video frame image at multiple scales through a subtitle area recognition model to obtain the target feature map of the video frame image;

[0143] The step of "performing a convolution operation on the target feature map to obtain the text unit heat map of the video frame image" includes:

[0144] Performing a convolution operation on the target feature map through a subtitle area recognition model to obtain the text unit heat map of the video frame image.

[0145] It should be noted that the subtitle area recognition model can be specifically trained by other devices and then provided to the video detection device, or it can also be trained by the video detection device itself.

[0146] If it is trained by the video detection device itself, before the step of "performing downsampling and upsampling processing on the video frame image at multiple scales through a subtitle area recognition model to obtain the target feature map of the video frame image", the video detection method may further include:

[0147] Obtaining training data, where the training data includes sample images and the target subtitle areas corresponding to the sample images;

[0148] Performing downsampling and upsampling processing on the sample image at multiple scales through a preset subtitle area recognition model to obtain the target feature map of the sample image;

[0149] Performing a convolution operation on the target feature map of the sample image to obtain the text unit heat map of the sample image;

[0150] Determining the reference subtitle area of the sample image based on the text unit heat map;

[0151] Based on the reference subtitle region and the target subtitle region, adjust the parameters of the preset subtitle region recognition model to obtain a subtitle region recognition model.

[0152] Among them, the step of "based on the reference subtitle region and the target subtitle region, adjusting the parameters of the preset subtitle region recognition model to obtain a subtitle region recognition model" can specifically be:

[0153] Calculate the degree of regional overlap between the reference subtitle region and the target subtitle region;

[0154] Based on the degree of regional overlap, adjust the parameters of the preset subtitle region recognition model to obtain a subtitle region recognition model.

[0155] Among them, the degree of regional overlap can specifically be represented by the intersection over union. The intersection over union (IoU, Intersection over Union) refers to the ratio of the intersection of two regions to the union, and the value is between [0, 1]. It can be used to represent the coincidence degree of two sets.

[0156] Among them, the training process can be to first calculate the degree of regional overlap between the reference subtitle region and the target subtitle region, and then use the backpropagation algorithm to adjust the parameters of the preset subtitle region recognition model. Based on the degree of regional overlap between the reference subtitle region and the target subtitle region, optimize the parameters of the preset subtitle region recognition model to make the degree of regional overlap between the reference subtitle region and the target subtitle region greater than the preset overlap degree, and obtain a trained subtitle region recognition model. Among them, the preset overlap degree can be set according to the actual situation, and this application has no limitation on this. For example, it can be set according to the requirements for subtitle region detection. If the requirements for subtitle region detection are higher, the preset overlap degree is higher.

[0157] 102. Perform speech recognition on the audio information to obtain the target text corresponding to the audio information, where the target text includes at least one text unit.

[0158] In this embodiment, the audio information can be converted into the corresponding target text through automatic speech recognition (ASR, Automatic Speech Recognition) technology. Specifically, the neural network can be used to perform speech recognition on the audio information. The neural network can be a convolutional neural network, a visual geometry group network, a residual network, a densely connected convolutional network, etc. However, it should be understood that the neural network in this embodiment is not limited to the above-listed several types. In addition, other machine learning methods can also be used to obtain the feature information corresponding to the audio information, such as various acoustic features such as Mel spectrogram and pitch, and it is not limited to the method of artificial neural network.

[0159] Optionally, in this embodiment, the step of "performing speech recognition on the audio information to obtain the target text corresponding to the audio information" may include:

[0160] Performing semantic extraction on the audio information to obtain the audio semantic feature information of the audio information;

[0161] Based on the audio semantic feature information, determining the prediction probability of translating the audio information into each candidate text;

[0162] Based on the prediction probability, determining the target text corresponding to the audio information from the candidate texts.

[0163] Among them, the step of "determining the target text corresponding to the audio information from the candidate texts based on the prediction probability" may include:

[0164] Taking the candidate text with the highest prediction probability as the target text corresponding to the audio information.

[0165] As Figure 1g shown, it is the network structure diagram of the speech recognition model. This speech recognition model can specifically be a fully convolutional speech recognition model, which can be divided into three parts, namely a learnable front end, an acoustic model, a language model, and a decoder.

[0166] Among them, the learnable front end can be used to process the original audio and extract some key features. Specifically, it can include a 2x1 convolutional module, a complex convolutional module with 40 filters, a squared absolute value operation module, a low-pass filter module, as well as a logarithmic compression module and a mean squared error normalization module. The acoustic model is a convolutional neural network with gated linear units. This network is fed with the output of the learnable front end. It can adopt the Auto Segmentation (ASG) standard and predict letters from the audio stream through training. Among them, the ASG standard is a loss function.

[0167] Among them, the language model can adopt GCNN-14B (Hierarchical Graph Convolutional Neural Network), generate candidate texts according to the input from the acoustic model, and then generate the final word sequence, that is, the target text, through a Beam-Search encoder.

[0168] 103. According to the text units of the target text and the text units of the subtitle text, performing text alignment on the target text and the subtitle text to obtain a text alignment result.

[0169] Among them, text alignment of the target text and the subtitle text can be solved by means of string matching.

[0170] In some embodiments, text alignment between the target text and the subtitle text can be performed through naive matching, or text alignment between the target text and the subtitle text can be performed through string fuzzy matching. This embodiment does not limit this, and specific methods can be adopted according to the actual situation.

[0171] Among them, fuzzy matching is to perform approximate matching according to the given requirements. It allows a certain number of characters / text units in the matching string to not match, that is, fuzzy matching has a certain error tolerance rate. The string matching in fuzzy matching, that is, the basic idea of SequenceMatcher (sequence matching) is to find the longest continuous matching subsequence that does not contain "garbage" elements; these "garbage" elements can be meaningless, such as blank lines or blanks; then, the same idea is recursively applied to the left subsequence and the right subsequence of the matching subsequence. For example, if the target text is acfgb and the subtitle text is gdrekacfab, the spaces in the subtitle text can be removed, and the subtitle text can be converted to gdrekacfab, and then it is matched with the target text. In fuzzy matching, the target text acfgb can be regarded as matching the acfab in the subtitle text. Acfgb and acfab are matching subsequences. That is to say, it allows some characters in the matching subsequence to not match (for example, the error tolerance rate can be 1 character not matching). Then, the unmatched subsequence in the subtitle text can be continued to be matched with the target text. Specifically, the left subsequence gdrek of the matching subsequence acfab in the subtitle text in the previous step is matched with the target text. Under the condition of an error tolerance rate of allowing 1 character not to match, the left subsequence still does not match the target text. Therefore, for the target text acfgb and the subtitle text gdrekacfab, their matching subsequence is acfab, and the non-matching subsequence is gdrek. Furthermore, based on the text alignment result (that is, the matching result), the integrity of the target video segment is judged. In some embodiments, if the non-matching subsequence of the target text and the subtitle text exceeds the preset length, the target video segment can be regarded as incomplete, and the preset length can be set according to the actual situation. In actual use, fuzzy matching can give the result that is closest to the correct match in the case of missed or misdetected calls in the OCR or ASR part.

[0172] Among them, naive matching is the naive string matching method, and its matching process can be as Figure 1h shown. Suppose the target text corresponding to the audio information is aab, and the subtitle text is acaabc. First, align the target text and the subtitle text based on the first character, as Figure 1hAs shown in (a) of [reference], in the first step, each character of the target text (aab) can be matched with the character at the corresponding position in the subtitle text. It can be seen from the figure that the second character 'a' of the target text fails to match the second character 'c' of the subtitle text (acaabc); in the second step: shift the target text one position to the right, as shown in (b) of [reference], and then repeat the steps in the first step. It can be seen from the figure that the first character fails to match; shift the target text one more position to the right and repeat the steps in the first step, as shown in (c) of [reference], until the matching is completed. It can be seen from the figure that when s = 2, that is, after the target text is shifted two positions to the right (s is the number of shifted positions), the target text 'aab' successfully matches the subsequence 'aab' in the subtitle text, and the matching subsequence 'aab' of the target text and the subtitle text, as well as the non-matching subsequences 'ac' (left subsequence) and 'c' (right subsequence) are obtained. Based on the matching result, the integrity of the target video segment is determined. In some embodiments, the integrity of the target video segment can be determined based on the length of the non-matching subsequence. Figure 1h As shown in (b) of [reference], repeat the steps in the first step again. It can be seen from the figure that the first character fails to match; shift the target text one more position to the right and repeat the steps in the first step, as shown in [reference], Figure 1h until the matching is completed. It can be seen from the figure that when s = 2, that is, after the target text is shifted two positions to the right (s is the number of shifted positions), the target text 'aab' successfully matches the subsequence 'aab' in the subtitle text, and the matching subsequence 'aab' of the target text and the subtitle text, as well as the non-matching subsequences 'ac' (left subsequence) and 'c' (right subsequence) are obtained. Based on the matching result, the integrity of the target video segment is determined. In some embodiments, the integrity of the target video segment can be determined based on the length of the non-matching subsequence.

[0173] Optionally, in this embodiment, the step of "performing text alignment on the target text and the subtitle text according to the text units of the target text and the text units of the subtitle text to obtain a text alignment result" may include:

[0174] Determine the target text unit at the target position in the target text;

[0175] Match the target text unit with the text units starting from the starting alignment position in the subtitle text;

[0176] When there is a text unit in the text units starting from the starting alignment position in the subtitle text that matches the target text unit, update the target position in the target text and the starting alignment position in the subtitle text, and update the target text unit to the text unit at the updated target position;

[0177] Return to execute the step of matching the target text unit with the text units starting from the starting alignment position in the subtitle text until all the text units in the target text are matched, and obtain a text alignment result.

[0178] Among them, when the target text and the subtitle text start to be matched, the target position in the target text can be the position of the first text unit in the target text, and the first text unit is the target text unit. The starting alignment position of the subtitle text is the position corresponding to the first text unit of the subtitle text. First, match the first text unit of the target text with each text unit of the subtitle text to find the text unit in the subtitle text that matches the first text unit of the target text. When there is a text unit in the subtitle text that matches the first text unit of the target text, the target position of the target text can be updated to the position corresponding to the second text unit in the target text. The second text unit is the target text unit. At the same time, the starting alignment position of the subtitle text is updated to the position corresponding to the adjacent text unit (i.e., the next text unit) of the text unit that matches the first text unit of the target text. Then, match the second text unit of the target text with the text units starting from the starting alignment position in the subtitle text, and so on.

[0179] 104. Detect the integrity of the target video segment according to the text alignment result.

[0180] Among them, the target video segment is a video segment in the video to be detected, which can specifically be the start segment, the end segment, or a composite video segment of the start segment and the end segment. This embodiment does not limit this. Determine the integrity of the video to be detected based on the integrity of the target video segment.

[0181] Among them, the text alignment result of the subtitle text and the target text can include the matching subsequence and the non-matching subsequence of the subtitle text and the target text. Specifically, by comparing the subtitle text and the target text in the start segment of the video to be detected, the matching subsequence and the non-matching subsequence (such as the left subsequence of the matching subsequence) of the subtitle text and the target text can be obtained. Based on the non-matching left subsequence, it can be determined whether the start segment of the video to be detected is complete. For example, if the non-matching left subsequence exceeds the preset length, it can be considered that the start segment is incomplete. In some embodiments, it can also be that when there is a non-matching subsequence, it is considered that the target video segment is incomplete. For example, if the subtitle text is "Sister, you're welcome" and the target text corresponding to the audio information is "You're welcome", it can be determined that there is a situation of missing words in the audio of the start segment, and the two words "Sister" in the subtitle text are not successfully matched, and this start segment is incomplete.

[0182] In some other embodiments, it can also be through the comparison of the subtitle text and the target text in the end segment of the video to be detected to obtain the matching subsequence and the non-matching subsequence (such as the right subsequence of the matching subsequence) of the subtitle text and the target text. Based on the non-matching right subsequence, it can be determined whether the end segment of the video to be detected is complete.

[0183] In the related technologies of current video integrity detection, there are mainly manual detection and audio (single-modal) based detection methods. Among them, the main problem of manual review is high cost. In the market with explosive growth of video content, the efficiency of manual review is low, and it is necessary to allocate a corresponding review team as the content volume increases. The audio-based detection method mainly extracts the audio track of the video, then generates the spectrogram corresponding to the audio track, statistically analyzes the frequency domain distribution, detects the voice position, and judges whether there is no voice duration at the beginning or end according to the threshold, so as to realize the detection of video integrity. However, in the audio of a large amount of content, there are usually a lot of background music, environmental noise, etc., and the voice environment is complex, and the effect still needs to be improved.

[0184] The video detection method provided by this application can automatically identify whether the beginning and end of the video are complete, quickly detect the problem of incomplete video content, and the detection process does not require manual participation, saving manpower and improving the review efficiency. At the same time, compared with the audio single-modal based detection method, the effect of detecting incomplete video content is more accurate.

[0185] The following table shows the comparison results of the video detection method of this application and the audio single-modal based detection method:

[0186]

[0187] Among them, F1 represents the comprehensive evaluation index, and the precision, recall rate, and F1 statistic the effect of the "incomplete" category.

[0188] The video detection method of this application can be applied to the integrity detection scenarios of video content such as UGC (User Generated Content, user-generated content, that is, user original content) and PGC (Professional Generated Content, professionally produced content). For example, it can be used in the background review process of video content publishing platforms to perform quality detection on videos, and can intercept or mark incomplete low-quality video content.

[0189] As can be seen from the above, this embodiment can obtain the audio information and subtitle text corresponding to the target video segment in the video to be detected, and the subtitle text includes at least one text unit; perform speech recognition on the audio information to obtain the target text corresponding to the audio information, and the target text includes at least one text unit; perform text alignment on the target text and the subtitle text according to the text units of the target text and the text units of the subtitle text to obtain a text alignment result; detect the integrity of the target video segment according to the text alignment result. The embodiment of this application can perform video content integrity detection based on the audio and subtitles of the target video segment, improve the efficiency of video integrity detection, and reduce the detection cost.

[0190] According to the method described in the previous embodiments, the following will further elaborate with the example that the video detection device is specifically integrated in the server.

[0191] An embodiment of the present application provides a video detection method. As Figure 2a shown, the specific process of this video detection method can be as follows:

[0192] 201. The server receives the video to be detected sent by the terminal.

[0193] Among them, the video to be detected is a video that needs to detect the integrity of the video content. In a specific scenario, the videos on the video playback platform need to be subject to quality review. Video integrity is an important item in quality review. Video incompleteness usually includes situations such as abrupt or incomplete beginnings or endings of the video, missing words, subtitles, and unfinished content.

[0194] 202. The server obtains the audio information corresponding to the target video segment in the video to be detected, and performs speech recognition on the audio information to obtain the target text corresponding to the audio information, where the target text includes at least one text unit.

[0195] Among them, the target video segment is a video segment in the video to be detected, and specifically can be the beginning segment, the ending segment, etc. of the video to be detected. In one embodiment, the beginning segment and the ending segment in the video to be detected can be used for video content integrity detection, and the integrity of the video to be detected can be determined based on the integrity of the beginning segment and the ending segment. For example, after performing video content integrity detection on the beginning segment based on the audio information and subtitle text corresponding to the beginning segment, if it is found that there is unfinished content in the subtitle of the beginning segment, it can be determined that the beginning segment is incomplete, that is, the video content of this video to be detected is incomplete.

[0196] Optionally, in this embodiment, the step of "performing speech recognition on the audio information to obtain the target text corresponding to the audio information" may include:

[0197] Performing semantic extraction on the audio information to obtain the audio semantic feature information of the audio information;

[0198] Based on the audio semantic feature information, determining the prediction probability of translating the audio information into each candidate text;

[0199] Based on the prediction probability, determining the target text corresponding to the audio information from the candidate texts.

[0200] 203. The server extracts video frames from the target video segment in the video to be detected to obtain at least one video frame image of the target video segment; by performing feature extraction on the video frame image, the subtitle area of the video frame image is determined.

[0201] Optionally, in this embodiment, the step of "determining the subtitle region of the video frame image by performing feature extraction on the video frame image" may include:

[0202] Performing downsampling and upsampling processing on the video frame image at multiple scales to obtain a target feature map of the video frame image;

[0203] Performing a convolution operation on the target feature map to obtain a text unit heat map of the video frame image;

[0204] Based on the text unit heat map, determining the subtitle region of the video frame image.

[0205] Among them, feature extraction can be performed on the video frame image through a neural network, and this neural network can be a convolutional neural network (CNN, Convolutional Neural Networks), a visual geometry group network (VGGNet, Visual Geometry Group Network), a residual network (ResNet, Residual Network), a dense connection convolutional network (DenseNet, Dense Convolutional Network), etc. However, it should be understood that the neural network in this embodiment is not limited to the several types listed above.

[0206] Optionally, in this embodiment, the step of "performing downsampling and upsampling processing on the video frame image at multiple scales to obtain a target feature map of the video frame image" may include:

[0207] Performing multiple downsampling processes on the video frame image to obtain downsampling feature maps of the video frame image at multiple scales;

[0208] Performing multiple upsampling processes on the downsampling feature map of the target scale to obtain upsampling fusion feature maps of the video frame image at multiple scales, where the input of the upsampling at each scale is a fusion feature obtained by fusing the upsampling feature map and the downsampling feature map of adjacent scales;

[0209] Determining the target feature map of the video frame image from the upsampling fusion feature maps of each scale.

[0210] Optionally, in this embodiment, the text unit heat map includes a character heat map and an inter-character heat map; the step of "based on the text unit heat map, determining the subtitle region of the video frame image" may include:

[0211] Selecting character regions from the character heat map according to the heat values of the heat points in the character heat map;

[0212] Select an inter - character region from the inter - character heat map according to the heat value of the heat points in the inter - character heat map;

[0213] Based on the character region and the inter - character region, determine the subtitle region of the video frame image.

[0214] 204. The server recognizes the subtitles in the subtitle region to obtain the subtitle text corresponding to the target video segment, and the subtitle text includes at least one text unit.

[0215] Optionally, in this embodiment, the step of "recognizing the subtitles in the subtitle region to obtain the subtitle text corresponding to the target video segment" may include:

[0216] Extract features from the subtitle region to obtain a feature sequence of the subtitle region, and the feature sequence includes at least one feature information;

[0217] Predict each feature information in the feature sequence according to the front - and - back feature information in the feature sequence to obtain the subtitle text corresponding to the target video segment.

[0218] Among them, the subtitles in the subtitle region can be recognized through a neural network, and the neural network can be a convolutional neural network, a recurrent neural network, a densely connected convolutional network, etc. However, it should be understood that the neural network in this embodiment is not limited to the above - listed types.

[0219] 205. The server aligns the target text and the subtitle text according to the text units of the target text and the text units of the subtitle text to obtain a text alignment result.

[0220] In some embodiments, the target text and the subtitle text can be text - aligned by naive matching or by string fuzzy matching. This embodiment does not limit this, and specific methods can be adopted according to the actual situation.

[0221] 206. The server detects the integrity of the target video segment according to the text alignment result.

[0222] Among them, the target video segment is a video segment in the video to be detected. Specifically, it can be a start segment, an end segment, or a composite video segment of the start segment and the end segment. This embodiment does not limit this. The integrity of the video to be detected is determined based on the integrity of the target video segment.

[0223] Among them, the text alignment result of the subtitle text and the target text may include the matching subsequence and the unmatched subsequence of the subtitle text and the target text. Specifically, by comparing the subtitle text and the target text in the start segment of the video to be detected, the matching subsequence and the unmatched subsequence (such as the left subsequence of the matching subsequence) of the subtitle text and the target text can be obtained. Based on the unmatched left subsequence, it can be determined whether the start segment of the video to be detected is complete. For example, if the unmatched left subsequence exceeds a preset length, it can be considered that the start segment is incomplete. In some embodiments, it can also be that when there is an unmatched subsequence, the target video segment is considered incomplete. For example, if the subtitle text is "Sister, you're welcome" and the target text corresponding to the audio information is "You're welcome", it can be determined that there is a situation of missing words in the audio of the start segment, and the two words "Sister" in the subtitle text are not successfully matched, and this start segment is incomplete.

[0224] In some other embodiments, it can also be by comparing the subtitle text and the target text in the end segment of the video to be detected, the matching subsequence and the unmatched subsequence (such as the right subsequence of the matching subsequence) of the subtitle text and the target text can be obtained. Based on the unmatched right subsequence, it can be determined whether the end segment of the video to be detected is complete.

[0225] The video detection method provided in this embodiment can identify the integrity of video content, which mainly includes audio text extraction, video subtitle extraction, matching of audio text and subtitle text, and an incomplete judgment method. As Figure 2b shown, it is the flowchart of video integrity detection, which is described in detail as follows:

[0226] 2001. Input the video to be detected and determine the target video segment of the video to be detected;

[0227] 2002. Translate the audio information corresponding to the target video segment to obtain the target text corresponding to the audio information;

[0228] 2003. Extract video frames from the target video segment to obtain at least one video frame image of the target video segment;

[0229] 2004. Identify the subtitle area of the video frame image;

[0230] 2005. Identify the subtitles in the subtitle area to obtain the subtitle text corresponding to the target video segment;

[0231] 2006. Perform aggregation processing on the subtitle text, that is, remove duplicates from the subtitle text;

[0232] 2007. Align the target text corresponding to the audio information and the aggregated subtitle text;

[0233] In 2008, based on the text alignment result, determine whether the video content of the target video segment is incomplete;

[0234] In 2009, based on the determination result of the integrity of the target video segment, determine and output the determination result of the integrity of the video to be detected.

[0235] In a specific embodiment, when identifying the subtitle area of a video frame image, for one video frame image, multiple subtitle areas may be identified; as Figure 2c shown, at this time, it is necessary to determine whether the identified subtitle area is a valid subtitle area. Specifically, the subtitle area that meets the preset valid conditions can be determined as the valid subtitle area, and the valid subtitle area can be used as the target subtitle area to be text-recognized. Among them, the preset valid conditions can be set according to the actual situation, and this embodiment does not limit this. For example, the preset valid conditions may include: a) the height of the subtitle area is less than 28 pixel points and greater than 5 pixel points, b) the aspect ratio of the length and width of the subtitle area is not less than 2, c) the rotation angle of the subtitle area does not exceed 15 degrees, d) the position of the subtitle area in the video frame image is not more than 5% of the height of the video frame image from the edge of the video frame image. The subtitle area that simultaneously meets these 4 conditions is the valid subtitle area.

[0236] Figure 2c In, subtitle area recognition is performed on the video frame image, and two candidate subtitle areas are obtained. It is necessary to select a valid subtitle area from the candidate subtitle areas and use the valid subtitle area as the subtitle area to be text-recognized. Specifically, it can be selected according to the preset valid conditions in the above embodiment.

[0237] As Figure 2d shown, when performing integrity detection on the target video segment of the video to be detected, video frames can be extracted from the target video segment to obtain multiple video frame images, and the subtitle text in the video frame images can be recognized; and speech recognition is performed on the audio information corresponding to the target video segment to obtain the target text corresponding to the audio information; then the subtitle text and the target text are text-aligned (i.e., matched), and based on the matching result, it is determined whether the video content of the target video segment is complete. For example, if the target text converted from the audio information of the target video segment is "My sister has a gift and wants to give it to my elder sister. My sister is being polite.", but the subtitle text recognized from the target video segment is "Wants to give it to my elder sister. My sister is being polite.", then there is a mismatched text between the target text and the subtitle text: "My sister has a gift", so the subtitle of this target video segment has unexpressed content and is incomplete.

[0238] As described above, in this embodiment, the server can receive the video to be detected sent by the terminal; obtain the audio information corresponding to the target video segment in the video to be detected, and perform speech recognition on the audio information to obtain the target text corresponding to the audio information, where the target text includes at least one text unit; extract video frames from the target video segment in the video to be detected to obtain at least one video frame image of the target video segment; determine the subtitle area of the video frame image by performing feature extraction on the video frame image; recognize the subtitles in the subtitle area to obtain the subtitle text corresponding to the target video segment, where the subtitle text includes at least one text unit; perform text alignment on the target text and the subtitle text according to the text units of the target text and the text units of the subtitle text to obtain a text alignment result; and detect the integrity of the target video segment according to the text alignment result. The embodiment of the present application can perform video content integrity detection based on the audio and subtitles of the target video segment, improve the efficiency of video integrity detection, and reduce the detection cost.

[0239] To better implement the above method, the embodiment of the present application further provides a video detection device, as Figure 3a shown. The video detection device may include an acquisition unit 301, an identification unit 302, an alignment unit 303, and a detection unit 304, as follows:

[0240] (1) Acquisition unit 301;

[0241] The acquisition unit 301 is configured to acquire the audio information and subtitle text corresponding to the target video segment in the video to be detected, where the subtitle text includes at least one text unit.

[0242] Optionally, in some embodiments of the present application, the acquisition unit 301 may include an acquisition subunit 3011, an extraction subunit 3012, an extraction subunit 3013, and an identification subunit 3014, as shown in Figure 3b and as follows:

[0243] The acquisition subunit 3011 is configured to acquire the audio information corresponding to the target video segment in the video to be detected;

[0244] The extraction subunit 3012 is configured to extract video frames from the target video segment in the video to be detected to obtain at least one video frame image of the target video segment;

[0245] The extraction subunit 3013 is configured to determine the subtitle area of the video frame image by performing feature extraction on the video frame image;

[0246] The identification subunit 3014 is configured to recognize the subtitles in the subtitle area to obtain the subtitle text corresponding to the target video segment.

[0247] Optionally, in some embodiments of the present application, the extraction subunit 3013 may specifically be configured to perform downsampling and upsampling processing on the video frame image at multiple scales to obtain a target feature map of the video frame image; perform a convolution operation on the target feature map to obtain a text unit heat map of the video frame image; and determine a subtitle area of the video frame image based on the text unit heat map.

[0248] Optionally, in some embodiments of the present application, the step of "performing downsampling and upsampling processing on the video frame image at multiple scales to obtain a target feature map of the video frame image" may include:

[0249] Performing multiple downsampling processes on the video frame image to obtain downsampled feature maps of the video frame image at multiple scales;

[0250] Performing multiple upsampling processes on the downsampled feature map of the target scale to obtain upsampled fusion feature maps of the video frame image at multiple scales, where the input for upsampling at each scale is a fusion feature obtained by fusing the upsampled feature map and the downsampled feature map of adjacent scales;

[0251] Determining a target feature map of the video frame image from the upsampled fusion feature maps of each scale.

[0252] Optionally, in some embodiments of the present application, the text unit heat map includes a character heat map and an inter-character heat map; the step of "determining a subtitle area of the video frame image based on the text unit heat map" may include:

[0253] Selecting character areas from the character heat map according to the heat values of the heat points in the character heat map;

[0254] Selecting inter-character areas from the inter-character heat map according to the heat values of the heat points in the inter-character heat map;

[0255] Determining a subtitle area of the video frame image based on the character areas and the inter-character areas.

[0256] Optionally, in some embodiments of the present application, the recognition subunit 3014 may specifically be configured to extract features from the subtitle area to obtain a feature sequence of the subtitle area, where the feature sequence includes at least one feature information; and predict each feature information in the feature sequence according to the front and back feature information in the feature sequence to obtain the subtitle text corresponding to the target video segment.

[0257] Optionally, in some embodiments of the present application, the step of "performing downsampling and upsampling processing on the video frame image at multiple scales to obtain the target feature map of the video frame image" may include:

[0258] Performing downsampling and upsampling processing on the video frame image at multiple scales through a subtitle area recognition model to obtain the target feature map of the video frame image;

[0259] The step of "performing a convolution operation on the target feature map to obtain the text unit heat map of the video frame image" may include:

[0260] Performing a convolution operation on the target feature map through a subtitle area recognition model to obtain the text unit heat map of the video frame image.

[0261] (2) Recognition unit 302;

[0262] The recognition unit 302 is configured to perform speech recognition on the audio information to obtain the target text corresponding to the audio information, and the target text includes at least one text unit.

[0263] Optionally, in some embodiments of the present application, the recognition unit 302 may include an audio extraction subunit 3021, a first determination subunit 3022, and a second determination subunit 3023, see Figure 3c , as follows:

[0264] The audio extraction subunit 3021 is configured to perform semantic extraction on the audio information to obtain the audio semantic feature information of the audio information;

[0265] The first determination subunit 3022 is configured to determine the prediction probability of translating the audio information into each candidate text based on the audio semantic feature information;

[0266] The second determination subunit 3023 is configured to determine the target text corresponding to the audio information from the candidate texts based on the prediction probability.

[0267] (3) Alignment unit 303;

[0268] The alignment unit 303 is configured to perform text alignment on the target text and the subtitle text according to the text units of the target text and the subtitle text to obtain a text alignment result.

[0269] Optionally, in some embodiments of the present application, the alignment unit 303 may include a third determination subunit 3031, a matching subunit 3032, an update subunit 3033, and a return subunit 3034, see Figure 3d , as follows:

[0270] The third determination subunit 3031 is configured to determine a target text unit at a target position in the target text;

[0271] The matching subunit 3032 is configured to match the target text unit with the text units starting from the starting alignment position in the subtitle text;

[0272] The updating subunit 3033 is configured to, when there is a text unit in the text units starting from the starting alignment position in the subtitle text that matches the target text unit, update the target position in the target text and the starting alignment position in the subtitle text, and update the target text unit to the text unit at the updated target position;

[0273] The returning subunit 3034 is configured to return to execute the step of matching the target text unit with the text units starting from the starting alignment position in the subtitle text until all the text units in the target text are matched, so as to obtain a text alignment result.

[0274] (4) The detection unit 304;

[0275] The detection unit 304 is configured to detect the integrity of the target video segment according to the text alignment result.

[0276] Optionally, in some embodiments of the present application, the video detection device may further include a training unit 305, see Figure 3e ; the training unit 305 is configured to train a subtitle area recognition model. Specifically, the training unit 305 may be configured to:

[0277] Obtain training data, where the training data includes a sample image and a target subtitle area corresponding to the sample image;

[0278] Perform downsampling and upsampling processing on the sample image at multiple scales through a preset subtitle area recognition model to obtain a target feature map of the sample image;

[0279] Perform a convolution operation on the target feature map of the sample image to obtain a text unit heat map of the sample image;

[0280] Determine a reference subtitle area of the sample image based on the text unit heat map;

[0281] Adjust the parameters of the preset subtitle area recognition model based on the reference subtitle area and the target subtitle area to obtain a subtitle area recognition model.

[0282] As can be seen from the above, in this embodiment, the acquisition unit 301 can acquire the audio information and subtitle text corresponding to the target video segment in the video to be detected, where the subtitle text includes at least one text unit; the recognition unit 302 performs speech recognition on the audio information to obtain the target text corresponding to the audio information, where the target text includes at least one text unit; the alignment unit 303 aligns the target text and the subtitle text according to the text units of the target text and the text units of the subtitle text to obtain a text alignment result; the detection unit 304 detects the integrity of the target video segment according to the text alignment result. The embodiment of the present application can perform video content integrity detection based on the audio and subtitles of the target video segment, improve the efficiency of video integrity detection, and reduce the detection cost.

[0283] The embodiment of the present application also provides an electronic device, as Figure 4 shown, which shows a schematic structural diagram of the electronic device involved in the embodiment of the present application. Specifically:

[0284] The electronic device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404, and other components. Those skilled in the art can understand that Figure 4 the structural diagram of the electronic device shown in

[0285] does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Among them:

[0286] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device. In addition, the memory 402 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 can also include a memory controller to provide the processor 401 with access to the memory 402.

[0287] The electronic device further includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0288] The electronic device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0289] Although not shown, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:

[0290] Obtain the audio information and subtitle text corresponding to the target video segment in the video to be detected, where the subtitle text includes at least one text unit; perform speech recognition on the audio information to obtain the target text corresponding to the audio information, where the target text includes at least one text unit; perform text alignment on the target text and the subtitle text according to the text units of the target text and the text units of the subtitle text to obtain a text alignment result; detect the integrity of the target video segment according to the text alignment result.

[0291] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.

[0292] As can be seen from the above, in this embodiment, the audio information and subtitle text corresponding to the target video segment in the video to be detected can be obtained, and the subtitle text includes at least one text unit; speech recognition is performed on the audio information to obtain the target text corresponding to the audio information, and the target text includes at least one text unit; according to the text units of the target text and the text units of the subtitle text, text alignment is performed on the target text and the subtitle text to obtain a text alignment result; according to the text alignment result, the integrity of the target video segment is detected. The embodiments of the present application can perform video content integrity detection based on the audio and subtitles of the target video segment, improve the efficiency of video integrity detection, and reduce the detection cost.

[0293] Those of ordinary skill in the art can understand that all or part of the steps in the above-mentioned various methods can be completed by instructions, or by controlling relevant hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0294] For this reason, the embodiments of the present application provide a storage medium, in which multiple instructions are stored, and these instructions can be loaded by a processor to execute the steps in any one of the video detection methods provided by the embodiments of the present application. For example, these instructions can execute the following steps:

[0295] Obtain the audio information and subtitle text corresponding to the target video segment in the video to be detected, where the subtitle text includes at least one text unit; perform speech recognition on the audio information to obtain the target text corresponding to the audio information, where the target text includes at least one text unit; according to the text units of the target text and the text units of the subtitle text, perform text alignment on the target text and the subtitle text to obtain a text alignment result; according to the text alignment result, detect the integrity of the target video segment.

[0296] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.

[0297] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0298] Since the instructions stored in this storage medium can execute the steps in any one of the video detection methods provided by the embodiments of the present application, the beneficial effects that can be achieved by any one of the video detection methods provided by the embodiments of the present application can be realized. For details, refer to the previous embodiments, which will not be elaborated here.

[0299] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations of the above video detection aspect.

[0300] The above has introduced in detail a video detection method, device, electronic device and storage medium provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A video detection method, characterized in that, comprising: obtaining audio information and subtitle text corresponding to a target video segment in a video to be detected, where the subtitle text includes at least one text unit; the target video segment is the start segment and the end segment of the video to be detected; performing speech recognition on the audio information to obtain target text corresponding to the audio information, where the target text includes at least one text unit; performing text alignment on the target text and the subtitle text according to the text units of the target text and the text units of the subtitle text to obtain a text alignment result; detecting the integrity of the target video segment according to the text alignment result.

2. The method according to claim 1, characterized in that, the obtaining audio information and subtitle text corresponding to a target video segment in a video to be detected includes: obtaining audio information corresponding to a target video segment in a video to be detected; extracting video frames from the target video segment in the video to be detected to obtain at least one video frame image of the target video segment; determining a subtitle area of the video frame image by performing feature extraction on the video frame image; recognizing subtitles in the subtitle area to obtain subtitle text corresponding to the target video segment.

3. The method according to claim 2, characterized in that, the determining a subtitle area of the video frame image by performing feature extraction on the video frame image includes: performing downsampling and upsampling processing on the video frame image at multiple scales to obtain a target feature map of the video frame image; performing a convolution operation on the target feature map to obtain a text unit heat map of the video frame image; determining a subtitle area of the video frame image based on the text unit heat map.

4. The method according to claim 3, characterized in that, the performing downsampling and upsampling processing on the video frame image at multiple scales to obtain a target feature map of the video frame image includes: performing multiple downsampling processes on the video frame image to obtain downsampled feature maps of the video frame image at multiple scales; performing multiple upsampling processes on the downsampled feature map of a target scale to obtain upsampled fusion feature maps of the video frame image at multiple scales, where the input for upsampling at each scale is a fusion feature obtained by fusing the upsampled feature map and the downsampled feature map of adjacent scales; determining a target feature map of the video frame image from the upsampled fusion feature maps of each scale.

5. The method according to claim 3, characterized in that, the text unit heat map includes a character heat map and an inter-character heat map; the determining a subtitle area of the video frame image based on the text unit heat map includes: selecting a character area from the character heat map according to the heat value of heat points in the character heat map; selecting an inter-character area from the inter-character heat map according to the heat value of heat points in the inter-character heat map; determining a subtitle area of the video frame image based on the character area and the inter-character area.

6. The method according to claim 2, characterized in that, Identifying the subtitles in the subtitle area to obtain the subtitle text corresponding to the target video segment includes: Extracting features from the subtitle area to obtain a feature sequence of the subtitle area, where the feature sequence includes at least one feature information; Predicting each feature information in the feature sequence according to the front and back feature information in the feature sequence to obtain the subtitle text corresponding to the target video segment.

7. The method according to claim 1, wherein, Performing speech recognition on the audio information to obtain the target text corresponding to the audio information includes: Performing semantic extraction on the audio information to obtain the audio semantic feature information of the audio information; Based on the audio semantic feature information, determining the prediction probability of translating the audio information into each candidate text; Based on the prediction probability, determining the target text corresponding to the audio information from the candidate texts.

8. The method according to claim 3, wherein, Performing downsampling and upsampling processing on the video frame image at multiple scales to obtain the target feature map of the video frame image includes: Performing downsampling and upsampling processing on the video frame image at multiple scales through a subtitle area recognition model to obtain the target feature map of the video frame image; Performing a convolution operation on the target feature map to obtain the text unit heat map of the video frame image includes: Performing a convolution operation on the target feature map through a subtitle area recognition model to obtain the text unit heat map of the video frame image.

9. The method according to claim 8, wherein, Before performing downsampling and upsampling processing on the video frame image at multiple scales through a subtitle area recognition model to obtain the target feature map of the video frame image, it further includes: Obtaining training data, where the training data includes sample images and the target subtitle areas corresponding to the sample images; Performing downsampling and upsampling processing on the sample images at multiple scales through a preset subtitle area recognition model to obtain the target feature map of the sample images; Performing a convolution operation on the target feature map of the sample images to obtain the text unit heat map of the sample images; Based on the text unit heat map, determining the reference subtitle area of the sample images; Based on the reference subtitle area and the target subtitle area, adjusting the parameters of the preset subtitle area recognition model to obtain a subtitle area recognition model.

10. The method according to claim 1, wherein, Aligning the target text and the subtitle text according to the text units of the target text and the text units of the subtitle text to obtain a text alignment result includes: Determining the target text unit at the target position in the target text; Matching the target text unit with the text units starting from the starting alignment position in the subtitle text; When there is a text unit in the text units starting from the starting alignment position in the subtitle text that matches the target text unit, update the target position in the target text and the starting alignment position in the subtitle text, and update the target text unit to the text unit at the updated target position; Return to execute the step of matching the target text unit with the text units starting from the starting alignment position in the subtitle text until all text units in the target text are matched to obtain a text alignment result.

11. A video detection device, characterized in that, comprising: an acquisition unit, configured to acquire audio information and subtitle text corresponding to a target video segment in a video to be detected, where the subtitle text includes at least one text unit; the target video segment is the start segment and the end segment of the video to be detected; an identification unit, configured to perform speech recognition on the audio information to obtain a target text corresponding to the audio information, where the target text includes at least one text unit; an alignment unit, configured to perform text alignment on the target text and the subtitle text according to the text units of the target text and the text units of the subtitle text to obtain a text alignment result; a detection unit, configured to detect the integrity of the target video segment according to the text alignment result.

12. The device according to claim 11, characterized in that, the acquisition unit includes an acquisition subunit, an extraction subunit, an extraction subunit and an identification subunit: the acquisition subunit is configured to acquire audio information corresponding to a target video segment in a video to be detected; the extraction subunit is configured to extract video frames of the target video segment in the video to be detected to obtain at least one video frame image of the target video segment; the extraction subunit is configured to determine a subtitle area of the video frame image by performing feature extraction on the video frame image; the identification subunit is configured to identify subtitles in the subtitle area to obtain subtitle text corresponding to the target video segment.

13. The device according to claim 12, characterized in that, the extraction subunit is specifically configured to perform downsampling and upsampling processing on the video frame image at multiple scales to obtain a target feature map of the video frame image; perform a convolution operation on the target feature map to obtain a text unit heat map of the video frame image; and determine the subtitle area of the video frame image based on the text unit heat map.

14. The device according to claim 13, characterized in that, the performing downsampling and upsampling processing on the video frame image at multiple scales to obtain a target feature map of the video frame image includes: performing multiple downsampling processes on the video frame image to obtain downsampling feature maps of the video frame image at multiple scales; performing multiple upsampling processes on the downsampling feature map of the target scale to obtain upsampling fusion feature maps of the video frame image at multiple scales, where the input of the upsampling at each scale is a fusion feature obtained by fusing the upsampling feature map and the downsampling feature map of adjacent scales; Determine the target feature map of the video frame image from the upsampled and fused feature maps at various scales.

15. The apparatus according to claim 13, wherein, the text unit heat map includes a character heat map and an inter-character heat map; determining the subtitle area of the video frame image based on the text unit heat map includes: Select character areas from the character heat map according to the heat values of the heat points in the character heat map; Select inter-character areas from the inter-character heat map according to the heat values of the heat points in the inter-character heat map; Based on the character areas and the inter-character areas, determine the subtitle area of the video frame image.

16. The apparatus according to claim 12, wherein, the recognition sub-unit is specifically configured to extract features from the subtitle area to obtain a feature sequence of the subtitle area, the feature sequence includes at least one feature information; according to the front and back feature information in the feature sequence, predict each feature information in the feature sequence to obtain the subtitle text corresponding to the target video segment.

17. The apparatus according to claim 11, wherein, the recognition unit includes an audio extraction sub-unit, a first determination sub-unit and a second determination sub-unit, as follows: The audio extraction sub-unit is configured to perform semantic extraction on the audio information to obtain audio semantic feature information of the audio information; The first determination sub-unit is configured to determine the prediction probability of translating the audio information into each candidate text based on the audio semantic feature information; The second determination sub-unit is configured to determine the target text corresponding to the audio information from the candidate texts based on the prediction probability.

18. The apparatus according to claim 13, wherein, performing downsampling and upsampling processing on the video frame image at multiple scales to obtain the target feature map of the video frame image includes: Performing downsampling and upsampling processing on the video frame image at multiple scales through a subtitle area recognition model to obtain the target feature map of the video frame image; Performing a convolution operation on the target feature map to obtain the text unit heat map of the video frame image includes: Performing a convolution operation on the target feature map through a subtitle area recognition model to obtain the text unit heat map of the video frame image.

19. The apparatus according to claim 18, wherein, before performing downsampling and upsampling processing on the video frame image at multiple scales through a subtitle area recognition model to obtain the target feature map of the video frame image, there is also a training unit, and the training unit is used to train the subtitle area recognition model, and the training unit is used for: Obtain training data, the training data includes sample images and the target subtitle areas corresponding to the sample images; Performing downsampling and upsampling processing on the sample images at multiple scales through a preset subtitle area recognition model to obtain the target feature map of the sample images; Performing a convolution operation on the target feature map of the sample images to obtain the text unit heat map of the sample images; Determine a reference caption area of the sample image based on the text unit heat map; Adjust parameters of the preset caption area recognition model based on the reference caption area and the target caption area to obtain a caption area recognition model.

20. The apparatus according to claim 11, wherein, the alignment unit includes a third determination subunit, a matching subunit, an update subunit, and a return subunit, as follows: the third determination subunit is configured to determine a target text unit at a target position in the target text; the matching subunit is configured to match the target text unit with text units starting from the starting alignment position in the caption text; the update subunit is configured to, when there is a text unit in the text units starting from the starting alignment position in the caption text that matches the target text unit, update the target position in the target text and the starting alignment position in the caption text, and update the target text unit to the text unit at the updated target position; the return subunit is configured to return to perform the step of matching the target text unit with the text units starting from the starting alignment position in the caption text until all text units in the target text are matched to obtain a text alignment result.

21. An electronic device, wherein, it includes a memory and a processor; the memory stores an application program, and the processor is configured to run the application program in the memory to execute the operations in the video detection method according to any one of claims 1 to 10.

22. A storage medium, wherein, the storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the video detection method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method and apparatus for determining video segment

    CN107027060A

  • Monitoring audio-visual content with captions

    CN107306342A