Video scene identification method and device, and electronic equipment

By extracting the depth map of video frames and using pre-trained CNN models for video scene recognition, the problems of high complexity and low accuracy in the prior art are solved, and fast and accurate video scene recognition is achieved.

CN120451855APending Publication Date: 2025-08-08WONDERSHARE TECH (HUNAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510382083.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing video scene recognition methods mainly rely on the convolutional neural network of human imaging area or original input, and cannot effectively process non-human-dominated videos or complex scenes, resulting in high recognition complexity and low accuracy.

Method used

By extracting the depth map of the video frame and using the pre-trained convolutional neural network model for scene type recognition, filtering out factors that are independent of distance, combining the accuracy of the depth map technology and the efficiency of the CNN model, fast and accurate scene type recognition is achieved.

Benefits of technology

It reduces the complexity of scene recognition, improves the recognition accuracy, and can accurately identify scene types in the video, such as close-up, close-up, medium-scene, long-range, etc.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451855A_ABST
    Figure CN120451855A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video scene identification method. The method comprises the following steps: extracting a depth map of a video frame of a video; and based on the extracted depth map of the video frame, using a pre-trained convolutional neural network model to identify the scene of the video. According to the video scene identification method provided by the embodiment of the invention, scene identification is carried out based on the depth information of the video, factors irrelevant to the distance are filtered out, the complexity of scene identification is reduced, and meanwhile, the precision of scene identification can be improved. The embodiment of the invention further provides a video scene recognition device and electronic equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of video processing technology, and more specifically, to a video scene recognition method, device, and electronic device. Background Art

[0002] AI technology is widely used in video algorithms, such as AI-based video content retrieval, AI-based video content generation, and AI-based automatic video editing. These algorithms require video scene recognition to provide basic support. For example, video scene analysis algorithms are required for video content retrieval based on user expectations of scene sizes.

[0003] However, existing video scene recognition methods mainly rely on convolutional neural networks with human imaging area or raw input, which have limitations when processing non-human-dominated videos or complex scenes. Summary of the Invention

[0004] In response to the problems existing in the above-mentioned prior art, the embodiments of the present application provide a video scene recognition method, device, and electronic device, which perform scene recognition based only on the depth information of the video, filter out factors unrelated to distance, reduce the complexity of scene recognition, and at the same time improve the accuracy of scene recognition.

[0005] In a first aspect, an embodiment of the present application provides a method for video scene recognition, comprising the following steps:

[0006] extracting a depth map of a video frame of the video; and

[0007] Based on the extracted depth map of the video frame, a pre-trained convolutional neural network model is used to identify the scene of the video.

[0008] Furthermore, extracting the depth map of the video frame of the video includes:

[0009] Input the video frames of the video into the depth estimation model frame by frame to obtain relative depth information of each frame;

[0010] Inputting the video frames into the subject recognition model frame by frame to obtain subject target information of each frame; and

[0011] The relative depth information is corrected using the subject target information to obtain a depth map of the video frame.

[0012] Furthermore, after inputting the video frames of the video into the depth estimation model frame by frame to obtain the relative depth information of each frame, the method further includes:

[0013] Normalization is performed on the relative depth information.

[0014] Furthermore, after the relative depth information is corrected by the subject target information to obtain the depth map of the video frame, the method further includes:

[0015] A center region cropping operation is performed on the depth map to remove edge information of the depth map.

[0016] Furthermore, the pre-training process of the convolutional neural network model includes:

[0017] Extracting the video frames according to a specific frame rate and extracting depth maps of the video frames as a training sample set; and

[0018] Construct a convolutional neural network model, and input the training sample set to train the convolutional neural network model.

[0019] Furthermore, before extracting the depth map of the video frame of the video, the method further includes:

[0020] The video is decoded into video frames, and sizes of the video frames are unified.

[0021] Furthermore, after identifying the scene of the video based on the extracted depth map of the video frame using a pre-trained convolutional neural network model, the method further includes:

[0022] The video scene recognition results of the video frames are summarized and the video scene recognition results at the timestamp level are output.

[0023] In a second aspect, an embodiment of the present application further provides a video scene recognition device, comprising:

[0024] a depth map extraction module, configured to extract a depth map of a video frame of the video; and

[0025] The scene recognition module is used to identify the scene of the video based on the extracted depth map of the video frame using a pre-trained convolutional neural network model.

[0026] In a third aspect, an embodiment of the present application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the video scene recognition method according to the first aspect above when executing the program.

[0027] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is used to implement the video scene recognition method according to the first aspect above.

[0028] The embodiments of the present application bring the following beneficial effects:

[0029] In the video scene recognition method provided in the embodiment of the present application, it is first necessary to extract a depth map of each video frame from the video. The depth map reveals the relative distance between objects in the scene and the camera, providing key information for scene recognition. Secondly, it is necessary to use a pre-trained convolutional neural network (CNN) model to analyze the extracted depth map and accurately identify the scene of the video, such as close-up, near shot, medium shot, and long shot. The video scene recognition method provided in the embodiment of the present application combines the accuracy of depth map technology and the efficiency of CNN models to achieve fast and accurate recognition of video scenes. At the same time, scene recognition is performed only based on the depth information of the video, filtering out factors unrelated to distance and reducing the complexity of scene recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0031] Figure 1 A schematic diagram of the process framework of the video scene recognition method provided in an embodiment of the present application;

[0032] Figure 2 A structural block diagram of a video scene recognition device provided in an embodiment of the present application;

[0033] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0034] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments described in this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0036] In the specification and claims of this application and the above-mentioned drawings, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to the specific circumstances.

[0037] Reference Figure 1 , Figure 1 FIG. 1 is a flow chart of a video scene recognition method according to an embodiment of the present application. Figure 1 As shown, the video scene recognition method of the embodiment of the present application includes the following steps:

[0038] S101: extracting a depth map of a video frame of the video; and

[0039] S102: Based on the extracted depth map of the video frame, a pre-trained convolutional neural network model is used to identify the scene of the video.

[0040] In the video scene recognition method provided in the embodiment of the present application, it is first necessary to extract a depth map of each video frame from the video. The depth map reveals the relative distance between objects in the scene and the camera, providing key information for scene recognition. Secondly, it is necessary to use a pre-trained convolutional neural network (CNN) model to analyze the extracted depth map and accurately identify the scene of the video, such as close-up, near shot, medium shot, and long shot. The video scene recognition method provided in the embodiment of the present application combines the accuracy of depth map technology and the efficiency of CNN models to achieve fast and accurate recognition of video scenes. At the same time, scene recognition is performed only based on the depth information of the video, filtering out factors unrelated to distance and reducing the complexity of scene recognition.

[0041] Furthermore, in some embodiments of the present application, extracting a depth map of a video frame of the video includes:

[0042] Input the video frames of the video into the depth estimation model frame by frame to obtain relative depth information of each frame;

[0043] Inputting the video frames into the subject recognition model frame by frame to obtain subject target information of each frame; and

[0044] The relative depth information is corrected using the subject target information to obtain a depth map of the video frame.

[0045] That is, first, the video frames of the video are input into the depth estimation model frame by frame. The model can analyze the content of each frame and output the corresponding relative depth information. This information reflects the relative distance between different parts of the picture and the camera. Then, in order to further improve the accuracy of the depth information, the video frames are input into the subject recognition model frame by frame again. The model can identify the main target in the picture, such as people, animals or important objects, and output the main target information. Finally, the relative depth information is corrected using the main target information. The purpose of this step is to ensure that the depth map can more accurately reflect the depth relationship between the main target and the background, thereby improving the accuracy of subsequent scene recognition. Through this sophisticated depth map extraction process, this method can more effectively utilize the video content and provide more reliable data support for subsequent scene recognition.

[0046] Furthermore, in some embodiments of the present application, after inputting the video frames of the video into the depth estimation model frame by frame to obtain the relative depth information of each frame, the method further includes:

[0047] Normalization is performed on the relative depth information.

[0048] Specifically, normalization is a data preprocessing method that aims to scale the data range to a specific interval (such as 0 to 1) to facilitate subsequent processing and analysis. In the embodiments of the present application, normalization can ensure that the relative depth information remains numerically consistent, avoiding differences in depth value ranges caused by different video frames or different scenes, which may affect the accuracy of subsequent scene recognition.

[0049] Furthermore, in some embodiments of the present application, after the relative depth information is corrected using the subject target information to obtain the depth map of the video frame, the method further includes:

[0050] A center region cropping operation is performed on the depth map to remove edge information of the depth map.

[0051] Specifically, edge information often contains noise or unimportant background details, which can interfere with subsequent scene recognition. Center region cropping allows us to focus on the important area at the center of the image, improving the quality of the depth map and the accuracy of subsequent scene recognition.

[0052] Furthermore, in some embodiments of the present application, the pre-training process of the convolutional neural network model includes:

[0053] Extracting the video frames according to a specific frame rate and extracting depth maps of the video frames as a training sample set; and

[0054] Construct a convolutional neural network model, and input the training sample set to train the convolutional neural network model.

[0055] Specifically, in the video scene recognition method provided in the embodiments of this application, pre-training of a convolutional neural network (CNN) model is a key step. This process first extracts video frames from the video at a specific frame rate and extracts a depth map for each frame to form a training sample set. The depth map reflects the relative distance between objects in the scene and the camera and is crucial for scene recognition.

[0056] Next, a CNN model is constructed and trained on the training set. Through multiple layers of convolution and pooling, the CNN model learns to extract effective features from the depth map and classify images based on these features. During training, the model continuously adjusts its parameters to minimize prediction error. After numerous iterations of training, the CNN model is able to accurately identify different scene types. This pre-training process not only improves the model's recognition accuracy but also ensures its stability and efficiency when processing new videos. The trained CNN model can then be used in real-world video scene recognition tasks, quickly outputting recognition results.

[0057] Furthermore, in some embodiments of the present application, before extracting the depth map of the video frame of the video, the method includes:

[0058] The video is decoded into video frames, and sizes of the video frames are unified.

[0059] Specifically, before extracting the depth map of the video frame, the video needs to be preprocessed to ensure the smooth progress of subsequent steps. That is, the video is decoded and split into a series of independent video frames. This step is the basis of video processing, which enables the video content to be analyzed frame by frame, making it possible for subsequent depth map extraction and scene type identification. Next, in order to maintain consistency and efficiency in processing, the size of all video frames needs to be unified. Since different videos may be recorded at different resolutions, the size of the video frames may vary. Unifying the size not only ensures that the depth algorithm can be applied correctly, but also improves processing speed and accuracy. In the process of unifying the size, a suitable resolution is usually selected as the standard, and then other video frames are adjusted to this resolution through scaling, cropping, and other methods.

[0060] Furthermore, in some embodiments of the present application, after identifying the scene of the video based on the extracted depth map of the video frame using a pre-trained convolutional neural network model, the method further includes:

[0061] The video scene recognition results of the video frames are summarized and the video scene recognition results at the timestamp level are output.

[0062] Specifically, after analyzing the video's scene types using a pre-trained convolutional neural network model based on the extracted depth map of the video frame, the analysis results need to be summarized and organized. Specifically, the scene types of each video frame are arranged in chronological order and each scene type is marked with a corresponding timestamp to clearly display the changes in scene types in the video and the time of their occurrence.

[0063] Finally, the output is the video shot type analysis result including timestamps. This result not only intuitively shows the distribution of various shots in the video, but also provides the exact time information of their appearance, providing strong support for video editing, special effects addition and other fields, while also providing data resources for subsequent applications and research.

[0064] Figure 2 FIG is a structural block diagram of the video scene recognition device 200 provided in an embodiment of the present application. Figure 2 As shown, the video scene recognition device 200 of the embodiment of the present application includes: a depth map extraction module 210 and a scene recognition module 220, wherein:

[0065] A depth map extraction module 210 for extracting a depth map of a video frame of the video; and

[0066] The scene recognition module 220 is configured to recognize the scene of the video based on the extracted depth map of the video frame using a pre-trained convolutional neural network model.

[0067] In the video scene recognition device provided in the embodiment of the present application, it is first necessary to extract a depth map of each video frame from the video. The depth map reveals the relative distance between objects in the scene and the camera, providing key information for scene recognition. Secondly, it is necessary to use a pre-trained convolutional neural network (CNN) model to analyze the extracted depth map and accurately identify the scene of the video, such as close-up, near shot, medium shot, and long shot. The video scene recognition device provided in the embodiment of the present application combines the accuracy of depth map technology and the efficiency of CNN models to achieve fast and accurate recognition of video scenes. At the same time, scene recognition is performed only based on the depth information of the video, filtering out factors unrelated to distance and reducing the complexity of scene recognition.

[0068] It should be noted that the specific implementation of the video scene recognition device in the embodiment of the present application is similar to the specific implementation of the video scene recognition method in the embodiment of the present application. Please refer to the description of the method part for details and will not be repeated here.

[0069] Figure 3 Schematic diagram of the structure of an electronic device 300 according to an embodiment of the present application.

[0070] like Figure 3As shown, the electronic device 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage part 302 to the random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0071] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, and the like; an output section 307 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 308 including a hard disk; and a communication section 309 including a network interface card such as a LAN card or a modem. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. Removable media 311, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 310 as needed, so that computer programs read therefrom can be installed into the storage section 308 as needed.

[0072] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 309, and / or installed from a removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, the above-mentioned functions defined in the electronic device of the present application are executed.

[0073] It should be noted that the computer-readable medium shown in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electronic device, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0074] In this application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction-executing electronic device, apparatus, or device. Furthermore, in this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction-executing electronic device, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination thereof.

[0075] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the processing receiving device, method and computer program product according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the aforementioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based electronic device that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0076] The units or modules involved in the embodiments described in this application can be implemented by software or hardware. The units or modules described can also be set in a processor, and the processor is used to implement the video scene recognition method when executing the program:

[0077] extracting a depth map of a video frame of the video; and

[0078] Based on the extracted depth map of the video frame, a pre-trained convolutional neural network model is used to identify the scene of the video.

[0079] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently and not be incorporated into the electronic device. The computer-readable storage medium stores one or more programs, and when the aforementioned programs are used by one or more processors to execute the video scene recognition method described in the present application:

[0080] extracting a depth map of a video frame of the video; and

[0081] Based on the extracted depth map of the video frame, a pre-trained convolutional neural network model is used to identify the scene of the video.

[0082] As another aspect, the present application further provides a computer program product, which may be included in the electronic device described in the above embodiment; or may exist independently and not be incorporated into the electronic device. The above computer program product stores one or more programs, and when the above programs are used by one or more processors to execute the video scene recognition method described in the present application:

[0083] extracting a depth map of a video frame of the video; and

[0084] Based on the extracted depth map of the video frame, a pre-trained convolutional neural network model is used to identify the scene of the video.

[0085] The above description is only a preferred embodiment of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made based on the contents of the present application specification and drawings, or direct / indirect application in other related technical fields, are included in the patent protection scope of the present application.

Claims

1. A video scene recognition method, characterized in that: The following steps are involved: extracting a depth map of a video frame of the video; and Based on the extracted depth map of the video frame, a pre-trained convolutional neural network model is used to identify the scene of the video.

2. The video scene recognition method according to claim 1, characterized in that: The extracting a depth map of a video frame of the video includes: Input the video frames of the video into the depth estimation model frame by frame to obtain relative depth information of each frame; Inputting the video frames into the subject recognition model frame by frame to obtain subject target information of each frame; and The relative depth information is corrected using the subject target information to obtain a depth map of the video frame.

3. The video scene recognition method according to claim 1, characterized in that: After inputting the video frames of the video into the depth estimation model frame by frame to obtain the relative depth information of each frame, the method further includes: Normalization is performed on the relative depth information.

4. The video scene recognition method according to claim 3, characterized in that: After the relative depth information is corrected using the subject target information to obtain the depth map of the video frame, the method further includes: A center region cropping operation is performed on the depth map to remove edge information of the depth map.

5. The video scene recognition method according to claim 1, characterized in that: The pre-training process of the convolutional neural network model includes: Extracting the video frames according to a specific frame rate and extracting depth maps of the video frames as a training sample set; and Construct a convolutional neural network model, and input the training sample set to train the convolutional neural network model.

6. The video scene recognition method according to claim 1, characterized in that: Before extracting the depth map of the video frame of the video, the method includes: The video is decoded into video frames, and sizes of the video frames are unified.

7. The video scene recognition method according to claim 1, characterized in that: After identifying the scene of the video based on the extracted depth map of the video frame using a pre-trained convolutional neural network model, the method further includes: The video scene recognition results of the video frames are summarized and the video scene recognition results at the timestamp level are output.

8. A video scene recognition device, characterized in that: include: A depth map extraction module, configured to extract a depth map of a video frame of the video; and The scene recognition module is used to identify the scene of the video based on the extracted depth map of the video frame using a pre-trained convolutional neural network model.

9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the video scene recognition method according to any one of claims 1 to 7 when executing the program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is used to implement the video scene recognition method according to any one of claims 1-7.