A video enhancement method and apparatus
By performing semantic analysis and high-dimensional feature extraction on video frames, personalized video enhancement parameters are determined, which solves the problem of poor enhancement effect in existing technologies and achieves efficient enhancement effect of video frame content adaptation.
Patent Information
- Application Number
- CN202510929288.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-07
AI Technical Summary
Existing video enhancement methods based on fixed rules lack flexibility, resulting in poor enhancement effects and an inability to adapt to the diversity of video frame content.
By performing semantic analysis on video frames, high-dimensional semantic features are extracted. Personalized video enhancement parameters are determined based on the content of the video frames, and corresponding video enhancement processing is performed.
It achieves enhanced video frame content adaptation, improving the naturalness and accuracy of video quality.
Smart Images

Figure CN120431001B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video processing, and particularly relates to a video enhancement method and device. BACKGROUND
[0002] With the development of the times, people's requirements for video visual effects are constantly improving. In order to improve the video visual effect, video enhancement technology is usually used to enhance the video to improve the video quality.
[0003] The current video enhancement method is mainly a fixed rule-based video enhancement method, such as a histogram equalization-based brightness adjustment method and a high-pass filter-based sharpness enhancement method. The histogram equalization-based brightness adjustment method improves the brightness by equalizing the gray scale distribution of the image, and the high-pass filter-based sharpness enhancement method improves the definition by enhancing the edges.
[0004] Although the fixed rule-based video enhancement method is simple to implement, its enhancement effect is not good. SUMMARY
[0005] Therefore, the present application provides a video enhancement method and device to solve the problem of poor enhancement effect of the current fixed rule-based video enhancement method. The technical solutions are as follows:
[0006] The first aspect of the present application provides a video enhancement method, comprising:
[0007] obtaining a target video;
[0008] processing the target video into a plurality of target video frames;
[0009] obtaining high-dimensional semantic features of each target video frame by performing semantic analysis on each target video frame;
[0010] determining video enhancement parameters of each target video frame according to the high-dimensional semantic features of each target video frame;
[0011] enhancing each target video frame according to the video enhancement parameters of each target video frame to obtain an enhanced video frame of each target video frame;
[0012] composing the enhanced video frames of the target video frames into a video in order to obtain an enhanced video of the target video.
[0013] In a possible implementation manner, the high-dimensional semantic features of any target video frame can represent the scene information of the target video frame.
[0014] The scene information of the target video frame includes some or all of the following semantic information: global scene-level semantic information, local object-level semantic information, and action behavior-level semantic information.
[0015] In a possible implementation, the processing of the target video into a plurality of target video frames comprises:
[0016] frame the target video to obtain a plurality of original video frames;
[0017] preprocess the plurality of original video frames respectively to obtain a plurality of target video frames; wherein the preprocessing of any original video frame comprises one or more of the following processes: removing high-frequency noise in the original video frame, adjusting the resolution of the original video frame to a target resolution, and converting the color space of the original video frame to a target color space.
[0018] In a possible implementation, the obtaining of the high-dimensional semantic feature of each target video frame by performing semantic analysis on each target video frame comprises:
[0019] for each target video frame:
[0020] extracting, by using a feature extraction module of a pre-constructed semantic analysis model, multi-scale local features of the target video frame;
[0021] fusing, by using a feature fusion module of the semantic analysis model, features of different scales in the multi-scale local features of the target video frame to obtain multi-scale fusion features of the target video frame;
[0022] performing, by using an attention module of the semantic analysis model, attention weight weighting calculation on the multi-scale fusion features of the target video frame to obtain attention-weighted features of the target video frame;
[0023] identifying, by using an information recognition module of the semantic analysis model, global scene-level semantic information, local object-level semantic information, and action behavior-level semantic information of the target video frame according to the attention-weighted features of the target video frame;
[0024] constructing the high-dimensional semantic feature of the target video frame according to the global scene-level semantic information, the local object-level semantic information, and the action behavior-level semantic information of the target video frame.
[0025] In a possible implementation, the video enhancement method further comprises:
[0026] obtaining statistical features of each target video frame;
[0027] the determining of the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame comprises:
[0028] The video enhancement parameter of each target video frame is determined according to the high-dimensional semantic feature of the target video frame and in combination with the statistical feature of the target video frame.
[0029] In a possible implementation, the video enhancement parameter of any target video frame includes a brightness adjustment parameter, a contrast adjustment parameter and a sharpness adjustment parameter of the target video frame; the statistical feature of any target video frame includes average brightness, contrast and noise level of the target video frame; and the determination of the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of the target video frame and in combination with the statistical feature of the target video frame includes:
[0030] For each target video frame:
[0031] The brightness adjustment parameter of the target video frame is determined according to the mean value of the high-dimensional semantic feature of the target video frame and in combination with the average brightness of the target video frame;
[0032] The contrast adjustment parameter of the target video frame is determined according to the variance of the high-dimensional semantic feature of the target video frame and in combination with the contrast of the target video frame;
[0033] The sharpness adjustment parameter of the target video frame is determined according to the gradient of the high-dimensional semantic feature of the target video frame and in combination with the noise level of the target video frame.
[0034] In a possible implementation, the determination of the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of the target video frame and in combination with the statistical feature of the target video frame includes:
[0035] The video enhancement parameter of each target video frame is determined according to the high-dimensional semantic feature of the target video frame, in combination with the statistical feature of the target video frame and the type of a target video playback device, wherein the target video playback device is a device for playing an enhanced video of the target video.
[0036] In a possible implementation, the video enhancement parameter of any target video frame includes a brightness adjustment parameter, a contrast adjustment parameter and a sharpness adjustment parameter of the target video frame; the statistical feature of any target video frame includes average brightness, contrast and noise level of the target video frame; and the determination of the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of the target video frame, in combination with the statistical feature of the target video frame and the type of a target video playback device includes:
[0037] For each target video frame:
[0038] determine the brightness adjustment parameter of the target video frame according to the mean of the high-dimensional semantic feature of the target video frame, in combination with the average brightness of the target video frame and the device adaptation item corresponding to the type of the target video playing device in the brightness dimension;
[0039] determine the contrast adjustment parameter of the target video frame according to the variance of the high-dimensional semantic feature of the target video frame, in combination with the contrast of the target video frame and the device adaptation item corresponding to the type of the target video playing device in the contrast dimension;
[0040] determine the sharpness adjustment parameter of the target video frame according to the gradient of the high-dimensional semantic feature of the target video frame, in combination with the noise level of the target video frame and the device adaptation item corresponding to the type of the target video playing device in the sharpness dimension.
[0041] In a possible implementation, the video enhancement parameter of any target video frame includes the brightness adjustment parameter, the contrast adjustment parameter and the sharpness adjustment parameter of the target video frame;
[0042] The enhancing each target video frame according to the video enhancement parameter of each target video frame to obtain an enhanced video frame of each target video frame includes:
[0043] For each target video frame:
[0044] adjusting the brightness and the contrast of the target video frame pixel by pixel according to the brightness adjustment parameter and the contrast adjustment parameter of the target video frame to obtain an adjusted video frame of the target video frame;
[0045] adjusting the sharpness of the adjusted video frame of the target video frame pixel by pixel according to the sharpness adjustment parameter of the target video frame to obtain an enhanced video frame of the target video frame.
[0046] The second aspect of the present application provides a video enhancement device, comprising: a video acquisition module, a video processing module, a high-dimensional semantic feature acquisition module, a video enhancement parameter determination module, a video frame enhancement module and a video frame combination module;
[0047] The video acquisition module is configured to acquire a target video.
[0048] The video processing module is configured to process the target video into a plurality of target video frames.
[0049] The high-dimensional semantic feature acquisition module is configured to obtain a high-dimensional semantic feature of each target video frame by performing semantic analysis on each target video frame.
[0050] The video enhancement parameter determination module is configured to determine a video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame.
[0051] The video frame enhancement module is configured to enhance each target video frame according to the video enhancement parameter of each target video frame to obtain an enhanced video frame of each target video frame.
[0052] The video frame combination module is configured to combine the enhanced video frames of the target video frames in sequence to obtain an enhanced video of the target video.
[0053] According to the video enhancement method provided by the present application, after obtaining a target video, the target video is first processed into a plurality of target video frames, then the semantic analysis is performed on each target video frame to obtain a high-dimensional semantic feature of each target video frame, then the video enhancement parameter of each target video frame is determined according to the high-dimensional semantic feature of each target video frame, then each target video frame is enhanced according to the video enhancement parameter of each target video frame to obtain an enhanced video frame of each target video frame, and finally the enhanced video frames of the target video frames are combined in sequence to obtain an enhanced video of the target video. Since the video enhancement method provided by the present application determines the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame, the video enhancement parameter of each target video frame is a video enhancement parameter that is adapted to the content of the target video frame, and then the target video frame is enhanced according to the video enhancement parameter that is adapted to the content of the target video frame, a better video enhancement effect can be obtained. BRIEF DESCRIPTION OF DRAWINGS
[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.
[0055] Figure 1 It is a schematic diagram of a system architecture related to the present application;
[0056] Figure 2 It is a schematic diagram of a hardware structure of a terminal provided by the embodiment of the present application;
[0057] Figure 3 It is a schematic diagram of a hardware structure of a server provided by the embodiment of the present application;
[0058] Figure 4 It is a flowchart of a video enhancement method provided by the embodiment of the present application;
[0059] Figure 5 A flowchart of a process for obtaining high-dimensional semantic features of each target video frame by performing semantic analysis on each target video frame is provided for the embodiments of the present application.
[0060] Figure 6 A structural diagram of a video enhancement device is provided for the embodiments of the present application. DETAILED DESCRIPTION
[0061] The embodiments of the present application are described below in conjunction with the accompanying drawings. The terms used in the implementation part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0062] The embodiments of the present application are described below in conjunction with the accompanying drawings. The skilled person can know that, as technology develops and new scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0063] The terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is only a distinguishing way used in the description of the embodiments of the present application to describe the objects with the same attributes. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that the processes, methods, systems, products or equipment containing a series of units do not have to be limited to those units, but can include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0064] In one possible implementation, as shown in Figure 1 The system architecture involved in the present application can include a terminal 101 and a server 102, and the terminal 101 can interact with the server 102 through a network (wired network or wireless network). The server 102 can include one or more servers (as an example, one server is included in the server 102). Figure 1 The terminal 101 can obtain a video to be enhanced, send the video to be enhanced to the server 102, and the server 102 can perform video enhancement processing on the video to be enhanced and send the enhanced video to the terminal 101.
[0065] In another possible implementation, the system architecture involved in the present application can include a terminal. The terminal has strong data processing capability, and can obtain a video to be enhanced and perform video enhancement processing on the video to be enhanced.
[0066] Next, the product form of the above terminal is described.
[0067] The terminal described above can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, a robot, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and the like, and embodiments of the present application do not limit the terminal to any of these.
[0068] Figure 2 An optional hardware structure diagram of the terminal is shown.
[0069] Reference Figure 2 As shown, the terminal can include a radio frequency unit 210, a memory 220, an input unit 230, a display unit 240, a camera 250 (optional), an audio circuit 260 (optional), a speaker 261 (optional), a microphone 262 (optional), a headphone jack 263 (optional), a processor 270, an external interface 280, a power supply 290, and the like. Those skilled in the art can understand that the terminal can include more or fewer components than those shown, or combine some components, or different components. Figure 2 The above is merely an example of the terminal and does not limit the terminal, which can include more or fewer components than those shown, or combine some components, or different components.
[0070] The input unit 230 can be used to receive input digital or character information, and generate key signal input related to user settings and function control of the terminal. Specifically, the input unit 230 can include a touch screen 231 (optional) and / or other input devices 232. The touch screen 231 can collect user touch operations (such as user operations on or near the touch screen using a finger, a joint, a stylus, or any suitable object) and drive the corresponding connection device according to the pre-set program. The touch screen can detect the user's touch action on the touch screen, convert the touch action into a touch signal and send it to the processor 270, and can receive commands from the processor 270 and execute them; the touch signal at least includes touch point coordinate information. The touch screen 231 can provide an input interface and an output interface between the terminal and the user. In addition, the touch screen can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch screen 231, the input unit 230 can also include other input devices. Specifically, the other input devices 232 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control buttons, on-off buttons, etc.), trackballs, mice, joysticks, and the like.
[0071] The display unit 240 can be used to display information input by the user or information provided to the user, various menus of the terminal, interactive interfaces, file display and / or playback of any multimedia file.
[0072] The memory 220 can be used to store instructions and data. The memory 220 may primarily include an instruction storage area and a data storage area. The data storage area can store various types of data, such as multimedia files and text. The instruction storage area can store software units such as operating systems, applications, and instructions required for at least one function, or subsets or extended sets thereof. It may also include non-volatile random access memory. It provides the processor 270 with hardware, software, and data resources for managing the computing device, supporting control software and applications. It is also used for storing multimedia files, as well as storing running programs and applications.
[0073] The processor 270 is the control center of the terminal, connecting various parts of the terminal through various interfaces and lines. It executes instructions stored in the memory 220 and calls data stored in the memory 220 to perform various functions and process data, thereby controlling the terminal as a whole. Optionally, the processor 270 may include one or more processing units; preferably, the processor 270 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 270. In some embodiments, the processor and memory can be implemented on a single chip; in some embodiments, they can also be implemented on separate chips. The processor 270 can also be used to generate corresponding operation control signals, send them to the corresponding components of the computing device, read and process data in the software, especially read and process data and programs in the memory 220, so that the various functional modules therein perform corresponding functions, thereby controlling the corresponding components to act according to the instructions.
[0074] The memory 220 can be used to store software code related to the video enhancement method. The processor 270 can execute the software code in the memory 220, and can also schedule other units (such as the input unit 230 and the display unit 240 mentioned above) to achieve the corresponding functions.
[0075] The radio frequency unit 210 (optional) can be used for receiving and sending signals in the process of information or communication, for example, receiving the downlink information of the base station and processing by the processor 270; in addition, sending the uplink data to the base station. Generally, the radio frequency unit 210 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the radio frequency unit 210 can also communicate with network devices and other devices through wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0076] In the embodiments of the present application, the radio frequency unit 210 can send data to other devices, and can also receive data sent by other devices. It should be understood that the radio frequency unit 210 is optional, which can be replaced by other communication interfaces, for example, can be a network interface.
[0077] The terminal also includes a power supply 290 (such as a battery) for supplying power to each component. Preferably, the power supply can be logically connected to the processor 270 through a power management system, so as to realize the functions of managing charging, discharging, and power consumption management through the power management system.
[0078] The terminal also includes an external interface 280, which can be a standard Micro USB interface, or can be a multi-pin connector, and can be used for connecting the terminal with other devices for communication, or can be used for connecting a charger to charge the terminal.
[0079] Although not shown, the terminal can also include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, different function sensors, etc., which will not be described here.
[0080] Next, the product form of the above server is described.
[0081] Figure 3 A structural schematic diagram of the above server is provided, as shown in Figure 3As shown, the server can include a bus 301, a processing device 302, a communication interface 303 and a storage device 304. The processing device 302, the storage device 304 and the communication interface 303 communicate through the bus 301.
[0082] The bus 301 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the middle, but it does not mean that there is only one bus or one type of bus.
[0083] The processing device 302 can be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.
[0084] The storage device 304 can include a volatile memory, such as a random access memory (RAM). The storage device 304 can also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a mechanical hard disk drive (HDD) or a solid state drive (SSD).
[0085] The storage device 304 can be used to store software code related to the video enhancement method, and the processing device 302 can call the software code stored in the storage device 304, or schedule other units to realize the corresponding functions.
[0086] The processor 270 in the terminal and the processing device 302 in the server can be a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor, a microcontroller, or the like), or a combination of these hardware circuits, for example, can be a hardware system with an instruction execution function, such as a CPU, a DSP, or the like, or a hardware system without an instruction execution function, such as an ASIC, an FPGA, or the like, or a combination of the hardware system without the instruction execution function and the hardware system with the instruction execution function.
[0087] The existing video enhancement method based on fixed rules is consistent in processing each video frame of the video, and each video frame of the video usually involves multiple scenes, and the content of different video frames can be quite different. The video enhancement method based on fixed rules lacks flexibility and has poor enhancement effect.
[0088] In view of the problems of the existing video enhancement method based on fixed rules, the present application provides a video enhancement method with good effect. Next, the video enhancement method provided by the present application will be introduced through the following embodiments.
[0089] Please refer to Figure 4 , which shows a flowchart of the video enhancement method provided by the embodiment of the present application. The video enhancement method can include the following steps:
[0090] Step S401: Obtain a target video.
[0091] The target video is a video to be enhanced.
[0092] Step S402: Process the target video into a plurality of target video frames.
[0093] Specifically, the target video is first framed to obtain a plurality of original video frames, and then the plurality of original video frames are preprocessed respectively to obtain a plurality of target video frames (the target video frame is a video frame after preprocessing of the original video frame).
[0094] The preprocessing of any original video frame can include one or more of the following processes, but is not limited to: denoising, resolution standardization, color space conversion.
[0095] The denoising processing is used to remove high-frequency noise in the original video frame to reduce the influence on subsequent processing; the resolution normalization processing is used to adjust the resolution of the original video frame to a target resolution (such as 1080p, 720p, etc.) to facilitate subsequent processing; and the color space conversion processing is used to convert the color space of the original video frame to a target color space, such as converting the original video frame from the RGB color space to the grayscale color space or the YUV color space.
[0096] In addition, when the target video frame is divided, the timestamp information of each original video frame can be extracted to preserve the time sequence. It can be understood that the timestamp information of the target video frame obtained by preprocessing an original video frame is the timestamp information of the original video frame.
[0097] Step S403: Obtain the high-dimensional semantic feature of each target video frame by performing semantic analysis on each target video frame.
[0098] The high-dimensional semantic feature of any target video frame is an abstract representation of the target video frame mapped into a high-dimensional space, and the high-dimensional semantic feature of any target video frame represents the scene information of the target video frame. The scene information of the target video frame can include part or all (preferably all) of the following semantic information: global scene-level semantic information (such as global scene category, etc.), local object-level semantic information (such as object category, object attribute, etc.), and action behavior-level semantic information (such as action type, action speed, action direction, etc.).
[0099] Step S404: Determine the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame.
[0100] For each target video frame, the video enhancement parameter of the target video frame is determined according to the high-dimensional semantic feature of the target video frame.
[0101] In this embodiment, the video enhancement parameter of any target video frame is determined according to the high-dimensional semantic feature of the target video frame, and thus the video enhancement parameter of the target video frame is a video enhancement parameter that is adapted to the content of the target video frame. Since the high-dimensional semantic features of different target video frames are usually different, the video enhancement parameters of different target video frames are also usually different.
[0102] In one possible implementation, the video enhancement parameter of each target video frame can be determined only according to the high-dimensional semantic feature of the target video frame.
[0103] To further improve the video enhancement effect, so that the enhancement effect is more accurate, natural, and more scene-adaptive, in another possible implementation, the statistical features of each target video frame can be obtained, and then the video enhancement parameters of each target video frame are determined according to the high-dimensional semantic features of each target video frame, while combining the statistical features of each target video frame.
[0104] To be able to adapt to the target video playback device (a device that plays the enhanced video of the target video), in another possible implementation, the video enhancement parameters of each target video frame can be determined according to the high-dimensional semantic features of each target video frame, while combining the statistical features of each target video frame and the type of the target video playback device.
[0105] Of course, the video enhancement parameters of each target video frame can also be determined according to the high-dimensional semantic features of each target video frame, while combining the type of the target video playback device.
[0106] Step S405: According to the video enhancement parameters of each target video frame, each target video frame is enhanced to obtain an enhanced video frame of each target video frame.
[0107] For each target video frame, according to the video enhancement parameters of the target video frame, the target video frame is enhanced to obtain an enhanced video frame of the target video frame.
[0108] Step S406: The enhanced video frames of the target video frames are synthesized into a video in order to obtain an enhanced video of the target video.
[0109] Specifically, the enhanced video frames of the target video frames can be synthesized into a video according to the timestamp information of the target video frames to obtain an enhanced video of the target video. Video encoding (such as H.264 or H.265 format) can be performed during the synthesis of the video to be compatible with multiple devices.
[0110] After the video is synthesized, the synthesized video, i.e., the enhanced video of the target video, can be saved in a specified format, such as MP4 format, AVI format, etc.
[0111] The video enhancement method provided in the embodiments of the present application comprises the following steps: after a target video is obtained, the target video is processed into a plurality of target video frames; high-dimensional semantic features of each target video frame are obtained by performing semantic analysis on each target video frame; video enhancement parameters of each target video frame are determined according to the high-dimensional semantic features of each target video frame; each target video frame is enhanced according to the video enhancement parameters of each target video frame, so as to obtain an enhanced video frame of each target video frame; and finally, an enhanced video of the target video is obtained by synthesizing the enhanced video frames of the target video frames in sequence. Since the video enhancement method provided in the embodiments of the present application determines the video enhancement parameters of each target video frame according to the high-dimensional semantic features of each target video frame, the video enhancement parameters of each target video frame are video enhancement parameters that are adapted to the content of the target video frame, and the target video frame is enhanced according to the video enhancement parameters that are adapted to the content of the target video frame, so that a better enhancement effect can be obtained.
[0112] In another embodiment of the present application, the specific implementation process of "step S403: obtaining high-dimensional semantic features of each target video frame by performing semantic analysis on each target video frame" in the above embodiment is introduced.
[0113] Referring to Figure 5 , a flowchart for obtaining high-dimensional semantic features of each target video frame by performing semantic analysis on each target video frame is shown, which can comprise the following steps:
[0114] Step S501: for each target video frame, a feature extraction module of a pre-constructed semantic analysis model is used to extract multi-scale local features of the target video frame.
[0115] The semantic analysis model can be a deep learning model, and the feature extraction module of the semantic analysis model can be, but is not limited to, a convolutional network.
[0116] The target video frame is input into the feature extraction module of the semantic analysis model, and the feature extraction module extracts multi-scale local features (such as texture features, edge features, etc.) of the input target video frame and outputs them.
[0117] Step S502: a feature fusion module of the semantic analysis model is used to fuse features of different scales in the multi-scale local features of the target video frame, so as to obtain multi-scale fusion features of the target video frame.
[0118] The multi-scale local features output by the feature extraction module are input into a feature fusion module of the semantic analysis model. The feature fusion module fuses features of different scales in the input multi-scale fusion features, and outputs multi-scale fusion features. Fusing features of different scales means fusing features of different semantic levels. The multi-scale fusion features obtained by fusing features of different scales contain both local information (such as texture information, edge information, etc.) of the target video frame and global semantic information (such as scene information, object information, etc.) of the target video frame.
[0119] The feature fusion module of the semantic analysis model can be, but is not limited to, a feature pyramid network (FPN).
[0120] In step S503, the attention module of the semantic analysis model is used to perform attention weight weighting processing on the multi-scale fusion features of the target video frame, to obtain attention-weighted features of the target video frame.
[0121] The attention weight weighting processing on the multi-scale fusion features of the target video frame can highlight key features in the multi-scale fusion features of the target video frame.
[0122] In step S504, the information recognition module of the semantic analysis model is used to recognize global scene-level semantic information, local object-level semantic information, and action behavior-level semantic information of the target video frame according to the attention-weighted features of the target video frame.
[0123] The global scene-level semantic information can be, but is not limited to, a global scene category (such as indoor, outdoor, daytime, nighttime, etc.), the local object-level semantic information can be, but is not limited to, an object category (such as a person, a table, a chair, a computer, etc.), an object attribute (such as the height, size, position, shape, color, etc. of an object), and the action behavior-level semantic information can be, but is not limited to, an action type, an action speed, an action direction, etc.
[0124] The attention-weighted features of the target video frame are input into the information recognition module of the semantic analysis model for information recognition. The information recognition module outputs global scene-level semantic information, local object-level semantic information, and action behavior-level semantic information of the target video frame. The information recognition module of the semantic analysis model can be, but is not limited to, a fully connected layer.
[0125] It should be noted that the embodiment is not limited to recognizing global scene-level semantic information, local object-level semantic information, and action behavior-level semantic information of the target video frame, but can also recognize part of the global scene-level semantic information, local object-level semantic information, and action behavior-level semantic information of the target video frame.
[0126] Step S505: constructing high-dimensional semantic features of the target video frame according to the global scene-level semantic information, the local object-level semantic information and the action behavior-level semantic information of the target video frame.
[0127] The high-dimensional semantic features of the target video frame can represent the global scene-level semantic information, the local object-level semantic information and the action behavior-level semantic information of the target video frame.
[0128] The embodiments of the present application promote the target video frame from pixel level to semantic level through the semantic analysis model, and provide a strong foundation for determination of the video enhancement parameter and video enhancement.
[0129] The above embodiments mention that the video enhancement parameter of each target video frame can be determined according to the high-dimensional semantic features of each target video frame, and in combination with the statistical features of each target video frame, and in another embodiment of the present application, the process is introduced.
[0130] In a possible implementation, the video enhancement parameter of any target video frame can include a brightness adjustment parameter, a contrast adjustment parameter and a sharpness adjustment parameter of the target video frame, and the statistical features of any target video frame can include an average brightness, a contrast and a noise level of the target video frame.
[0131] The process of determining the video enhancement parameter of each target video frame according to the high-dimensional semantic features of each target video frame, and in combination with the statistical features of each target video frame can include: for each target video frame, determining the brightness adjustment parameter, the contrast adjustment parameter and the sharpness adjustment parameter of the target video frame through the following steps a1, step a2 and step a3.
[0132] Step a1: determining the brightness adjustment parameter of the target video frame according to the mean of the high-dimensional semantic features of the target video frame, and in combination with the average brightness of the target video frame.
[0133] Specifically, the mean of the high-dimensional semantic features of the target video frame can be calculated, and the brightness adjustment parameter of the target video frame can be determined according to the average brightness of the target video frame and the mean of the high-dimensional semantic features of the target video frame.
[0134] More specifically, the brightness adjustment parameter of the target video frame can be determined in the following manner:
[0135] B = B0 + α1 · mean (F sem ) + α2 · L mean (1).
[0136] Wherein, F sem represents the high-dimensional semantic features of the target video frame, mean (F sem) represents the mean value of the high-dimensional semantic feature of the target video frame, L mean represents the average brightness of the target video frame, B0 is a preset initial brightness, which is the expected baseline brightness, a1 represents a semantic feature adjustment coefficient, a2 represents a statistical feature adjustment coefficient, and B is the brightness adjustment parameter of the target video frame.
[0137] Step a2, according to the variance of the high-dimensional semantic feature of the target video frame, and in combination with the contrast of the target video frame, the contrast adjustment parameter of the target video frame is determined.
[0138] Specifically, the variance of the high-dimensional semantic feature of the target video frame is calculated, and according to the contrast of the target video frame and the variance of the high-dimensional semantic feature of the target video frame, the contrast adjustment parameter of the target video frame is determined.
[0139] More specifically, the contrast adjustment parameter of the target video frame can be determined in the following manner:
[0140] C = C0+ β1·var(F sem )+ β2·C idx (2).
[0141] Wherein, var(F sem ) represents the variance of the high-dimensional semantic feature F sem of the target video frame, which is a quantitative description of the diversity and distribution of the semantic information inside the target video frame, reflecting the scene complexity, C idx represents the contrast of the target video frame, C0 is a preset initial contrast, β1 represents a semantic feature adjustment coefficient, β2 represents a statistical feature adjustment coefficient, and C is the contrast adjustment parameter of the target video frame.
[0142] Step a3, according to the gradient of the high-dimensional semantic feature of the target video frame, and in combination with the noise level of the target video frame, the sharpness adjustment parameter of the target video frame is determined.
[0143] Wherein, the noise level of the target video frame refers to the intensity of the noise in the target video frame. Considering that some video frames may have serious noise (such as night video), by evaluating the noise level, the amplitude of the sharpness enhancement can be more accurately controlled.
[0144] Specifically, the gradient of the high-dimensional semantic feature of the target video frame is calculated, and according to the noise level of the target video frame and the gradient of the high-dimensional semantic feature of the target video frame, the sharpness adjustment parameter of the target video frame is determined.
[0145] More specifically, the sharpness adjustment parameter of the target video frame can be determined in the following manner:
[0146] S = S0+ γ1·grad(F sem )+ γ2·N level (3)。
[0147] wherein grad(F sem ) represents the gradient of the high-dimensional semantic feature F sem of the target video frame, which reflects the mutation degree of semantic information, N level represents the noise level of the target video frame, S0represents the preset initial sharpness, γ1represents the semantic feature adjustment coefficient, γ2represents the statistical feature adjustment coefficient, and S is the sharpness adjustment parameter of the target video frame.
[0148] For each target video frame, the brightness adjustment parameter, the contrast adjustment parameter, and the sharpness adjustment parameter that are adapted to the scene of the target video frame can be obtained via the above steps a1-a3.
[0149] The embodiment is not limited to the execution order of steps a1-a3, and steps a1-a3 can be executed in series or in parallel.
[0150] In a possible implementation, steps a1-a3 can be implemented based on an enhancement parameter generation model, that is, the enhancement parameter generation model determines the brightness adjustment parameter, the contrast adjustment parameter, and the sharpness adjustment parameter of each target video frame according to the process of steps a1-a3.
[0151] The above embodiment mentions that, in order to be adapted to a target video playback device (a device that plays the enhanced video of the target video), in another possible implementation, the video enhancement parameter of each target video frame can be determined according to the high-dimensional semantic feature of each target video frame, in combination with the statistical feature of each target video frame and the type of the target video playback device. In another embodiment of the present application, this process is introduced.
[0152] In a possible implementation, the video enhancement parameter of any target video frame can include the brightness adjustment parameter, the contrast adjustment parameter, and the sharpness adjustment parameter of the target video frame, and the statistical feature of any target video frame can include the average brightness, the contrast, and the noise level of the target video frame.
[0153] The process of determining the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame, in combination with the statistical feature of each target video frame and the type of the target video playback device can include: for each target video frame, determining the brightness adjustment parameter, the contrast adjustment parameter, and the sharpness adjustment parameter of the target video frame through steps b1-b3 as follows.
[0154] Step b1: Based on the mean of the high-dimensional semantic features of the target video frame, and combined with the average brightness of the target video frame and the device adaptation item corresponding to the brightness dimension of the target video playback device type, determine the brightness adjustment parameters of the target video frame.
[0155] Specifically, the mean of the high-dimensional semantic features of the target video frame can be calculated, and the brightness adjustment parameters of the target video frame can be determined based on the average brightness of the target video frame, the mean of the high-dimensional semantic features of the target video frame, and the device adaptation item corresponding to the brightness dimension of the target video playback device type.
[0156] More specifically, the brightness adjustment parameters of the target video frame can be determined as shown in the following formula:
[0157] B = B0 + α1·mean(F) sem )+α2·L mean +δ B device (4).
[0158] Among them, F sem This represents the high-dimensional semantic features of the target video frame, mean(F sem L represents the mean of the high-dimensional semantic features of the target video frame. mean δ represents the average brightness of the target video frame. B device The device adaptation item corresponding to the type of the target video playback device in the brightness dimension (the compensation parameter introduced in the brightness dimension according to the type of the target video playback device), B0 is the preset initial brightness, which is the expected baseline brightness, α1 represents the semantic feature adjustment coefficient, α2 represents the statistical feature adjustment coefficient, and B is the brightness adjustment parameter of the target video frame.
[0159] Step b2: Based on the variance of the high-dimensional semantic features of the target video frame, and in conjunction with the contrast of the target video frame and the device adaptation item corresponding to the type of the target video playback device in the contrast dimension, determine the contrast adjustment parameters of the target video frame.
[0160] Specifically, the variance of the high-dimensional semantic features of the target video frame can be calculated, and the contrast adjustment parameters of the target video frame can be determined based on the contrast of the target video frame, the variance of the high-dimensional semantic features of the target video frame, and the device adaptation item corresponding to the type of the target video playback device in the contrast dimension.
[0161] More specifically, the contrast adjustment parameters for the target video frame can be determined as shown in the following formula:
[0162] C = C0 + β1·var(F) sem )+β2·C idx +δC device (5)。
[0163] wherein, var(F sem ) represents the variance of the high-dimensional semantic feature F sem of the target video frame, C idx represents the contrast of the target video frame, δ C device represents the device adaptation item corresponding to the type of the target video playing device in the contrast dimension (a compensation parameter introduced according to the type of the target video playing device in the contrast dimension), C0 represents a preset initial contrast, β1 represents a semantic feature adjustment coefficient, β2 represents a statistical feature adjustment coefficient, and C represents the contrast adjustment parameter of the target video frame.
[0164] Step b3, determining the sharpness adjustment parameter of the target video frame according to the gradient of the high-dimensional semantic feature of the target video frame, in combination with the noise level of the target video frame and the device adaptation item corresponding to the type of the target video playing device in the sharpness dimension.
[0165] Specifically, the gradient of the high-dimensional semantic feature of the target video frame is calculated, and the sharpness adjustment parameter of the target video frame is determined according to the noise level of the target video frame, the gradient of the high-dimensional semantic feature of the target video frame and the device adaptation item corresponding to the type of the target video playing device in the sharpness dimension.
[0166] More specifically, the sharpness adjustment parameter of the target video frame can be determined in the following manner:
[0167] S = S0 + γ1·grad(F sem ) + γ2·N level + δ S device (6).
[0168] wherein, grad(F sem ) represents the gradient of the high-dimensional semantic feature F sem of the target video frame, which reflects the mutation degree of semantic information, N level represents the noise level of the target video frame, δ C device represents the device adaptation item corresponding to the type of the target video playing device in the sharpness dimension (a compensation parameter introduced according to the type of the target video playing device in the sharpness dimension), S0 represents a preset initial sharpness, γ1 represents a semantic feature adjustment coefficient, γ2 represents a statistical feature adjustment coefficient, and S represents the sharpness adjustment parameter of the target video frame.
[0169] The following shows examples of device adaptation items corresponding to the luminance dimension, the contrast dimension and the sharpness dimension of several types of target video playing devices:
[0170] Table 1: Examples of types of target video playing devices and corresponding device adaptation items
[0171]
[0172] The initial values of the device adaptation items corresponding to the types of target video playing devices can be preset, and can be dynamically adjusted according to requirements in subsequent applications.
[0173] In addition, it should be noted that the specific values of the device adaptation items given in Table 1 are only examples, and the specific values of the device adaptation items corresponding to the types of target video playing devices can be set according to actual application scenarios.
[0174] For each target video frame, the brightness adjustment parameter, the contrast adjustment parameter, and the sharpness adjustment parameter that are adapted to the scene of the target video frame and adapted to the type of target video playing device can be obtained through the above steps b1-b3.
[0175] The embodiment does not limit the execution order of steps b1-b3, and steps b1-b3 can be executed in series or in parallel.
[0176] In one possible implementation, steps b1-b3 can be implemented based on an enhancement parameter generation model, that is, the enhancement parameter generation model determines the brightness adjustment parameter, the contrast adjustment parameter, and the sharpness adjustment parameter of each target video frame according to the process of steps b1-b3.
[0177] In another embodiment of the present application, the specific implementation process of "step S405: enhancing each target video frame according to the video enhancement parameter of each target video frame to obtain an enhanced video frame of each target video frame" in the above embodiment is introduced.
[0178] The implementation process of enhancing each target video frame according to the video enhancement parameter of each target video frame to obtain an enhanced video frame of each target video frame can include:
[0179] Step c1: for each target video frame, performing brightness adjustment and contrast adjustment on the target video frame pixel by pixel according to the brightness adjustment parameter and the contrast adjustment parameter of the target video frame to obtain an adjusted video frame of the target video frame.
[0180] Specifically, for each target video frame, the brightness adjustment parameter and the contrast adjustment parameter of the target video frame can be used to adjust each pixel value of the target video frame in the following manner:
[0181] (7)。
[0182] wherein P pixel orig represents any pixel value of the target video frame, C represents the contrast adjustment parameter of the target video frame, B represents the brightness adjustment parameter of the target video frame, P pixel new represents P pixel orig the adjusted pixel value.
[0183] Step c2, performing sharpness adjustment on the adjusted video frame of the target video frame pixel by pixel according to the sharpness adjustment parameter of the target video frame to obtain an enhanced video frame of the target video frame.
[0184] Specifically, each pixel value of the adjusted video frame of the target video frame can be further adjusted according to the sharpness adjustment parameter of the target video frame in the following manner:
[0185] (8).
[0186] wherein ΔP represents the pixel gradient of the video frame obtained in step c1, S represents the sharpness adjustment parameter of the target video frame, P sharp represents P pixel new the pixel value after sharpness adjustment.
[0187] Through the above process, the enhanced video frame of each target video frame can be obtained.
[0188] After obtaining the enhanced video frame of each target video frame, the enhanced video frames of each target video frame can be synthesized into a video according to the timestamp information of each target video frame to obtain an enhanced video of the target video.
[0189] Optionally, after obtaining the enhanced video of the target video, the enhanced video of the target video can be displayed for the user to view the video enhancement effect.
[0190] Optionally, after displaying the enhanced video of the target video, feedback data of the user on the enhanced video of the target video can be obtained, and then the parameter generation model can be optimized according to the feedback data of the user on the enhanced video of the target video.
[0191] The video enhancement method provided in the embodiments of the present application can obtain high-dimensional semantic features of each target video frame of a target video through a semantic analysis model, can further determine video enhancement parameters that are adapted to the content (i.e., scene information) of the target video frame according to the high-dimensional semantic features, and can further enhance the target video frame according to the video enhancement parameters that are adapted to the scene information of the target video frame. In addition, the video enhancement method provided in the embodiments of the present application can further introduce a device adaptation item according to the type of a target video playing device when determining the video enhancement parameters. The semantic features of the content of a video can be deeply understood through the semantic analysis model, and the video can be accurately optimized. The video enhancement parameters of each target video frame are determined according to the high-dimensional semantic features of each target video frame, and each target video frame is enhanced according to the video enhancement parameters, so that multi-scene dynamic adaptation can be achieved. The introduction of the device adaptation item can achieve multi-device dynamic adaptation. The video enhancement method provided in the embodiments of the present application can significantly improve the comprehensive performance of video quality and meet the diversified needs of users.
[0192] The embodiments of the present application further provide a device for executing the video enhancement method provided in the above embodiments. Please refer to Figure 6 , Figure 6 FIG. 1 is a structural schematic diagram of a video enhancement device provided in the embodiments of the present application. The video enhancement device can include a video acquisition module 601, a video processing module 602, a high-dimensional semantic feature acquisition module 603, a video enhancement parameter determination module 604, a video frame enhancement module 605, and a video frame combination module 606.
[0193] The video acquisition module 601 is configured to acquire a target video.
[0194] The video processing module 602 is configured to process the target video into a plurality of target video frames.
[0195] The high-dimensional semantic feature acquisition module 603 is configured to obtain high-dimensional semantic features of each target video frame by performing semantic analysis on each target video frame.
[0196] The video enhancement parameter determination module 604 is configured to determine video enhancement parameters of each target video frame according to the high-dimensional semantic features of each target video frame.
[0197] The video frame enhancement module 605 is configured to enhance each target video frame according to the video enhancement parameters of each target video frame to obtain an enhanced video frame of each target video frame.
[0198] The video frame combination module 606 is configured to combine the enhanced video frames of the target video frames in sequence to obtain an enhanced video of the target video.
[0199] In a possible implementation, the high-dimensional semantic feature of any target video frame can represent scene information of the target video frame; the scene information of the target video frame includes part or all of the following semantic information: global scene-level semantic information, local object-level semantic information, and action behavior-level semantic information.
[0200] In a possible implementation, the video processing module 602 can include a video framing module and a video frame preprocessing module.
[0201] The video framing module is configured to frame a target video to obtain a plurality of original video frames.
[0202] The video frame preprocessing module is configured to preprocess the plurality of original video frames respectively to obtain a plurality of target video frames.
[0203] The preprocessing performed by the video frame preprocessing module on any original video frame includes one or more of the following processes: removing high-frequency noise in the original video frame, adjusting a resolution of the original video frame to a target resolution, and converting a color space of the original video frame to a target color space.
[0204] In a possible implementation, when obtaining the high-dimensional semantic feature of each target video frame by performing semantic analysis on each target video frame, the high-dimensional semantic feature acquisition module 603 is specifically configured to:
[0205] For each target video frame:
[0206] extracting, by using a feature extraction module of a pre-constructed semantic analysis model, multi-scale local features of the target video frame;
[0207] fusing, by using a feature fusion module of the semantic analysis model, features of different scales in the multi-scale local features of the target video frame to obtain multi-scale fusion features of the target video frame;
[0208] performing attention weight weighting calculation on the multi-scale fusion features of the target video frame by using an attention module of the semantic analysis model to obtain attention-weighted features of the target video frame;
[0209] identifying, by using an information recognition module of the semantic analysis model, global scene-level semantic information, local object-level semantic information, and action behavior-level semantic information of the target video frame according to the attention-weighted features of the target video frame;
[0210] constructing a high-dimensional semantic feature of the target video frame according to the global scene-level semantic information, the local object-level semantic information, and the action behavior-level semantic information of the target video frame.
[0211] In a possible implementation, the video enhancement apparatus further includes a video frame statistical feature acquisition module. The video frame statistical feature acquisition module is configured to acquire the statistical feature of each target video frame.
[0212] The video enhancement parameter determination module 604, when determining the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame, is specifically configured to:
[0213] The video enhancement parameter determination module 604, when determining the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame, is specifically configured to:
[0214] In a possible implementation, the video enhancement parameter of any target video frame includes a brightness adjustment parameter, a contrast adjustment parameter, and a sharpness adjustment parameter of the target video frame; and the statistical feature of any target video frame includes an average brightness, a contrast, and a noise level of the target video frame. The video enhancement parameter determination module 604, when determining the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame and in combination with the statistical feature of each target video frame, is specifically configured to:
[0215] For each target video frame:
[0216] The video enhancement parameter determination module 604, when determining the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame, is specifically configured to:
[0217] The video enhancement parameter determination module 604, when determining the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame, is specifically configured to:
[0218] The video enhancement parameter determination module 604, when determining the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame, is specifically configured to:
[0219] In a possible implementation, the video enhancement parameter determination module 604, when determining the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame and in combination with the statistical feature of each target video frame, is specifically configured to:
[0220] The video enhancement parameter determination module 604, when determining the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame, is specifically configured to:
[0221] In a possible implementation, the video enhancement parameter of any target video frame includes a brightness adjustment parameter, a contrast adjustment parameter and a sharpness adjustment parameter of the target video frame; and the statistical feature of any target video frame includes an average brightness, a contrast and a noise level of the target video frame. The video enhancement parameter determination module 604, when determining the video enhancement parameter of each target video frame according to the high-dimensional semantic feature of each target video frame, in combination with the statistical feature of each target video frame and the type of the target video playback device, is specifically configured to:
[0222] For each target video frame:
[0223] determine the brightness adjustment parameter of the target video frame according to the mean of the high-dimensional semantic feature of the target video frame, in combination with the average brightness of the target video frame and the device adaptation item corresponding to the brightness dimension of the type of the target video playback device;
[0224] determine the contrast adjustment parameter of the target video frame according to the variance of the high-dimensional semantic feature of the target video frame, in combination with the contrast of the target video frame and the device adaptation item corresponding to the contrast dimension of the type of the target video playback device;
[0225] determine the sharpness adjustment parameter of the target video frame according to the gradient of the high-dimensional semantic feature of the target video frame, in combination with the noise level of the target video frame and the device adaptation item corresponding to the sharpness dimension of the type of the target video playback device.
[0226] In a possible implementation, the video enhancement parameter of any target video frame includes a brightness adjustment parameter, a contrast adjustment parameter and a sharpness adjustment parameter of the target video frame. The video frame enhancement module 605, when enhancing each target video frame according to the video enhancement parameter of each target video frame to obtain an enhanced video frame of each target video frame, is specifically configured to:
[0227] For each target video frame:
[0228] perform brightness adjustment and contrast adjustment on the target video frame pixel by pixel according to the brightness adjustment parameter and the contrast adjustment parameter of the target video frame, to obtain an adjusted video frame of the target video frame;
[0229] perform sharpness adjustment on the adjusted video frame of the target video frame pixel by pixel according to the sharpness adjustment parameter of the target video frame, to obtain an enhanced video frame of the target video frame.
[0230] The video enhancement device provided in the embodiments of the present application, after obtaining a target video, first processes the target video into a plurality of target video frames, then performs semantic analysis on each target video frame to obtain high-dimensional semantic features of each target video frame, then determines video enhancement parameters of each target video frame according to the high-dimensional semantic features of each target video frame, then enhances each target video frame according to the video enhancement parameters of each target video frame to obtain enhanced video frames of each target video frame, and finally synthesizes the enhanced video frames of the target video frames into a video in sequence to obtain an enhanced video of the target video. Since the video enhancement device provided in the embodiments of the present application determines the video enhancement parameters of each target video frame according to the high-dimensional semantic features of each target video frame, the video enhancement parameters of each target video frame are video enhancement parameters that are adapted to the content of the target video frame, and then the target video frame is enhanced according to the video enhancement parameters that are adapted to the content of the target video frame, a better enhancement effect can be obtained.
[0231] The embodiments of the present application further provide an electronic device, which can include at least one processor and a memory connected with the processor.
[0232] The processor can be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application, etc.; the memory can include a high-speed RAM memory, and can further include a non-volatile memory, etc., such as at least one disk memory.
[0233] The memory is configured to store a computer program, and the processor is configured to execute the computer program, so that the electronic device can implement the steps of the video enhancement method provided in the above embodiments.
[0234] The embodiments of the present application further provide a computer storage medium, which carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the steps of the video enhancement method provided in the above embodiments.
[0235] The embodiments of the present application further provide a computer program product, which includes computer readable instructions, and when the computer readable instructions run on an electronic device, the electronic device implements the steps of the video enhancement method provided in the above embodiments.
[0236] In addition, it should be noted that the apparatus embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. In addition, the connection relationship between the modules in the apparatus embodiments provided in the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0237] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be realized by means of software and the necessary general hardware, and of course can also be realized by special hardware including special integrated circuits, special CPUs, special memories, special components and the like. Generally, functions completed by computer programs can be easily realized by corresponding hardware, and the specific hardware structure for realizing the same function can also be various, such as analog circuit, digital circuit or special circuit. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of software products, which are stored in readable storage media, such as computer floppy disks, U disks, mobile hard disks, ROM, RAM, magnetic or optical disks, etc., including a plurality of instructions for making a computer device (which can be a personal computer, a training device, or a network device, etc.) execute the methods described in various embodiments of the present application.
[0238] In the above embodiments, all or part can be realized by software, hardware, firmware or any combination thereof. When realized by software, it can be realized in the form of a computer program product in whole or in part.
[0239] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as a training device, a data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
Claims
1. A method of video enhancement, characterized by, The method comprises the following steps: obtaining a target video; processing the target video into a plurality of target video frames; obtaining high-dimensional semantic features of each target video frame by performing semantic analysis on each target video frame; determining video enhancement parameters of each target video frame according to the high-dimensional semantic features of each target video frame, wherein the video enhancement parameters of any target video frame comprise brightness adjustment parameters, contrast adjustment parameters and sharpness adjustment parameters of the target video frame; enhancing each target video frame according to the video enhancement parameters of each target video frame to obtain an enhanced video frame of each target video frame; synthesizing the enhanced video frames of the target video frames in sequence to obtain an enhanced video of the target video; the video enhancement method further comprises: obtaining statistical features of each target video frame, wherein the statistical features of any target video frame comprise average brightness, contrast and noise level of the target video frame; the determination of the video enhancement parameters of each target video frame according to the high-dimensional semantic features of each target video frame comprises: determining the video enhancement parameters of each target video frame according to the high-dimensional semantic features of each target video frame while combining the statistical features of each target video frame; the determination of the video enhancement parameters of each target video frame according to the high-dimensional semantic features of each target video frame while combining the statistical features of each target video frame comprises: for each target video frame: calculating the mean value of the high-dimensional semantic features of the target video frame, directly combining the mean value of the high-dimensional semantic features of the target video frame and the average brightness of the target video frame to determine the brightness adjustment parameters of the target video frame; calculating the variance of the high-dimensional semantic features of the target video frame, directly combining the variance of the high-dimensional semantic features of the target video frame and the contrast of the target video frame to determine the contrast adjustment parameters of the target video frame; calculating the gradient of the high-dimensional semantic features of the target video frame, directly combining the gradient of the high-dimensional semantic features of the target video frame and the noise level of the target video frame to determine the sharpness adjustment parameters of the target video frame.
2. The video enhancement method of claim 1, wherein, The high-dimensional semantic features of any target video frame can represent the scene information of the target video frame. The scene information of the target video frame comprises some or all of the following semantic information: global scene-level semantic information, local object-level semantic information and action behavior-level semantic information.
3. The video enhancement method of claim 1, wherein, The processing of the target video into a plurality of target video frames comprises: frame dividing the target video to obtain a plurality of original video frames; respectively pre-processing the plurality of original video frames to obtain a plurality of target video frames; wherein the pre-processing of any original video frame comprises one or more of the following processes: removing high-frequency noise in the original video frame, adjusting the resolution of the original video frame to a target resolution, and converting the color space of the original video frame to a target color space.
4. The video enhancement method of claim 1, wherein, The obtaining of the high-dimensional semantic features of each target video frame by performing semantic analysis on each target video frame comprises: for each target video frame: extracting multi-scale local features of the target video frame by using a feature extraction module of a pre-constructed semantic analysis model; Fusing different scale features in the multi-scale local features of the target video frame by using a feature fusion module of the semantic analysis model to obtain multi-scale fused features of the target video frame; Performing attention weight weighting calculation on the multi-scale fused features of the target video frame by using an attention module of the semantic analysis model to obtain attention weighted features of the target video frame; Identifying global scene level semantic information, local object level semantic information, and action behavior level semantic information of the target video frame according to the attention weighted features of the target video frame by using an information recognition module of the semantic analysis model; Constructing high-dimensional semantic features of the target video frame according to the global scene level semantic information, the local object level semantic information, and the action behavior level semantic information of the target video frame.
5. The video enhancement method of claim 1, wherein, The video enhancement parameters of each target video frame are determined according to the high-dimensional semantic features of each target video frame, and the statistical features of each target video frame are combined. The video enhancement parameters of each target video frame are determined according to the high-dimensional semantic features of each target video frame, and the statistical features of each target video frame are combined.
6. The video enhancement method of claim 5, wherein, The video enhancement parameters of each target video frame are determined according to the high-dimensional semantic features of each target video frame, and the statistical features of each target video frame are combined. For each target video frame: The luminance adjustment parameter of the target video frame is determined according to the mean value of the high-dimensional semantic features of the target video frame, the average luminance of the target video frame, and the device adaptation item corresponding to the luminance dimension of the type of the target video playback device; The contrast adjustment parameter of the target video frame is determined according to the variance of the high-dimensional semantic features of the target video frame, the contrast of the target video frame, and the device adaptation item corresponding to the contrast dimension of the type of the target video playback device; The sharpness adjustment parameter of the target video frame is determined according to the gradient of the high-dimensional semantic features of the target video frame, the noise level of the target video frame, and the device adaptation item corresponding to the sharpness dimension of the type of the target video playback device.
7. The video enhancement method of claim 1, wherein, The target video frame is enhanced according to the video enhancement parameter of the target video frame to obtain an enhanced video frame of the target video frame. For each target video frame: The adjusted video frame of the target video frame is obtained by performing luminance adjustment and contrast adjustment on the target video frame pixel by pixel according to the luminance adjustment parameter and the contrast adjustment parameter of the target video frame; The enhanced video frame of the target video frame is obtained by performing sharpness adjustment on the adjusted video frame of the target video frame pixel by pixel according to the sharpness adjustment parameter of the target video frame.
8. A video enhancement apparatus, characterized by comprising: It includes: A video acquisition module, a video processing module, a high-dimensional semantic feature acquisition module, a video enhancement parameter determination module, a video frame enhancement module, and a video frame combination module; The video acquisition module is configured to acquire a target video; The video processing module is configured to process the target video into a plurality of target video frames; The high-dimensional semantic feature acquisition module is configured to obtain high-dimensional semantic features of each target video frame by performing semantic analysis on each target video frame; The video enhancement parameter determination module is configured to determine video enhancement parameters of each target video frame according to the high-dimensional semantic features of each target video frame, wherein the video enhancement parameters of any target video frame include a brightness adjustment parameter, a contrast adjustment parameter, and a sharpness adjustment parameter of the target video frame; The video frame enhancement module is configured to enhance each target video frame according to the video enhancement parameters of each target video frame to obtain an enhanced video frame of each target video frame; The video frame combination module is configured to combine the enhanced video frames of the target video frames in sequence to obtain an enhanced video of the target video; The device further includes a video frame statistical feature acquisition module; The video frame statistical feature acquisition module is configured to obtain statistical features of each target video frame, wherein the statistical features of any target video frame include an average brightness, a contrast, and a noise level of the target video frame; The video enhancement parameter determination module is configured to determine the video enhancement parameters of each target video frame according to the high-dimensional semantic features of each target video frame and in combination with the statistical features of each target video frame; The video enhancement parameter determination module is configured to: For each target video frame, determine a brightness adjustment parameter of the target video frame according to a mean value of the high-dimensional semantic features of the target video frame and in combination with an average brightness of the target video frame; determine a contrast adjustment parameter of the target video frame according to a variance of the high-dimensional semantic features of the target video frame and in combination with a contrast of the target video frame; and determine a sharpness adjustment parameter of the target video frame according to a gradient of the high-dimensional semantic features of the target video frame and in combination with a noise level of the target video frame.
Citation Information
Patent Citations
Image processing method, equipment and medium
CN113763296A
Remote sensing target detection method based on multi-scale feature collaboration and semantic complementation
CN119399424A