Video processing method and device, electronic equipment and storage medium
By integrating video feature information and environment feature information to determine matching stylized templates, the problem of time-consuming search for stylized templates by users in the template library is solved, and the efficiency of video processing is improved.
Patent Information
- Application Number
- CN202510351536.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-20
AI Technical Summary
When there are many stylized templates in the template library, users need to spend more time finding stylized templates that meet their needs, resulting in low video processing efficiency.
By obtaining video feature information and environmental feature information of the shooting environment, the two are integrated to determine the matching stylized template, reducing the user's search time in the template library.
Improves the efficiency of video processing, reduces the time for users to find stylized templates in the template library, and simplifies video processing steps.
Smart Images

Figure CN120186413A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of video processing, and particularly relates to a video processing method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of electronic devices, electronic devices can perform stylization processing on captured videos. For example, lighting effects, rain effects, etc. can be added to the videos. When a user needs to perform stylization processing on a certain video, the user can select a stylization template from multiple stylization templates displayed on the electronic device to add the effect corresponding to the stylization template to the video.
[0003] However, when there are many stylization templates in the template library, the user needs to search back and forth in the template library to find a stylization template that meets the requirements. Therefore, the steps of video processing are cumbersome and time-consuming, and thus the efficiency of video processing is low. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a video processing method, apparatus, electronic device, and storage medium, which can improve the efficiency of video processing.
[0005] In a first aspect, the embodiments of this application provide a video processing method, which includes: obtaining video feature information corresponding to a first video; obtaining environmental feature information corresponding to the shooting environment information when shooting the first video; fusing the video feature information and the environmental feature information to obtain fused feature information; and determining at least one first stylization template that matches the fused feature information.
[0006] In a second aspect, the embodiments of this application provide a video processing apparatus, which includes: an obtaining module and a processing module. The obtaining module is used to obtain video feature information corresponding to a first video; and obtain environmental feature information corresponding to the shooting environment information when shooting the first video. The processing module is used to fuse the video feature information obtained by the obtaining module and the environmental feature information to obtain fused feature information; and determine at least one first stylization template that matches the fused feature information.
[0007] In a third aspect, the embodiments of this application provide an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0008] In a fourth aspect, the embodiments of this application provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0009] Fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run programs or instructions to implement the method described in the first aspect.
[0010] Sixth aspect, an embodiment of the present application provides a computer program / program product, which is stored in a storage medium, and the program / program product is executed by at least one processor to implement the method described in the first aspect.
[0011] In an embodiment of the present application, video feature information corresponding to the first video is obtained; environmental feature information corresponding to the shooting environment information when the first video is shot is obtained; the video feature information and the environmental feature information are fused to obtain fused feature information; at least one first stylized template that matches the fused feature information is determined. In this solution, since the electronic device can fuse the video feature information of the first video and the environmental feature information corresponding to the shooting environment information when the first video is shot, and determine a stylized template suitable for the first video based on the fused feature information, the problem that the user has to search back and forth in the template library to find a stylized template that meets the requirements is avoided, thereby improving the efficiency of video processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is one of the flow diagrams of a video processing method provided by an embodiment of the present application;
[0013] Figure 2 is the second flow diagram of a video processing method provided by an embodiment of the present application;
[0014] Figure 3 is the third flow diagram of a video processing method provided by an embodiment of the present application;
[0015] Figure 4 is the fourth flow diagram of a video processing method provided by an embodiment of the present application;
[0016] Figure 5 is the fifth flow diagram of a video processing method provided by an embodiment of the present application;
[0017] Figure 6 is an example diagram of video processing provided by an embodiment of the present application;
[0018] Figure 7 is the structural diagram of a video processing device provided by an embodiment of the present application;
[0019] Figure 8 is the first hardware structural diagram of an electronic device provided by an embodiment of the present application;
[0020] Figure 9 This is the second schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0021] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.
[0022] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0023] The terms "at least one (item)", "at least one" in the specification and claims of the present application refer to any one, any two or more combinations of the objects they contain. For example, at least one (item) of a, b, and c can represent: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two (items)" means two or more, and its meaning is similar to that of "at least one (item)".
[0024] Next, in conjunction with the accompanying drawings, the video processing method provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.
[0025] The video processing method in the embodiments of the present application can be applied to scenarios where video is stylized.
[0026] In the video processing method, device, electronic device, and storage medium provided by the embodiments of the present application, since the electronic device can fuse the video feature information of the first video and the environmental feature information corresponding to the shooting environment when shooting the first video, and determine a stylized template suitable for the first video based on the fused feature information, the problem that the user has to search back and forth in the template library to find a stylized template that meets the requirements is avoided, thereby improving the efficiency of video processing.
[0027] The execution subject of the video processing method provided by the embodiments of the present application may be a video processing device, which may be an electronic device, or a functional module or entity in the electronic device. Hereinafter, taking an electronic device as an example, the technical solution provided by the embodiments of the present application will be described.
[0028] The embodiments of the present application provide a video processing method. Figure 1 The flowchart of a video processing method provided by the embodiments of the present application is shown. This method can be applied to an electronic device. As Figure 1 shown, the video processing method provided by the embodiments of the present application may include the following steps 201 to 204.
[0029] Step 201: The electronic device acquires video feature information corresponding to the first video.
[0030] In some embodiments of the present application, the above first video may be any one of the following: a landscape video, a food video, an emotional video, a game video, etc. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not limit this.
[0031] In some embodiments of the present application, the above first video may be a video recorded by the user, or a video downloaded from a browser, or a video sent by a certain contact received. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not limit this.
[0032] In the embodiments of the present application, the above video feature information includes image feature information and audio feature information corresponding to each video frame.
[0033] In some embodiments of the present application, when the electronic device receives a second input for the first video, it may acquire the video feature information of the first video and the environmental feature information corresponding to the shooting environment when shooting the first video. Then, the electronic device may fuse the video feature information and the environmental feature information to obtain fused feature information, and then determine at least one first stylized template that matches the fused feature information. Finally, the electronic device may display the template identifier of each first stylized template in at least one first stylized template.
[0034] In some embodiments of the present application, the above second input is used to perform stylization processing on the first video.
[0035] In some examples, the above second input may be an input by the user to a video editing control when the first video is being displayed.
[0036] It can be understood that when the user's demand is to stylize the first video, the electronic device can obtain the video feature information of the first video and the environmental feature information corresponding to the shooting environment when the first video is shot, and then fuse these two feature information, so as to recommend and display a stylized template suitable for the first video to the user based on the fused feature information, so that the user can select a stylized template from the stylized templates recommended by the electronic device to process the first video.
[0037] In some embodiments of the present application, in combination with Figure 1 , such as Figure 2 shown, the above step 201 can be specifically implemented by the following step 201a and step 201b.
[0038] Step 201a: The electronic device obtains the image feature information corresponding to each video frame of the first video.
[0039] In some embodiments of the present application, the above image feature information may include at least one of the following: color feature, texture feature, shape feature, spatial relationship feature, etc.
[0040] In some examples, the above spatial relationship feature describes the relative positions and layouts between different objects in the video frame. The spatial relationship feature can help identify and locate the objects in the video frame and their mutual relationships.
[0041] In some embodiments of the present application, the image feature information corresponding to each video frame above can be used to characterize the key information in each video frame.
[0042] For example: Suppose the first video includes video frame A, and the image feature information of video frame A can be used to characterize that "video frame A includes sky, sea and a little girl, and the little girl stands by the sea".
[0043] In some embodiments of the present application, the electronic device can decompose the first video into continuous video frames through a video decoding algorithm.
[0044] In some embodiments of the present application, the electronic device can perform semantic segmentation on each video frame through semantic segmentation technology to accurately divide different objects and scene areas in each video frame, so as to obtain the image feature information corresponding to each video frame of the first video.
[0045] It should be noted that performing semantic segmentation technology on a video frame means performing pixel-level classification on the video frame to divide the video frame into different semantic regions, such as foreground objects, background, etc.
[0046] In some embodiments of the present application, when the electronic device performs semantic segmentation on each video frame through semantic segmentation technology, it can also combine an object detection algorithm to identify key objects (such as people, vehicles, buildings, etc.) in each video frame, so as to accurately obtain the image feature information corresponding to each video frame of the first video.
[0047] In some embodiments of the present application, the electronic device can extract image features from each video frame through deep learning algorithms such as convolutional neural networks, and identify information such as objects and spatial structures in each video frame. For example, elements such as people, buildings, and natural landscapes and their positional relationships can be accurately identified from each video frame.
[0048] Step 201b: The electronic device obtains the audio feature information corresponding to the first video.
[0049] In some embodiments of the present application, the above audio feature information may include at least one of the following: time-domain features, frequency-domain features, time-frequency domain features, statistical features, and filter bank features, etc.
[0050] In some examples, the above time-domain features mainly reflect the changes of the sound signal on the time axis. The time-domain waveform diagram can show the amplitude changes of the sound signal at different time points, while the energy diagram shows the energy distribution of the sound signal in different time periods.
[0051] In some examples, the above frequency-domain features mainly focus on the changes of the sound signal in the frequency domain. The Fourier transform is a commonly used frequency-domain analysis method, which can convert the sound signal from the time domain to the frequency domain. Common frequency-domain features include spectrogram, frequency features, and spectral envelope.
[0052] In some examples, the above time-frequency domain features are the results of joint time-frequency analysis of the sound signal. The short-time Fourier transform and continuous wavelet transform are commonly used time-frequency domain analysis methods. The time-frequency diagram can show the changes of the sound signal in time and frequency, and can better observe the time-frequency characteristics of the sound signal.
[0053] In some examples, the above statistical features describe the basic features of the sound by calculating statistical quantities of the audio signal, such as mean, variance, maximum value, minimum value, etc.
[0054] In some examples, the electronic device can use a set of filters to decompose the sound signal of the first video, and extract corresponding feature parameters on the output of each filter, that is, the above filter bank features, to better capture the features of the sound signal in different frequency ranges.
[0055] In some embodiments of the present application, the above audio feature information can be used to characterize the characteristics of the audio signal.
[0056] For example, the audio feature information of the first video can be used to characterize that the music rhythm of the first video is relatively lively.
[0057] In some embodiments of the present application, the electronic device can extract the sound signal of the first video through the audio processing module to obtain the audio feature information corresponding to the sound signal.
[0058] In some embodiments of the present application, after the electronic device extracts the sound signal of the first video through the audio processing module, it can perform preprocessing operations such as noise reduction and gain on the sound signal, and then obtain the audio feature information corresponding to the preprocessed sound signal.
[0059] In some embodiments of the present application, the electronic device can obtain the audio feature information corresponding to the sound signal of the first video through an audio analysis algorithm.
[0060] In some examples, the above audio analysis algorithm can be Fourier transform.
[0061] In some embodiments of the present application, the above audio feature information can be used to determine the scene where the electronic device is located when the first video is shot.
[0062] In some examples, when the audio feature information corresponding to the first video is used to characterize that the music rhythm of the first video is relatively lively, the electronic device can determine that the electronic device is in a lively scene when the first video is shot.
[0063] In some examples, when the audio feature information corresponding to the first video is used to characterize that the music rhythm of the first video is relatively low, the electronic device can determine that the electronic device is in a serious or quiet scene when the first video is shot.
[0064] It should be noted that the electronic device can execute step 201a first and then step 201b; or, execute step 201b first and then step 201a. The embodiments of the present application do not limit the execution order of step 201a and step 201b.
[0065] Step 202: The electronic device obtains the environmental feature information corresponding to the shooting environment information when the first video is shot.
[0066] In some embodiments of the present application, the above shooting environment information includes at least one of the following: gyroscope data collected by the electronic device when the first video is shot, the position information of the electronic device when the first video is shot, and the weather information of the location where the electronic device is located when the first video is shot.
[0067] In some embodiments of the present application, when shooting the first video, the electronic device can collect gyroscope data through the gyroscope of the electronic device.
[0068] In some embodiments of the present application, when shooting the first video, the electronic device may obtain the location information of the location where the electronic device is located through the Global Positioning System (GPS) module of the electronic device.
[0069] In some embodiments of the present application, the above weather information may include at least one of the following: weather (e.g., cloudy, overcast, sleet), temperature, air pressure, humidity, wind force level, etc. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not limit this.
[0070] In some embodiments of the present application, the electronic device may query the real-time weather information of the location where the electronic device is located when shooting the first video by connecting to the network. Alternatively, the electronic device may query the real-time weather information of the location where the electronic device is located when shooting the first video according to the location where the electronic device is located by calling the weather data API.
[0071] In some examples, the above weather data API may be any one of the following: OpenWeatherMap, Weather.com.
[0072] It should be noted that the electronic device may first execute step 201 and then execute step 202; or first execute step 202 and then execute step 201. The embodiments of the present application do not limit the execution order of step 201 and step 202.
[0073] In some embodiments of the present application, the above step 202 may be specifically implemented by at least one of the following steps 202a to 202c.
[0074] Step 202a: When the shooting environment information includes gyroscope data, the electronic device analyzes the gyroscope data to obtain attitude feature information.
[0075] In the embodiments of the present application, the above attitude feature information is used to characterize the attitude of the electronic device when shooting the first video.
[0076] In some embodiments of the present application, the electronic device may analyze the gyroscope data to determine the changes such as the movement and rotation of the electronic device when shooting the first video, so as to obtain the above attitude feature information. That is, the gyroscope data provides a basis for understanding the dynamic changes of the electronic device during shooting.
[0077] For example: the above attitude feature information may be used to characterize that the electronic device is smoothly translating to the right when shooting the first video.
[0078] Step 202b: When the shooting environment information includes location information, the electronic device encodes the location information to obtain location feature information.
[0079] In an embodiment of the present application, the above position feature information is used to represent the position of the electronic device when the first video is shot.
[0080] For example, the above position feature information can be used to represent that the electronic device is in an ancient town when the first video is shot.
[0081] In some embodiments of the present application, the above encoding of the position information by the electronic device can be understood as: the electronic device performs data calibration and format conversion on the position information to make it meet the requirements of subsequent processing.
[0082] Step 202c: When the shooting environment information includes weather information, the electronic device encodes the weather information to obtain weather feature information.
[0083] In an embodiment of the present application, the above weather feature information is used to represent the weather condition of the position where the electronic device is located when the first video is shot.
[0084] For example: Assume that the electronic device is in an ancient town when the first video is shot, then the above weather feature information can be used to represent that the weather condition of the ancient town where the electronic device is located when the first video is shot is "light rain, temperature 20°".
[0085] In some embodiments of the present application, the above encoding of the weather information by the electronic device can be understood as: the electronic device performs data calibration and format conversion on the weather information to make it meet the requirements of subsequent processing.
[0086] In an embodiment of the present application, the above environmental feature information includes at least one of the following: attitude feature information, position feature information, and weather feature information.
[0087] Step 203: The electronic device fuses the video feature information and the environmental feature information to obtain fused feature information.
[0088] In some embodiments of the present application, the above fused feature information is used to represent the key information of the first video.
[0089] For example: The above fused feature information can be used to represent "sunny, by the sea, a little girl is walking by the sea".
[0090] In some embodiments of the present application, the electronic device can perform weighted fusion on the video feature information and the environmental feature information to obtain the above fused feature information.
[0091] In some embodiments of the present application, the electronic device can perform weighted fusion on the video feature information and the environmental feature information through the following formula (1) to obtain the above fused feature information.
[0092] Formula 1: F = α·I + β·A + γ·G + δ·P + ∈·W
[0093] Wherein, F is the fused feature information, I is the image feature information, and α is the weight coefficient of the image feature information I; A is the audio feature information, and β is the weight coefficient of the audio feature information A; G is the pose feature information, and γ is the weight coefficient of the pose feature information G; P is the position feature information, and δ is the weight coefficient of the position feature information P; W is the weather feature information, and ∈ is the weight coefficient of the weather feature information W.
[0094] In some embodiments of the present application, the above image feature information I = [i1, i2, …, in], and in may be the image feature information of the nth video frame of the first video; the audio feature information A = [a1, a2, …, am], and am may be any one of the following: time-domain feature, frequency-domain feature, time-frequency domain feature, statistical feature, and filter bank feature; the pose feature information G = [g1, g2, …, gk], and gk may be the pose feature of the electronic device when the kth video frame of the first video is captured; the position feature information P = [p1, p2, …, pl], and pl may be the position information of the lth position point where the electronic device is located when the first video is captured; the weather feature information W = [w1, w2, …, ws], and ws may be any one of the following: weather feature, temperature feature, pressure feature, humidity feature, wind force level feature.
[0095] In the embodiments of the present application, α + β + γ + δ + ∈ = 1.
[0096] In some embodiments of the present application, the weight coefficients α, β, γ, δ, and ∈ are optimized and solved through a large amount of training data, so that the fused feature information can most effectively express the information of the shooting scene, providing strong support for subsequent scene understanding and stylization processing.
[0097] In some embodiments of the present application, the electronic device may adjust the weight coefficients of the video feature information and the environmental feature information based on the position information of the electronic device when the first video is captured.
[0098] Exemplarily, if the electronic device is at the seaside when the first video is captured, the video elements of the first video (e.g., ocean elements, beach elements) are the most critical for judging the shooting scene. Therefore, the electronic device may appropriately increase the value of the weight coefficient α of the image feature information.
[0099] Exemplarily, if the electronic device is at a music scene when the first video is captured, the audio feature information of the first video is very important for judging the atmosphere of the shooting scene. Therefore, the electronic device may appropriately increase the value of the weight coefficient β of the audio feature information.
[0100] In some embodiments of the present application, the electronic device can comprehensively analyze video feature information and environmental feature information through a machine learning model to provide rich and accurate context support for the stylization processing of the first video.
[0101] In some embodiments of the present application, the electronic device can determine the scene where the electronic device is located when shooting the first video by fusing audio feature information and pose feature information.
[0102] For example: when the audio feature information characterizes that the audio rhythm of the first video is fast and the pose feature information characterizes that the electronic device is moving fast, the electronic device can determine that the electronic device is in a scene of intense movement or lively activity when shooting the first video.
[0103] In some embodiments of the present application, the electronic device can determine the scene where the electronic device is located when shooting the first video by fusing the image feature information, audio feature information, and pose feature information of each video frame.
[0104] For example: when the image feature information of the video frame of the first video characterizes that there is a large area of water in the video frame, the audio feature information characterizes that there is the sound of waves in the first video, and the pose feature information characterizes that the electronic device is in a stable pose near the water surface, the electronic device can determine that the electronic device is by the sea when shooting the first video.
[0105] Step 204: The electronic device determines at least one first stylization template that matches the fused feature information.
[0106] For example: when the fused feature information characterizes "sunny day, park, big tree, noisy sound", the first stylization template can be a cartoon style template with the park as the theme. After processing the first video with this template, the big tree in the video can be turned into a tree sprite.
[0107] In some embodiments of the present application, in combination with Figure 1 , as Figure 3 shown, the above step 204 can be specifically implemented by the following step 204a and step 204b.
[0108] Step 204a: The electronic device matches the fused feature information with the template feature information corresponding to each stylization template stored in the template library.
[0109] It can be understood that there are multiple stylization templates stored in the template library of the electronic device, and each stylization template corresponds to a template feature information.
[0110] In some embodiments of the present application, the template feature information corresponding to the stylization template is used to characterize the design style, application scenario, content elements, etc. of the stylization template.
[0111] For example, the template feature information can be used to represent that "the corresponding stylized template is in a cartoon style, suitable for the scenario of portrait shooting, and includes cartoon headdress elements", etc.
[0112] Step 204b: The electronic device determines the stylized template corresponding to the template feature information that matches the fusion feature information as the first stylized template.
[0113] In some embodiments of the present application, when the matching degree between the template feature information corresponding to any stylized template and the fusion feature information is greater than a preset threshold, the electronic device may determine any stylized template as the first stylized template.
[0114] In this way, since the electronic device fuses the video feature information of the first video and the environmental feature information corresponding to the shooting environment when shooting the first video, it can search in the template library for the stylized template corresponding to the template feature information that matches the fusion feature information, that is, search for the stylized template suitable for the first video, and then recommend the stylized template to the user. Therefore, the efficiency of processing the video using the stylized template is improved.
[0115] In some embodiments of the present application, in combination Figure 1 , such as Figure 4 shown, the above step 204 can be specifically implemented by the following steps 204c to 204e.
[0116] Step 204c: The electronic device obtains user preference information.
[0117] In the embodiments of the present application, the above user preference information is used to indicate the template type of the stylized template that the user historically preferred.
[0118] In some embodiments of the present application, the template type of the above stylized template may be any one of the following: cartoon type, comic type, fantasy type, film type, rainy day type, etc. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not limit this.
[0119] In some embodiments of the present application, the electronic device may obtain the stylized template historically selected by the user, and then based on the stylized template historically selected by the user, determine the template type of the stylized template that the user historically preferred.
[0120] In some embodiments of the present application, the electronic device may establish a user preference database to record the stylized templates historically selected by the user, and then the electronic device may determine the template type of the stylized template that the user historically preferred based on the data recorded in the user preference database, that is, determine the above user preference information.
[0121] Step 204d: The electronic device adjusts the fused feature information based on the user preference information to obtain the adjusted fused feature information.
[0122] In some embodiments of the present application, the electronic device may adjust the feature parameters of the fused feature information based on the user preference information to obtain the adjusted fused feature information.
[0123] For example: Suppose the above fused feature information is used to represent "sunny day, by the sea, sound of waves, little girl walking by the sea". When the user preference information indicates that the template type of the stylized template preferred by the user history is the rainy day type, the electronic device adjusts the fused feature information based on the user preference information, and the obtained adjusted fused feature information can be used to represent "rainy day, by the sea, sound of waves, little girl walking by the sea".
[0124] Step 204e: The electronic device determines at least one first stylized template that matches the adjusted fused feature information.
[0125] In some embodiments of the present application, the electronic device may match the adjusted fused feature information with the template feature information corresponding to each stylized template stored in the template library, and then determine the stylized template corresponding to the template feature information that matches the adjusted fused feature information as the first stylized template.
[0126] For example: When the above fused feature information is used to represent "sunny day, park, big tree, noisy sound", and the adjusted fused feature information is used to represent "rainy day, park, big tree, noisy sound", the first stylized template that matches the adjusted fused feature information may be a rainy style template. After processing the first video with this template, a raindrop effect can be added to the video to imitate the rainy weather.
[0127] In this way, since the electronic device obtains scene information such as scene type and atmosphere through the fusion analysis of multimodal data (i.e., the video frames of the above first video, the sound signal of the first video, and the shooting environment information), it can combine the user's historical preference information to accurately recommend suitable stylized templates for the user. It not only considers the weather and geographical location, but also recommends based on the user's past habits. For example, if the shooting weather of the video is rainy and the user usually selects a bright style in the past, the electronic device preferentially recommends a "sunny day" stylized effect, replaces the gloomy sky on a rainy day with a clear sky, and adjusts the light, shadow, and color to create a bright atmosphere; if the shooting weather of the video is sunny and the user often selects a style with a strong sense of art, the electronic device recommends a "dusk" stylized effect to increase the warmth and artistic sense of the picture.
[0128] In the video processing method provided by the embodiments of the present application, since the electronic device can fuse the video feature information of the first video and the environmental feature information corresponding to the shooting environment when shooting the first video, and determine a stylized template suitable for the first video based on the fused feature information, the problem that the user has to search back and forth in the template library to find a stylized template that meets the requirements is avoided, thereby improving the efficiency of video processing.
[0129] In some embodiments of the present application, after the electronic device determines at least one first stylized template that matches the fused feature information, it can display the template identifier of each first stylized template in the at least one first stylized template.
[0130] In some examples, the electronic device can display the template identifier of each first stylized template in the at least one first stylized template in the interface for displaying the first video.
[0131] In this way, since the electronic device can display the template identifier of each first stylized template in the at least one first stylized template after determining at least one first stylized template suitable for the first video, the user can select a stylized template from the stylized templates recommended by the electronic device that are suitable for the first video to process the first video, avoiding the problem that the user has to search back and forth in the template library to find a stylized template that meets the requirements, thereby improving the efficiency of video processing.
[0132] In some embodiments of the present application, in combination Figure 1 , such as Figure 5 shown, after step 204 above, the video processing method provided by the embodiments of the present application further includes the following step 301 and step 302.
[0133] Step 301: The electronic device receives a first input of the template identifier of the second stylized template in the at least one first stylized template from the user.
[0134] In some embodiments of the present application, the above first input is used to select the second stylized template.
[0135] In some embodiments of the present application, the above first input includes but is not limited to: the user's touch input on the template identifier of the second stylized template through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or a click input, or other feasible inputs. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not limit this.
[0136] In some embodiments of the present application, the above specific gesture can be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture.
[0137] Step 302: In response to the first input, the electronic device processes the first video using the second stylized template to obtain a second video.
[0138] It should be noted that for the detailed steps of the electronic device processing the first video using the second stylized template to obtain a second video, reference can be made to the description in the related art of the electronic device processing a video using a stylized template to obtain a video with the template effect added, which will not be elaborated here.
[0139] In some embodiments of the present application, after the above step 204, the video processing method provided by the embodiments of the present application further includes the following step 401 and step 402.
[0140] Step 401: The electronic device processes the first video using the second stylized template in at least one first stylized template to obtain a second video.
[0141] It should be noted that for the detailed steps of the electronic device processing the first video using the second stylized template to obtain a second video, reference can be made to the description in the related art of the electronic device processing a video using a stylized template to obtain a video with the template effect added, which will not be elaborated here.
[0142] Step 402: The electronic device adjusts the animation effect of the video elements in the second video based on the fusion feature information to obtain a third video.
[0143] In some embodiments of the present application, the above video elements may be at least one of the following: light and shadow, plants, animals, people, daily necessities, special effect elements, etc. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not limit this.
[0144] In some embodiments of the present application, the above special effect elements may be at least one of the following: raindrop special effect, snowflake special effect, light and shadow special effect, expression special effect, etc. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not limit this.
[0145] In some embodiments of the present application, the above video elements may be the original video elements of the first video, or the newly added special effect elements after processing the first video using the second stylized template.
[0146] In some embodiments of the present application, the animation effect of the above video elements may be at least one of the following: the movement speed of the video elements, the display angle of the video elements, the display direction of the video elements, the display form of the video elements, the flicker frequency of the video elements, the light and shadow effect of the video elements, the jitter effect of the video elements.
[0147] In some examples, when the above-mentioned fused feature information indicates that the electronic device is tilted or rotated while shooting the first video, the electronic device can adjust the display direction and display angle of the video elements.
[0148] For example: Suppose the video element is light and shadow. When the fused feature information indicates that the electronic device is tilted or rotated while shooting the first video, the electronic device can correspondingly adjust the display direction and display angle of the light and shadow to enhance the dynamic sense of the visual effect.
[0149] In some examples, when the above-mentioned fused feature information indicates that the electronic device shakes violently while shooting the first video, the electronic device can adjust the jitter effect of the video elements.
[0150] In some examples, when the above-mentioned fused feature information indicates that the weather at the location where the electronic device is located while shooting the first video is cloudy, the electronic device can dim the light and shadow effect of the video elements (for example: reduce the brightness and contrast of the video elements).
[0151] In some examples, when the above-mentioned fused feature information indicates that the audio rhythm of the first video is accelerated, the electronic device can increase the flicker frequency of the video elements.
[0152] For example: When the video element is the stars in the sky, the electronic device can increase the flicker frequency of the stars by a certain proportion. Or, when the video element is a raindrop special effect, the electronic device can increase the falling speed of the raindrops by a certain proportion.
[0153] In this way, since the electronic device processes the first video using the second stylized template selected by the user and obtains the second video, the electronic device can also adjust the animation effect of the video elements in the second video based on the fused feature information. Therefore, the stylized effect of the video can be made more natural and coherent, thereby improving the reliability of video processing.
[0154] In some embodiments of the present application, the electronic device can process the data of each video frame, sound data, and sensor data (i.e., the above-mentioned gyroscope data, position information, etc.) in real time to ensure the coherence and naturalness of the stylized effect. An efficient data processing architecture is adopted to perform parallel processing on continuous video frames, audio signals, and sensor data to reduce processing latency.
[0155] In some embodiments of the present application, after obtaining the second video, the electronic device can track the motion trajectory of the shooting object in continuous video frames by means of optical flow method, so as to adjust the display position and display form of the special effect elements added to the shooting object according to the motion speed and motion direction of the shooting object, ensuring the natural and smooth transition of the special effect elements added to the shooting object between different video frames.
[0156] For example, when the object being photographed is a football, the special effect elements of the football (e.g., a ring of flames) should be presented coherently along with the ball's movement trajectory.
[0157] In some embodiments of the present application, the electronic device may adjust the video effect of the second video based on weather feature information.
[0158] In some examples, when the weather feature information indicates that the weather at the location where the electronic device is located when shooting the second video is rainy, the electronic device may simulate the raindrop effect through a particle system and add it to the video frames of the second video.
[0159] In some examples, when the weather feature information indicates that the weather at the location where the electronic device is located when shooting the second video is sunny, the electronic device may use an image enhancement algorithm to enhance the light and shadow contrast of the second video and highlight the three-dimensional and hierarchical sense of the scene.
[0160] In some embodiments of the present application, the electronic device may call the corresponding visual effect optimization module based on different weather feature information to adjust the video effect of the second video.
[0161] In some embodiments of the present application, the electronic device may adjust the video effect of the second video based on location feature information.
[0162] In some examples, when the location feature information indicates that the location where the electronic device is located when shooting the second video is an ancient town with historical and cultural characteristics, the electronic device may enhance the retro tone of the second video to highlight the simple charm of the ancient town.
[0163] Furthermore, when the location feature information indicates that the location where the electronic device is located when shooting the second video is an ancient town with historical and cultural characteristics and the weather feature information indicates that the weather at the location where the electronic device is located when shooting the second video is rainy, the electronic device may also add raindrop special effects and a hazy light and shadow effect to the second video to create a poetic atmosphere.
[0164] In some embodiments of the present application, after obtaining the second video, the electronic device may adjust the stylization parameters of the second video to optimize the video effect of the second video.
[0165] In some embodiments of the present application, the electronic device may establish a mapping relationship database between stylization parameters and location feature information and weather feature information. Then, after obtaining the second video, the electronic device may determine, based on the location feature information and weather feature information corresponding to the first video, the stylization parameters that match the location feature information and weather feature information corresponding to the first video from the database, and then process the second video using the stylization parameters.
[0166] For example, assume that the location feature information corresponding to the first video indicates that the shooting location of the first video is a desert, and the weather feature information corresponding to the first video indicates that the weather when the first video was shot is sunny. Then, the electronic device can retrieve from the database stylization parameters suitable for a sunny day in the desert, such as parameters for increasing warm tones and enhancing light and shadow contrast, for application to subsequent stylization processing.
[0167] In an embodiment of the present application, as Figure 6 shown, after the user shoots and uploads the first video, the electronic device can display a multi-modal recognition of the first video, and the multi-modal recognition includes: GPS recognition, weather recognition, video content semantic segmentation recognition, audio content recognition, gyroscope data recognition, etc.; then the electronic device can integrate and understand the multi-modal data, that is, F = α·I + β·A + γ·G + δ·P + ∈·W; then the electronic device can, based on F and the user preference database, search the template library for a stylization template suitable for the first video and recommend it to the user; then the user can select a stylization template from at least one stylization template recommended by the electronic device to trigger the electronic device to process the first video using the stylization template selected by the user to obtain the stylized first video; then the electronic device can optimize the dynamic effect of the stylized first video, that is, perform parallel processing on consecutive video frames, audio signals, and sensor data to ensure the coherence and naturalness of the stylization effect; then the electronic device can display a video preview after the dynamic effect optimization; the user can fine-tune the video after the dynamic effect optimization; finally, the electronic device outputs the finished video.
[0168] In an embodiment of the present application, the following effects can be achieved: 1. The stylization effect is more natural: By fusing multi-modal data and understanding the scene, the problem of unnatural and lack of coherence in the stylization effect in traditional technologies is solved. The multi-modal data provides rich information, enabling the stylization process to closely fit the actual scene and avoiding the disconnection between the style and the content. 2. The user experience is better: By combining geographical location and weather information, the stylization effect is dynamically adjusted to provide a more personalized and immersive user experience. Users can feel the high degree of fit between the stylization effect and the actual shooting environment, enhancing the fun and sense of satisfaction in creation. The intelligent recommendation function also greatly improves the efficiency of users in obtaining their desired styles. 3. The technology has stronger adaptability: It supports cross-platform and multi-device adaptation to ensure the consistency of the stylization effect on different devices. Through optimizing the algorithm and data processing flow, the technology of the present invention can stably operate on various mainstream mobile devices, providing users with a unified high-quality service. 4. Intelligent interaction design: Through an intelligent interaction interface, a more intelligent and user-friendly user experience is provided. Users can easily obtain personalized stylization suggestions, laying a foundation for subsequent user operations and style transformation.
[0169] In the embodiments of the present application, through multi-modal data fusion and scene understanding technology, combined with geographical location and weather information, the stylization effect of dynamic images is optimized, the visual and auditory coherence is enhanced, and at the same time, an intelligent user interaction experience is provided. This technology has high innovation and practicality, and can meet the needs of users for high-quality dynamic image processing.
[0170] The following illustrates the solution of the present application through specific embodiments:
[0171] Embodiment 1: Stylization processing of a seaside sunset scene
[0172] S1. The user shoots a video of a seaside sunset.
[0173] S2. The electronic device obtains the shooting location as the seaside through GPS positioning and queries that the current weather is sunny.
[0174] S3. The electronic device combines elements such as the ocean, beach, and sky in the video frame to dynamically adjust the stylization effect and increase the light and shadow effect of the setting sun's afterglow. Using semantic segmentation and object detection models, the areas of the ocean, beach, and sky are accurately identified, and then according to the characteristics of a sunny seaside sunset, the warm tone of the setting sun's afterglow is increased through an image enhancement algorithm, the contrast of the light and shadow is improved, and the effect of the golden sunlight shining on the sea surface is highlighted.
[0175] S4. The electronic device combines the audio rhythm and the device posture to dynamically adjust the blinking frequency and light and shadow jitter effect of the stylized elements. Analyze the audio rhythm. If the rhythm is slow, appropriately reduce the blinking frequency of the stylized elements such as the blinking shells on the beach to create a peaceful atmosphere; according to the device posture data, if the device shakes slightly, correspondingly adjust the jitter amplitude of the light and shadow to simulate the light and shadow changes under the real sea breeze.
[0176] S5. Assuming that the user often selects a dreamy style in the past, the electronic device recommends a seaside sunset style enhanced with dreamy elements through an intelligent interaction interface, such as adding blinking jellyfish on the sea surface and a faint fantasy halo appearing in the sky.
[0177] Embodiment 2: Stylization processing of an urban rainy day scene
[0178] S11. The user shoots a video of an urban street.
[0179] S12. The electronic device obtains the shooting location as the city through GPS positioning and queries that the current weather is rainy.
[0180] S13. The electronic device combines elements such as buildings, streets, and pedestrians in the video frame, dynamically adjusts the stylized effect, and adds raindrop effects and ground reflections. It identifies areas such as buildings, streets, and pedestrians through a semantic segmentation model, uses a particle system to simulate realistic raindrop effects and adds them to the corresponding areas; at the same time, it enhances the ground reflection effect through an image rendering algorithm to simulate the real scene of a wet and slippery ground on a rainy day.
[0181] S14. The electronic device combines the audio rhythm and the device posture to dynamically adjust the blinking frequency and light and shadow jitter effect of the stylized elements. If the audio rhythm is fast, it may indicate that pedestrians on the street are in a hurry. At this time, it increases the blinking frequency of stylized elements such as the lights of roadside store signs to create a busy atmosphere; according to the device posture data, such as the device rising and falling with the steps of pedestrians, it correspondingly adjusts the jitter effect of the light and shadow to make the picture more dynamic.
[0182] S15. If the electronic device determines that the user prefers a bright and fresh style and based on the rainy day scene, it recommends through the intelligent interaction interface to convert the rainy day scene into a style effect similar to that after the rain has just cleared, brightens the overall tone, and enhances the clarity of buildings and streets.
[0183] It should be noted that the above-mentioned various method embodiments, or various possible implementation manners in the various method embodiments, can be executed separately, or, on the premise of no contradiction, can also be executed in combination with each other, which can be specifically determined according to actual usage requirements, and the embodiments of the present application do not limit this.
[0184] It should be noted that for the video processing method provided in the embodiments of the present application, the execution subject can be a video processing device. In the embodiments of the present application, taking the video processing device executing the video processing method as an example, the video processing device provided in the embodiments of the present application is described.
[0185] Figure 7 Shows a possible structural schematic diagram of the video processing device involved in the embodiments of the present application. As Figure 7 shown, the video processing device 70 may include: an acquisition module 71 and a processing module 72.
[0186] Among them, the acquisition module 71 is used to acquire the video feature information corresponding to the first video; and acquire the environmental feature information corresponding to the shooting environment information when shooting the first video.
[0187] The processing module 72 is used to fuse the video feature information and the environmental feature information acquired by the acquisition module 71 to obtain the fused feature information; and determine at least one first stylized template that matches the fused feature information.
[0188] An embodiment of the present application provides a video processing device. Since an electronic device can fuse the video feature information of the first video and the environmental feature information corresponding to the shooting environment information when shooting the first video, and determine a stylized template suitable for the first video based on the fused feature information, it avoids the problem that the user has to search back and forth in the template library to find a stylized template that meets the requirements, thereby improving the efficiency of video processing.
[0189] In a possible implementation manner, the obtaining module 71 is specifically configured to obtain the image feature information corresponding to each video frame of the first video; and obtain the audio feature information corresponding to the first video; wherein, the video feature information includes the image feature information and the audio feature information corresponding to each video frame.
[0190] In a possible implementation manner, the shooting environment information includes at least one of the following: gyroscope data collected by the electronic device when shooting the first video, the position information of the electronic device when shooting the first video, and the weather information of the location where the electronic device is located when shooting the first video.
[0191] In a possible implementation manner, the obtaining module 71 is specifically configured to:
[0192] In the case that the shooting environment information includes gyroscope data, analyze the gyroscope data to obtain attitude feature information, and the attitude feature information is used to characterize the attitude of the electronic device when shooting the first video;
[0193] In the case that the shooting environment information includes position information, encode the position information to obtain position feature information, and the position feature information is used to characterize the position of the electronic device when shooting the first video;
[0194] In the case that the shooting environment information includes weather information, encode the weather information to obtain weather feature information, and the weather feature information is used to characterize the weather condition of the location where the electronic device is located when shooting the first video;
[0195] Wherein, the environmental feature information includes at least one of the following: attitude feature information, position feature information, and weather feature information.
[0196] In a possible implementation manner, the processing module 72 is specifically configured to match the fused feature information with the template feature information corresponding to each stylized template stored in the template library; and determine the stylized template corresponding to the template feature information that matches the fused feature information as the first stylized template.
[0197] In a possible implementation, the obtaining module 71 is further configured to obtain user preference information, where the user preference information is used to indicate the template type of the stylized template that the user historically preferred. The processing module 72 is specifically configured to adjust the fused feature information based on the user preference information obtained by the obtaining module 71 to obtain the adjusted fused feature information; and determine at least one first stylized template that matches the adjusted fused feature information.
[0198] In a possible implementation, the processing module 72 is further configured to, after determining at least one first stylized template that matches the fused feature information, process the first video using a second stylized template among the at least one first stylized templates to obtain a second video; and adjust the animation effect of the video elements in the second video based on the fused feature information to obtain a third video.
[0199] The video processing apparatus in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than the terminal. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an Augmented Reality (AR) / Virtual Reality (VR) device, a robot, a wearable device, an Ultra-Mobile Personal Computer (UMPC), a netbook, or a Personal Digital Assistant (PDA), etc., and may also be a server, a Network Attached Storage (NAS), a Personal Computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0200] The video processing apparatus in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0201] The video processing apparatus provided in the embodiments of the present application can implement each process implemented in the above method embodiments. To avoid repetition, it will not be described in detail here.
[0202] Optionally, as Figure 8As shown in the figure, an embodiment of the present application further provides an electronic device 900, including a processor 901 and a memory 902. A program or instruction that can run on the processor 901 is stored on the memory 902. When the program or instruction is executed by the processor 901, each step of the above method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described in detail here.
[0203] It should be noted that the electronic device in the embodiment of the present application includes the above-mentioned mobile electronic device and non-mobile electronic device.
[0204] Figure 9 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.
[0205] The electronic device 100 includes but is not limited to: a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110 and other components.
[0206] Those skilled in the art can understand that the electronic device 100 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 110 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 9 The structure of the electronic device shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, which will not be described in detail here.
[0207] Among them, the processor 110 is used to obtain the video feature information corresponding to the first video; and obtain the environmental feature information corresponding to the shooting environment information when shooting the first video; and fuse the video feature information and the environmental feature information to obtain fused feature information; and determine at least one first stylized template that matches the fused feature information.
[0208] An embodiment of the present application provides an electronic device. Since the electronic device can fuse the video feature information of the first video and the environmental feature information corresponding to the shooting environment information when shooting the first video, and determine a stylized template suitable for the first video based on the fused feature information, the problem that the user has to search back and forth in the template library to find a stylized template that meets the requirements is avoided, thereby improving the efficiency of video processing.
[0209] In some embodiments of the present application, the processor 110 is specifically used to obtain the image feature information corresponding to each video frame of the first video; and obtain the audio feature information corresponding to the first video; wherein, the video feature information includes the image feature information corresponding to each video frame and the audio feature information.
[0210] In some embodiments of the present application, the shooting environment information includes at least one of the following: gyroscope data collected by the electronic device when shooting the first video, location information of the electronic device when shooting the first video, and weather information of the location where the electronic device is located when shooting the first video.
[0211] In some embodiments of the present application, the processor 110 is specifically configured to:
[0212] When the shooting environment information includes gyroscope data, analyze the gyroscope data to obtain attitude feature information, where the attitude feature information is used to characterize the attitude of the electronic device when shooting the first video;
[0213] When the shooting environment information includes location information, encode the location information to obtain location feature information, where the location feature information is used to characterize the location where the electronic device is located when shooting the first video;
[0214] When the shooting environment information includes weather information, encode the weather information to obtain weather feature information, where the weather feature information is used to characterize the weather condition of the location where the electronic device is located when shooting the first video;
[0215] Among them, the environmental feature information includes at least one of the following: attitude feature information, location feature information, and weather feature information.
[0216] In some embodiments of the present application, the processor 110 is specifically configured to match the fusion feature information with the template feature information corresponding to each stylized template stored in the template library; and determine the stylized template corresponding to the template feature information that matches the fusion feature information as the first stylized template.
[0217] In some embodiments of the present application, the processor 110 is specifically configured to obtain user preference information, where the user preference information is used to indicate the template type of the stylized template that the user historically preferred; and based on the user preference information, adjust the fusion feature information to obtain adjusted fusion feature information; and determine at least one first stylized template that matches the adjusted fusion feature information.
[0218] In some embodiments of the present application, the processor 110 is further configured to, after determining at least one first stylized template that matches the fusion feature information, process the first video using a second stylized template among the at least one first stylized template to obtain a second video; and based on the fusion feature information, adjust the animation effect of the video elements in the second video to obtain a third video.
[0219] The electronic device provided by the embodiment of the present application can implement each process implemented by the above method embodiment and achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0220] For the beneficial effects of various implementation manners in this embodiment, reference may be specifically made to the beneficial effects of the corresponding implementation manners in the above method embodiment. To avoid repetition, it will not be elaborated here.
[0221] It should be understood that in the embodiment of the present application, the input unit 104 may include a Graphics Processing Unit (GPU) 1041 and a microphone 1042. The graphics processor 1041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also referred to as a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. The other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.
[0222] The memory 109 can be used to store software programs and various data. The memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 109 may include a volatile memory or a non-volatile memory, or the memory 109 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 109 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.
[0223] The processor 110 may include one or more processing units; optionally, the processor 110 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 110 either.
[0224] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above method embodiments is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.
[0225] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs.
[0226] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above method embodiment, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0227] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip.
[0228] The embodiments of the present application provide a computer program / program product. The computer program / program product is stored in a storage medium. The program / program product is executed by at least one processor to implement each process of the above method embodiment, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0229] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of other identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed. It may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0230] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method of the embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0231] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A video processing method, characterized in that: The method comprises: Obtaining video feature information corresponding to the first video; Obtaining environmental feature information corresponding to shooting environment information when the first video was shot; Fusing the video feature information with the environment feature information to obtain fused feature information; At least one first stylized template matching the fused feature information is determined.
2. The method according to claim 1, characterized in that The obtaining of video feature information corresponding to the first video includes: Obtaining image feature information corresponding to each video frame of the first video; Acquire audio feature information corresponding to the first video; The video feature information includes the image feature information and the audio feature information corresponding to each video frame.
3. The method according to claim 1, characterized in that The shooting environment information includes at least one of the following: gyroscope data collected by the electronic device when shooting the first video, location information of the electronic device when shooting the first video, and weather information of the location of the electronic device when shooting the first video.
4. The method according to any one of claims 1 to 3, characterized in that: The determining of at least one first stylized template matching the fused feature information includes: Matching the fused feature information with the template feature information corresponding to each stylized template stored in the template library; The stylized template corresponding to the template feature information matching the fused feature information is determined as the first stylized template.
5. The method according to any one of claims 1 to 3, characterized in that: The determining of at least one first stylized template matching the fused feature information includes: Acquire user preference information, where the user preference information is used to indicate the template type of the stylized template that the user historically prefers; Based on the user preference information, adjusting the fused feature information to obtain the adjusted fused feature information; Determine the at least one first stylized template that matches the adjusted fusion feature information.
6. The method according to claim 1, characterized in that After determining at least one first stylized template matching the fused feature information, the method further includes: Processing the first video using a second stylized template in the at least one first stylized template to obtain a second video; Based on the fusion feature information, the animation effects of the video elements in the second video are adjusted to obtain a third video.
7. A video processing device, characterized in that: The device comprises: an acquisition module and a processing module; The acquisition module is used to acquire video feature information corresponding to the first video; and acquire environmental feature information corresponding to shooting environment information when the first video is shot; The processing module is used to fuse the video feature information and the environmental feature information acquired by the acquisition module to obtain fused feature information; and determine at least one first stylized template that matches the fused feature information.
8. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the video processing method according to any one of claims 1 to 6 are implemented.
9. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the video processing method according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The computer program product is stored in a storage medium, and the program product is executed by at least one processor to implement the steps of the video processing method according to any one of claims 1 to 6.