Video shooting method, device and electronic equipment

By acquiring inertial measurement unit data in real time to determine the user's state, and by using different power consumption modes to collect and cache video frames, the problem of short battery life for high-quality shooting of electronic devices is solved, and a shooting experience with low power consumption pre-recording and instant response is achieved.

CN122138042APending Publication Date: 2026-06-02VIVO MOBILE COMM CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2026-03-04
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Electronic devices have short battery life in high-definition shooting mode, and pre-recording technology leads to high power consumption, making it difficult to balance high image quality, long battery life and instant response.

Method used

By acquiring the detection data of the inertial measurement unit in real time, user status information is determined, video frame sequences are collected using different power consumption modes, and low-power pre-recording is achieved by combining the buffer function.

Benefits of technology

While ensuring high image quality, it saves power consumption of electronic devices, solves the problem of capture delay, and provides an instant-response shooting experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122138042A_ABST
    Figure CN122138042A_ABST
Patent Text Reader

Abstract

This application discloses a video shooting method, apparatus, and electronic device, belonging to the field of electronic device technology. The method includes: upon receiving a user's command to initiate the shooting function of the electronic device, acquiring detection data from the inertial measurement unit of the electronic device; determining the user's target state information based on the detection data; controlling the camera module of the electronic device to shoot in a corresponding target power consumption mode based on the target state information, and buffering the obtained video frame sequence; upon receiving a user's command to terminate shooting, generating a target video based on the buffered video frame sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic equipment technology, specifically relating to a video shooting method, apparatus, and electronic equipment. Background Technology

[0002] With the rapid development of shooting technology, the shooting function of electronic devices has become their core application. Users expect to obtain a shooting experience with high image quality, long battery life and instant response by using the shooting function of electronic devices in scenarios such as recording life and assisting work.

[0003] Currently, electronic devices typically use a fixed image quality shooting mode. While the captured images or videos can guarantee a certain image quality, continuous high power consumption leads to short battery life and severe overheating. In order to solve the problem of capture delay, electronic devices can pre-record the shooting scene in advance, so as to capture the fleeting moment that the user needs. However, the high power consumption of the pre-recording technology of electronic devices also leads to short battery life.

[0004] Therefore, the shooting function of electronic devices in related technologies cannot simultaneously achieve high image quality, long battery life, and instant response, which restricts the user experience of the shooting function of electronic devices. Therefore, a new solution is urgently needed to resolve this contradiction. Summary of the Invention

[0005] The purpose of this application is to provide a video shooting method, apparatus, and electronic device that can save power consumption of electronic devices and respond to shooting in real time while ensuring high image quality, thus solving the problem of capture delay.

[0006] In a first aspect, embodiments of this application provide a video shooting method, the method comprising: Upon receiving a user's command to initiate the shooting function of the electronic device, the detection data of the inertial measurement unit of the AI ​​device is acquired; Based on the detection data, the user's target status information is determined; Based on the target state information, the camera module of the electronic device is controlled to shoot in the corresponding target power consumption mode, and the obtained video frame sequence is cached. Upon receiving a user's command to stop shooting, the target video is generated based on the cached video frame sequence.

[0007] Secondly, embodiments of this application provide a video recording device, the device comprising: The acquisition module is used to acquire the detection data of the inertial measurement unit of the electronic device when it receives a user's instruction to start the shooting function of the electronic device. The determination module is used to determine the user's target status information based on the detection data; The control module is used to control the camera module of the electronic device to shoot in the corresponding target power consumption mode according to the target state information, and to cache the obtained video frame sequence; The generation module is used to generate the target video based on the cached video frame sequence when a user's shooting termination command is received.

[0008] Thirdly, embodiments of this application provide an electronic device, which includes a processor and a memory, wherein the memory stores programs or instructions that can run on the processor, and the programs or instructions, when executed by the processor, implement the method as described in the first aspect.

[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the method described in the first aspect.

[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0012] In this embodiment, the user's target state information is determined by real-time acquisition of detection data from the inertial measurement unit of the electronic device. Based on this target state information, the corresponding target power consumption mode is selected to acquire video frame sequences. This allows for continuous monitoring of user state changes, employing different power consumption modes to acquire video frame sequences according to different user states. This ensures high image quality without requiring the electronic device to operate in a single high-power mode, thus saving power. Furthermore, by caching the acquired video frame sequences, combining the caching function required for pre-recording with dynamic power consumption control, a practical pre-recording function is achieved at a low power consumption level that the electronic device can tolerate. This ensures that even if the user does not press the shutter in time, the system has already cached the image before the crucial moment, solving the problem of capture delay and providing an instant-response shooting experience. Attached Figure Description

[0013] Figure 1 This is a flowchart illustrating a video capture method provided in some embodiments of this application; Figure 2These are schematic diagrams showing the shooting preview interface provided in some embodiments of this application; Figure 3 This is a flowchart illustrating a video capture method provided in some embodiments of this application; Figure 4 These are schematic diagrams illustrating the structure of a video recording device according to some embodiments of this application; Figure 5 These are schematic diagrams illustrating the structure of an electronic device according to some embodiments of this application; Figure 6 These are schematic diagrams illustrating the hardware structure of an electronic device according to some embodiments of this application. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0015] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and are not limited in number; for example, a first object can be one or N objects. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0016] The terminology used in the embodiments of this invention will be explained below.

[0017] Artificial Intelligence (AI) is the study of using computers to simulate certain human thought processes and intelligent behaviors, such as learning, reasoning, thinking, and planning. It mainly includes the principles of how computers achieve intelligence, the creation of computers that resemble the intelligence of the human brain, and enabling computers to achieve higher-level applications.

[0018] An Inertial Measurement Unit (IMU) is a device that measures an object's three-axis attitude angles (or angular rates) and acceleration. Typically, an IMU contains three single-axis accelerometers and three single-axis gyroscopes. The accelerometers detect the object's acceleration signals along the three independent axes of the carrier's coordinate system, while the gyroscopes detect the carrier's angular velocity signals relative to the navigation coordinate system. By measuring the object's angular velocity and acceleration in three-dimensional space, the object's attitude can be calculated. The IMU's role in electronic devices mainly includes three aspects: first, electronic image stabilization, primarily compensating for lens shake using gyroscope data; second, attitude interaction control, where shaking the electronic device triggers commands (e.g., nodding or shaking the head triggers commands in AI devices like AI glasses); and third, energy-saving wake-up, where the camera can be turned off when the accelerometer detects the electronic device is stationary to reduce power consumption.

[0019] The aforementioned gyroscope can measure the angular velocities of electronic devices around the X, Y, and Z axes, and then perform integral calculations based on the angular velocities of the electronic devices around the X, Y, and Z axes to obtain the attitude angle changes, as shown in the following formula (1): In the above formula (1), This refers to the angular velocity of an electronic device around the X, Y, or Z axis. For example, when a user turns their head, the change in the Z-axis angular velocity triggers an adjustment in the screen orientation.

[0020] The formula for calculating the variance of angular velocity is shown in formula (2) below: In the above formula (2), For the first The angular velocity value of each sampling point, in units of , for The average angular velocity of each sampling point This represents the number of samples.

[0021] AI frame interpolation technology is a technique that uses artificial intelligence algorithms to intelligently generate new frames between the original frames of a video or game, aiming to improve the smoothness of the visuals and the overall experience. Its core principle is to analyze the pixel motion patterns between adjacent frames, "predict" and insert intermediate transition frames, thereby converting low frame rate content into high frame rate content, such as converting 24 frames per second (fps) content to 60 fps or higher. The core algorithm used in AI frame interpolation technology is optical flow: specifically, it generates intermediate transition frames by calculating the motion vector of each pixel between two adjacent frames; for example, converting 30fps to 60fps requires inserting one frame between every two frames.

[0022] Currently, AI frame interpolation technology is usually integrated into artificial intelligence models. AI frame interpolation is achieved based on artificial intelligence models. The commonly used artificial intelligence models that integrate AI frame interpolation technology use a Pyramid Warping Cost Network (PWC-Net) to predict pixel displacement and then perform AI frame interpolation. The specific process is as follows: Step 1: Construct an image pyramid (i.e., multi-scale feature extraction); Step 2: Calculate the optical flow field of adjacent layers; Step 3: Optimize the optical flow accuracy through a convolutional network. The specific interpolation formula is shown in formula (3) below: In the above formula (3), For the first The pixel matrix of the original frame image at any given time; For the first The pixel matrix of the original frame image at any given time; For the insertion frame position weight, , Used to control the insertion of frames and The proportion of positions between them The closer to 0, the closer the inserted frame. , The closer to 0, the closer the inserted frame. ; As an optical flow correction term, it compensates for inter-frame motion deviations based on the optical flow method, corrects the distortion problem of simple linear interpolation, and improves the motion coherence of interpolated frames; For generated The position is inserted into the pixel matrix of the frame image, which is the final output of the padded frame result.

[0023] PWC-Net is a deep learning network architecture for stereo matching. The core idea of ​​PWC-Net is to process images of different scales by constructing a pyramid structure, and to improve the accuracy and efficiency of stereo matching by combining image warping and cost aggregation techniques. Its core components include: **Pyramid Structure:** By constructing an image pyramid, features at different scales can be captured, which is very effective for handling objects and scenes of different sizes. **Image Warping Component:** In stereo matching, pixels from one view are mapped onto another view to find the best matching point. **Cost Aggregation Component:** By aggregating cost information from different scales and locations, disparity, i.e., depth information, can be estimated more accurately.

[0024] Extended Reality (XR) is a collective term encompassing Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR). It achieves fusion by combining virtual content with real-world scenes through hardware devices.

[0025] Controls are graphical elements that users can directly manipulate or perceive during interaction. They are used to receive user input, trigger functions, or display real-time information. They serve as an interaction bridge between users and device or application functions, enabling command transmission and status feedback through a visual format.

[0026] Interface: Refers to the graphical interactive layer seen by users through the screen of an electronic device. Also known as the "user interface (UI)," it is the medium through which applications or operating systems interact and exchange information with users, converting the internal form of information into a form acceptable to the user. The user interface is source code written in specific computer languages ​​such as Java and XML. This source code is parsed and rendered on the electronic device, ultimately presenting content that the user can recognize. The most common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be visible interface elements displayed on the screen of an electronic device, such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and web widgets.

[0027] An application (APP) is a computer program developed and running on an operating system to accomplish one or more specific tasks. Applications run in user mode, allowing interaction with the user and featuring a visual user interface.

[0028] Photo preview interface: This is the real-time view that the user sees on the device screen before taking a photo. It is a visual interactive area that is rendered in real time after the image data captured by the camera sensor is processed. It can be a GUI, which can display visible interface elements such as buttons, navigation bars, and widgets.

[0029] The technical solution of this application embodiment can be applied to scenarios where it is necessary to use an electronic device to shoot a video, and the power consumption of the electronic device must be ensured during the video shooting process. For example, a user visits a zoo and wants to record the process of a peacock displaying its tail feathers using their AI glasses. Therefore, the user starts recording before the peacock displays its tail feathers. Initially, the user is not directly in front of the peacock, so after starting recording, the user moves to be directly in front of the peacock. At the 8-second mark, the user is directly in front of the peacock and remains still. Two seconds later, the peacock begins to display its tail feathers. The peacock displays its tail feathers for a total of 5 seconds. After the peacock displays its tail feathers, the user clicks the end-of-recording control and obtains a video of the peacock displaying its tail feathers.

[0030] The video shooting method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0031] Figure 1 This is a flowchart illustrating a video shooting method provided in an embodiment of this application. The subject executing this video shooting method can be an electronic device, which can be, but is not limited to, a personal computer (PC), smartphone, tablet computer, personal digital assistant (PDA), or XR device, etc. Here, the XR device can be, but is not limited to, a VR device or an AR device. In the following embodiments, an XR device is used as an example for illustration; specifically, the XR device can be AI glasses.

[0032] like Figure 1 As shown, the video shooting method provided in this application embodiment may include steps 110-140.

[0033] Step 110: Upon receiving a user's instruction to initiate the shooting function of the electronic device, acquire the detection data of the inertial measurement unit of the electronic device.

[0034] The shooting start command can be a command to activate the shooting function. Specifically, it can be a command generated upon receiving a user's command to activate the shooting function, such as... Figure 2 As shown, users can generate a shooting start command by clicking the shooting control 25 in the camera application's shooting preview interface.

[0035] like Figure 2 As shown, the photo preview interface includes a toolbar 21, a preview window 22, camera mode options, album shortcut controls 24, a photo control 25, and a camera flip control 26.

[0036] Toolbar 21 can be used to display one or more functional controls, such as Figure 2 As shown, the one or more functional controls may include a flash control 211 and a settings control 213, etc. It is understood that the toolbar may also include more or fewer functional controls.

[0037] The preview window 22 can be used to display a preview image obtained after the image data is rendered in real time. The preview window can also be called a viewfinder.

[0038] Camera mode options can be user-selectable preset function modules to adapt to different shooting scenarios or operational needs. By adjusting the combination of hardware parameters and software algorithms, specific shooting effects or creative intentions can be achieved. For example... Figure 2 As shown, the camera mode options may include portrait mode option 231, video recording mode option 232, and photo mode option 233, etc.

[0039] The detection data from the inertial measurement unit (IMU) can be data detected by the IMU itself, such as the angular velocity of the electronic device. Specifically, the IMU's sampling frequency can be 32 kHz, meaning it can sample 100 times every 3.125 milliseconds, and then perform calculations on these 100 data points. The specific sampling frequency of the IMU can be set according to user needs and is not limited in this embodiment.

[0040] In some embodiments of this application, the attitude angle of the electronic device can be calculated according to the above formula (1) based on the obtained angular velocity of the electronic device, thereby determining the attitude of the electronic device. The angular velocity variance can also be calculated according to the above formula (2).

[0041] Step 120: Determine the user's target status information based on the detection data.

[0042] The target state information can be the user's state information, such as a static state or a moving state. The specific moving state can be further subdivided into micro-movement state, walking state, etc. The specific division of the user's state information can be customized according to the user's needs.

[0043] In some embodiments of this application, when the detection data includes the angular velocity of the electronic device, step 120 may specifically include: Calculate the variance of the angular velocity; The user's target state information is determined based on the relationship between variance and a preset variance threshold.

[0044] Wherein, the preset variance threshold can be a pre-set threshold for the variance of the angular velocity within the first time period, and this preset variance threshold can be, for example, 2. Or 5 The specific value of the preset variance threshold can be set by the user according to their needs, and is not limited in this embodiment.

[0045] In some embodiments of this application, since variance reflects the speed and direction of an object’s rotation, when calculating the variance of angular velocity, it is a statistical measure of the degree of fluctuation of the measured or instantaneous value of angular velocity around its average value. Therefore, the user’s posture can be reflected by the variance. Thus, the variance of the angular velocity of the electronic device can be calculated in real time according to the above formula (2), and then the user’s target state information can be determined according to the variance and the preset variance threshold.

[0046] In the embodiments of this application, variance calculation is a standard mathematical operation with low computational load and low resource consumption of the device processor. Thus, by calculating the variance of the angular velocity of the electronic device in real time, the user's target state information can be determined, achieving low-power state monitoring and helping to extend the device's battery life.

[0047] In some embodiments of this application, determining the user's target state information based on the relationship between variance and a preset variance threshold may specifically include: If the variance is greater than the first preset variance threshold and the variance remains greater than the first preset variance threshold for a first preset duration, the user's target state information is determined to be in motion. If the variance is less than the second preset variance threshold and the variance remains less than the second preset variance threshold for a second preset duration, the user's target state information is determined to be in a static state.

[0048] The first preset variance threshold and the second preset variance threshold can be two pre-set variance thresholds, specifically the first preset variance threshold being greater than or equal to the second preset variance threshold.

[0049] The first preset duration can be the duration during which the variance of the electronic device's angular velocity exceeds a first preset variance threshold. For example, the first preset duration could be 3 seconds.

[0050] The second preset duration can be the duration during which the variance of the electronic device's angular velocity is less than a second preset variance threshold. For example, the second preset duration could be 3 seconds.

[0051] It should be noted that the values ​​of the first preset duration and the second preset duration mentioned above can be the same or different. The specific values ​​of the first preset duration and the second preset duration can be set by the user according to their needs, and are not limited in this embodiment.

[0052] In some embodiments of this application, a larger angular velocity variance indicates a faster user speed and more unstable motion, i.e., closer to a motion state. Therefore, if the variance of the electronic device's angular velocity is greater than a first preset variance threshold, and this variance remains greater than the first preset variance threshold for a first preset duration, it indicates that the user's instantaneous action changes significantly, and thus the user's target state information can be determined as a motion state.

[0053] As in the example above, the first preset variance threshold is 5. Taking the first preset duration of 3 seconds as an example, when the user just presses the button... Figure 2 After the camera control is activated (25), the camera begins recording the peacock's display of its tail feathers. During the first 8 seconds, the user moves around, attempting to move directly in front of the peacock. During these 8 seconds, the variance of the angular velocity per second is consistently higher than 5. and higher than 5 The duration lasted for 8 seconds, so it can be determined that the user's target state information during these 8 seconds was a motion state.

[0054] Because a smaller angular velocity variance indicates a slower user speed and more stable motion, meaning it's closer to a stationary state. Therefore, if the variance of the electronic device's angular velocity is less than a second preset variance threshold, and this variance remains below the second preset variance threshold for a second preset duration, it indicates that the user's instantaneous movement amplitude is small, thus confirming that the user's target state information is a stationary state.

[0055] As in the example above, the second preset variance threshold is 3. Taking the second preset duration of 3 seconds as an example, just after the user presses the button... Figure 2 After the camera control is activated (25), the camera begins recording the peacock's display of its tail feathers. For the first 8 seconds, the user's target state is in motion. Then, the user moves to the front of the peacock and remains stationary. After 2 seconds, the peacock displays its tail feathers. Throughout this process, the user remains stationary. The peacock displays its tail feathers for 5 seconds. Therefore, within the 7 seconds between the 2 seconds after the user moves to the front of the peacock and the 5 seconds of the tail feathers display, the variance of the angular velocity per second is less than 3. And below 3 The duration lasted for 7 seconds, so it can be determined that the user's target state information was static during these 7 seconds.

[0056] It should be noted that when a user's command to initiate the shooting function of the electronic device is detected, a timer can be activated. Specifically, two timers, Timer A and Timer B, can be activated. Timer A is responsible for counting the duration for which the variance is greater than a first preset variance threshold, and Timer B is responsible for counting the duration for which the variance is less than a second preset variance threshold. When the timer A reaches the time corresponding to the first preset duration, the user's target state information can be determined to be in motion. When the timer B reaches the time corresponding to the second preset duration, the user's target state information can be determined to be in a stationary state.

[0057] In the embodiments of this application, the stability and anti-interference of the user's target state information are improved by using both variance and duration to determine the target state information, effectively preventing misjudgment caused by momentary jitter or short pauses, and improving the accuracy of the target state information determination.

[0058] Step 130: Based on the target status information, control the camera module of the electronic device to shoot in the corresponding target power consumption mode, and cache the obtained video frame sequence.

[0059] The target power consumption mode can be a power consumption mode corresponding to the target state information. Different target state information corresponds to different target power consumption modes.

[0060] In some embodiments of this application, after determining the user's target state information, the camera module of the electronic device can be controlled to shoot in the corresponding target power consumption mode to obtain a video frame sequence, and then the obtained video frame sequence can be cached.

[0061] In some embodiments of this application, step 130 may specifically include: When the target state information is in motion, the camera module of the control electronic device takes pictures in the first power consumption mode, obtains the first video frame sequence, and stores the first video frame sequence in the first buffer. When the target state information is static, the camera module of the control electronic device takes pictures in the second power mode to obtain the second video frame sequence and stores the second video frame sequence in the second buffer.

[0062] The first power consumption mode can be the power consumption mode used by the camera to capture video frames when the target state information is in motion.

[0063] The second power consumption mode can be the power consumption mode used by the camera to capture video frames when the target state information is in a static state.

[0064] The power consumption of the first power mode is lower than that of the second power mode. In other words, the electronic device uses less power in the first power mode than in the second power mode; that is, the second power mode consumes more power than the first. For example, the first power mode could be 480P / 15fps, while the second power mode could be 1080P / 15fps.

[0065] The first video frame sequence may be a sequence of video frames captured by the camera module in the first power consumption mode.

[0066] The second video frame sequence can be a sequence of video frames captured by the camera module in the second power consumption mode.

[0067] The image quality of the first video frame sequence is lower than that of the second video frame sequence. In other words, the frame rate of the first video frame sequence is not as high as that of the second video frame sequence. For example, the frame rate of the first video frame sequence is 24 frames per second, while the frame rate of the second video frame sequence is 60 frames per second.

[0068] The first buffer can be a buffer used to cache the first video frame sequence. The second buffer can be a buffer used to cache the second video frame sequence. Here, the first buffer and the second buffer are different buffer areas in the buffer space of the electronic device.

[0069] In some embodiments of this application, since the user is in motion, the captured image is not necessarily the scene the user specifically wants to capture. Because a user typically wants to capture a specific scene while stationary, the image captured while the user is in motion does not need to be of high quality; it only needs to record the scene and provide the desired video. Therefore, when the target state information indicates motion, the camera module of the electronic device can be controlled to capture images in a first power consumption mode, thus obtaining a low-quality first video frame sequence, which is then stored in a first buffer.

[0070] As in the example above, when the user presses... Figure 2 After the user moves to the front of the peacock, the video frame sequence within 8 seconds can be collected using the first power mode. The video frame sequence within these 8 seconds is then cached in the first buffer.

[0071] It should be noted that when caching the first video frame sequence, the N video frames closest to the current time are retained. Here, N is a positive integer, and the value of N can be set by the user according to their needs. It is not limited in this embodiment.

[0072] When the target state information is static, the camera module of the controllable electronic device can shoot in the second power mode, thereby obtaining the second video frame sequence, which can then be stored in the second buffer.

[0073] As in the example above, when the user presses... Figure 2 After the user moves to the front of the peacock, the video frame sequence of 2 seconds after the user moves to the front of the peacock and 5 seconds after the peacock spreads its tail can be collected using the second power mode to obtain the video frame sequence within these 7 seconds, and then the video frame sequence within these 7 seconds is cached in the second buffer.

[0074] In the embodiments of this application, when the target state information is in motion, a low-power shooting mode is used to capture the image, resulting in a low-quality first video frame sequence, which significantly reduces the power consumption of the camera module. Furthermore, during motion, the image changes rapidly, and while the information increment from a high frame rate is limited, the data volume is enormous. Therefore, a low-power shooting mode can be used to capture the image, saving storage space while still being sufficient to record the motion trajectory and key moments. When the target state information is stationary, the image is stable, and a high frame rate can capture more details. In this case, storage space and computing power are used effectively, allowing different power-consuming shooting modes to be used based on the user's different states, and thus enabling caching and efficient resource utilization.

[0075] In some embodiments of this application, since the obtained first video frame sequence has low image quality, the synthesized video will have an impact on the viewing experience. Therefore, after obtaining the first video frame sequence, the method described above may further include: The first video frame sequence is interpolated to obtain the third video frame sequence; The step of storing the first video frame sequence into the first buffer includes: Store the third video frame sequence into the first buffer.

[0076] The third video frame sequence can be a video frame sequence obtained by interpolating the first video frame sequence. Specifically, it can be obtained by interpolating the first video frame sequence using AI frame interpolation technology.

[0077] Since the third video frame sequence is obtained by interpolating the first video frame sequence, its image quality is superior to that of the first video frame sequence. Specifically, the image quality of the third video frame sequence can be close to that of the second video frame sequence. Thus, when fused with the second video frame sequence to obtain a video, the difference in image quality between the two will not be too large; that is, the difference in image quality between the third and second video frame sequences is less than a first threshold. This first threshold can be a pre-set threshold representing the difference in image quality between the third and second video frame sequences. For example, the first threshold could be set to 1 frame / second. The specific value of the first threshold can be set according to user needs and is not limited in this embodiment.

[0078] In some embodiments of this application, since the image quality of the first video frame sequence is poor, the first video frame sequence can be interpolated to obtain a third video frame sequence with better image quality. Then, when storing the video frame sequence, the third video frame sequence is stored in the first buffer. In this way, when the video frame sequence is obtained from the first buffer for video synthesis, the third video frame sequence is used directly.

[0079] In the embodiments of this application, a high-quality third video frame sequence is obtained by interpolating the low-quality first video frame sequence captured in low-power mode. This is because low frame rate videos may exhibit stuttering or jumps in motion scenes. Frame interpolation, by generating intermediate frames, makes the motion trajectory appear more continuous and smooth. Thus, without increasing the original shooting power consumption and storage costs, the visual smoothness and viewing experience of the video are significantly improved. Furthermore, using a low frame rate during shooting directly saves power and storage space during the shooting phase. Frame interpolation, as a post-processing step, shifts high-power tasks from critical shooting moments, optimizing the overall power consumption distribution of the device.

[0080] Step 140: Upon receiving the user's instruction to stop shooting, generate the target video based on the cached video frame sequence.

[0081] The shooting termination command can be a command to end the shooting function. Specifically, it can be a command generated upon receiving a user's command to end the shooting function, for example, it could be... Figure 2 As shown, when a user clicks the camera control 25 for the first time to generate a shooting start command and begin shooting, clicking the camera control 25 again will generate a shooting end command.

[0082] The target video can be a video generated based on a cached sequence of video frames, such as the video obtained from the moment the user clicks the shooting control to the moment the user finishes recording the peacock spreading its tail feathers.

[0083] In some embodiments of this application, generating the target video based on the cached video frame sequence includes: The cached third video frame sequence and the second video frame sequence are concatenated to generate the target video.

[0084] In some embodiments of this application, after obtaining the third video frame sequence and the second video frame sequence, the third video frame sequence and the second video frame sequence can be spliced ​​together to obtain a target video with better image quality.

[0085] In the embodiments of this application, a target video with better image quality can be obtained by stitching together the third video frame sequence and the second video frame sequence. Thus, the system continuously "buffers" footage from the past few seconds or minutes through low-power pre-recording. When a static state is detected and high-quality shooting is triggered, the pre-recorded low-power video before the trigger can be seamlessly stitched together with the high-quality video after the trigger. In this way, the target video completely contains the entire process before and during the event, avoiding missing crucial beginnings due to system startup delays. This application's solution only briefly activates high-power mode when a "crucial static scene" is confirmed. Ultimately, through stitching, a complete video with a "smooth beginning and high-definition ending" is obtained, achieving high-quality recording of critical moments with minimal continuous resource consumption, fundamentally resolving the contradiction between continuous high-quality recording and limited device resources.

[0086] It should be noted that the user's target state information in the above embodiments is only illustrated using motion state and stationary state as examples. However, those skilled in the art should know that the user's target state information is not limited to motion state and stationary state, but may also include violent motion state, walking state, micro-motion state and stationary state. When the user's target state information is violent motion state, walking state or micro-motion state, the video frame sequence collected in the corresponding power consumption mode under the violent motion state, walking state or micro-motion state can be interpolated according to the interpolation method of motion state in the above embodiments, so that the image quality of the video frame sequence under the violent motion state, walking state or micro-motion state after interpolation is not much different from that under the stationary state.

[0087] Additionally, it should be noted that the data collected in the above embodiments only includes the detection data of the inertial measurement unit of the electronic device. However, those skilled in the art should know that the collected data is not limited to the detection data of the inertial measurement unit of the electronic device. For example, it may also include barometer data to detect changes in altitude in real time. In this way, the user's posture changes can be accurately determined based on the barometer data and the detection data of the inertial measurement unit of the electronic device, specifically whether it is mountain climbing or other sports.

[0088] To better understand the video shooting method provided in the embodiments of this application, the video shooting method of the embodiments of this application will be described below with specific scenarios.

[0089] Figure 3 This is a flowchart illustrating a video shooting method provided in an embodiment of this application, as shown below. Figure 3 As shown, the video shooting method provided in this application embodiment may include steps 31-44.

[0090] Step 31: Upon receiving the user's instruction to start the shooting function of the electronic device, acquire the detection data of the inertial measurement unit of the electronic device.

[0091] Step 31 is the same as step 110 in the above embodiment, and will not be described again here.

[0092] Step 32: Determine whether the user's shooting stop command has been received. If not, proceed to step 33; if yes, proceed to step 34.

[0093] Step 33: Assemble the captured video frame sequence in chronological order to obtain the video.

[0094] In steps 32-33, if a user's shooting termination command is received, it means that the user has only shot a short video segment. In this case, the captured video frame sequence can be directly spliced ​​together in chronological order to obtain the final video without any other operations on the captured video frame sequence.

[0095] Step 34: Calculate the variance of the angular velocity.

[0096] Step 35: Determine whether the variance is greater than the first preset variance threshold. If yes, proceed to step 36; otherwise, proceed to step 40.

[0097] Step 36: Determine whether the variance is greater than the first preset variance threshold for a first preset duration. If yes, proceed to step 37; otherwise, return to step 32.

[0098] Step 37: Determine that the user's target state information is in motion, and control the camera module of the electronic device to shoot in the first power consumption mode to obtain the first video frame sequence.

[0099] Steps 34-37 above are consistent with those in the above embodiment, namely calculating the variance of angular velocity; if the variance is greater than the first preset variance threshold and the variance is greater than the first preset variance threshold for a first preset duration, the user's target state information is determined to be in motion; if the target state information is in motion, the camera module of the electronic device is controlled to shoot in the first power consumption mode to obtain the first video frame sequence, which will not be described again here.

[0100] Step 38: Perform frame interpolation on the first video frame sequence to obtain the third video frame sequence.

[0101] Step 38 is the same as the process of interpolating the first video frame sequence to obtain the third video frame sequence in the above embodiment, and will not be described again here.

[0102] Step 39: Store the third video frame sequence into the first buffer.

[0103] Step 39 is the same as the process of storing the third video frame sequence into the first buffer in the above embodiment, and will not be described again here.

[0104] Step 40: Determine whether the variance is less than the second preset variance threshold. If yes, proceed to step 41; otherwise, return to step 32.

[0105] Step 41: Determine whether the variance is less than the second preset variance threshold for a second preset duration. If yes, proceed to step 42; otherwise, return to step 32.

[0106] Step 42: Determine that the user's target state information is a static state, and control the camera module of the electronic device to shoot in the second power consumption mode to obtain the second video frame sequence.

[0107] Steps 34 and 40-42 above are consistent with those in the above embodiment, namely, calculating the variance of angular velocity; if the variance is less than the second preset variance threshold and the variance remains less than the second preset variance threshold for a second preset duration, the user's target state information is determined to be stationary; if the target state information is stationary, the camera module of the electronic device is controlled to shoot in the second power consumption mode to obtain the second video frame sequence, which will not be described in detail here.

[0108] Step 43: Store the second video frame sequence into the second buffer.

[0109] Step 43 above is the same as the process of storing the second video frame sequence into the second buffer in the above embodiment, and will not be repeated here.

[0110] Step 44: Concatenate the cached third video frame sequence and the second video frame sequence to generate the target video.

[0111] Step 44 is the same as the process in the above embodiment of splicing the cached third video frame sequence and the second video frame sequence to generate the target video, and will not be described again here.

[0112] The video shooting method provided in this application can be executed by a video shooting device. This application uses a video shooting device as an example to illustrate the video shooting device provided in this application.

[0113] Figure 4 This is a schematic diagram illustrating the structure of a video recording device according to an exemplary embodiment. Figure 4 As shown, the video recording device 400 may include: The acquisition module 410 is used to acquire the detection data of the inertial measurement unit of the electronic device when it receives a user's shooting start command for the shooting function of the electronic device. The determination module 420 is used to determine the user's target status information based on the detection data; The control module 430 is used to control the camera module of the electronic device to shoot in the corresponding target power consumption mode according to the target state information, and to cache the obtained video frame sequence; The generation module 440 is used to generate a target video based on the cached video frame sequence when a user's shooting termination instruction is received.

[0114] In this embodiment, the user's target state information is determined by real-time acquisition of detection data from the inertial measurement unit of the electronic device. Based on this target state information, the corresponding target power consumption mode is selected to acquire video frame sequences. This allows for continuous monitoring of user state changes, employing different power consumption modes to acquire video frame sequences according to different user states. This ensures high image quality without requiring the electronic device to operate in a single high-power mode, thus saving power. Furthermore, by caching the acquired video frame sequences, combining the caching function required for pre-recording with dynamic power consumption control, a practical pre-recording function is achieved at a low power consumption level that the electronic device can tolerate. This ensures that even if the user does not press the shutter in time, the system has already cached the image before the crucial moment, solving the problem of capture delay and providing an instant-response shooting experience.

[0115] In some embodiments of this application, the detection data includes the angular velocity of the electronic device; the determining module is specifically used for: Calculate the variance of the angular velocity; Based on the relationship between the variance and the preset variance threshold, the user's target state information is determined.

[0116] In some embodiments of this application, the determining module is specifically used for: If the variance is greater than a first preset variance threshold and the variance remains greater than the first preset variance threshold for a first preset duration, the user's target state information is determined to be a motion state. If the variance is less than a second preset variance threshold and the variance remains less than the second preset variance threshold for a second preset duration, the user's target state information is determined to be a static state. Wherein, the first preset variance threshold is greater than or equal to the second preset variance threshold.

[0117] In some embodiments of this application, the control module is specifically used for: When the target state information is in motion, the camera module of the electronic device is controlled to shoot in a first power consumption mode to obtain a first video frame sequence, and the first video frame sequence is stored in a first buffer. When the target state information is in a static state, the camera module of the electronic device is controlled to shoot in a second power consumption mode to obtain a second video frame sequence, and the second video frame sequence is stored in a second buffer. Wherein, the power consumption of the first power consumption mode is lower than that of the second power consumption mode, and the image quality of the first video frame sequence is lower than that of the second video frame sequence.

[0118] In some embodiments of this application, the apparatus further includes: The frame interpolation module is used to perform frame interpolation processing on the first video frame sequence after obtaining the first video frame sequence to obtain a third video frame sequence, wherein the difference in image quality between the third video frame sequence and the second video frame sequence is less than a first threshold. The control module is specifically used for: The third video frame sequence is stored in the first buffer.

[0119] In some embodiments of this application, the generation module is specifically used for: The cached third video frame sequence and the second video frame sequence are concatenated to generate the target video.

[0120] The video recording device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0121] The video recording device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0122] The video shooting device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0123] Optionally, such as Figure 5 As shown, this application embodiment also provides an electronic device 500, including a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. When the program or instructions are executed by the processor 501, they implement the various steps of the above-described video shooting method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0124] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0125] Figure 6 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0126] The electronic device 600 includes, but is not limited to, components such as: radio frequency unit 601, network module 602, audio output unit 603, input unit 604, sensor 605, display unit 606, user input unit 607, interface unit 608, memory 609, and processor 610.

[0127] Those skilled in the art will understand that the electronic device 600 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 610 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0128] The processor 610 is configured to, upon receiving a user's instruction to initiate the shooting function of the electronic device, acquire detection data from the inertial measurement unit of the AI ​​device; determine the user's target state information based on the detection data; control the camera module of the electronic device to shoot in a corresponding target power consumption mode based on the target state information, and cache the obtained video frame sequence; and, upon receiving a user's instruction to terminate shooting, generate a target video based on the cached video frame sequence.

[0129] In this way, by acquiring real-time detection data from the inertial measurement unit of the electronic device, the user's target state information is determined. Based on this target state information, the corresponding target power consumption mode is selected to capture video frame sequences. This allows for continuous monitoring of user state changes, employing different power consumption modes to capture video frame sequences according to different user states. This ensures high image quality without requiring the electronic device to operate in a single high-power mode, thus saving power. Furthermore, by caching the acquired video frame sequences, combining the caching function required for pre-recording with dynamic power control, a practical pre-recording function is achieved at a low power consumption level that the electronic device can tolerate. This ensures that even if the user does not press the shutter in time, the system has already cached the image before the crucial moment, solving the problem of capture delay and providing an instant-response shooting experience.

[0130] Optionally, the detection data includes the angular velocity of the electronic device; the processor 610 is further configured to calculate the variance of the angular velocity; and determine the user's target state information based on the relationship between the variance and a preset variance threshold.

[0131] Optionally, the processor 610 is further configured to determine that the user's target state information is in motion when the variance is greater than a first preset variance threshold and the variance is greater than the first preset variance threshold for a first preset duration; and to determine that the user's target state information is in a stationary state when the variance is less than a second preset variance threshold and the variance is less than the second preset variance threshold for a second preset duration; wherein the first preset variance threshold is greater than or equal to the second preset variance threshold.

[0132] Optionally, the processor 610 is further configured to, when the target state information is in a moving state, control the camera module of the electronic device to shoot in a first power consumption mode to obtain a first video frame sequence and store the first video frame sequence in a first buffer; and when the target state information is in a stationary state, control the camera module of the electronic device to shoot in a second power consumption mode to obtain a second video frame sequence and store the second video frame sequence in a second buffer; wherein the power consumption corresponding to the first power consumption mode is lower than the power consumption corresponding to the second power consumption mode, and the image quality of the first video frame sequence is lower than the image quality of the second video frame sequence.

[0133] Optionally, the processor 610 is further configured to perform frame interpolation processing on the first video frame sequence to obtain a third video frame sequence, wherein the difference in image quality between the third video frame sequence and the second video frame sequence is less than a first threshold; and to store the third video frame sequence in a first buffer.

[0134] Optionally, the processor 610 is further configured to concatenate the cached third video frame sequence and the second video frame sequence to generate a target video.

[0135] It should be understood that, in this embodiment, the input unit 604 may include a graphics processing unit (GPU) 6041 and a microphone 6042. The GPU 6041 processes image data of still images or videos obtained by an image capture device (such as a color camera) in video capture mode or image capture mode. The display unit 606 may include a display panel 6061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 607 includes at least one of a touch panel 6071 and other input devices 6072. The touch panel 6071 is also called a touch screen. The touch panel 6071 may include a touch detection device and a touch controller. Other input devices 6072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0136] The memory 609 can be used to store software programs and various data. The memory 609 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 609 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 609 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0137] Processor 610 may include one or more processing units; optionally, processor 610 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 610.

[0138] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described video shooting method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0139] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0140] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described video shooting method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0141] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0142] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the video shooting method embodiment described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0143] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0144] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0145] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A video shooting method, characterized in that, The method includes: Upon receiving a user's command to initiate the shooting function of the electronic device, the system acquires the detection data from the inertial measurement unit of the electronic device. Based on the detection data, the user's target status information is determined; Based on the target state information, the camera module of the electronic device is controlled to shoot in the corresponding target power consumption mode, and the obtained video frame sequence is cached. Upon receiving a user's command to stop shooting, the target video is generated based on the cached video frame sequence.

2. The method according to claim 1, characterized in that, The detection data includes the angular velocity of the electronic device; The step of determining the user's target state information based on the detection data includes: Calculate the variance of the angular velocity; Based on the relationship between the variance and the preset variance threshold, the user's target state information is determined.

3. The method according to claim 2, characterized in that, The step of determining the user's target state information based on the relationship between the variance and a preset variance threshold includes: If the variance is greater than a first preset variance threshold and the variance remains greater than the first preset variance threshold for a first preset duration, the user's target state information is determined to be a motion state. If the variance is less than a second preset variance threshold and the variance remains less than the second preset variance threshold for a second preset duration, the user's target state information is determined to be a static state. Wherein, the first preset variance threshold is greater than or equal to the second preset variance threshold.

4. The method according to claim 1, characterized in that, The step of controlling the camera module of the electronic device to shoot in a corresponding target power consumption mode according to the target state information, and caching the obtained video frame sequence, includes: When the target state information is in motion, the camera module of the electronic device is controlled to shoot in a first power consumption mode to obtain a first video frame sequence, and the first video frame sequence is stored in a first buffer. When the target state information is in a static state, the camera module of the electronic device is controlled to shoot in a second power consumption mode to obtain a second video frame sequence, and the second video frame sequence is stored in a second buffer. Wherein, the power consumption of the first power consumption mode is lower than that of the second power consumption mode, and the image quality of the first video frame sequence is lower than that of the second video frame sequence.

5. The method according to claim 4, characterized in that, After obtaining the first video frame sequence, the method further includes: The first video frame sequence is subjected to frame interpolation to obtain a third video frame sequence, wherein the difference in image quality between the third video frame sequence and the second video frame sequence is less than a first threshold. The step of storing the first video frame sequence into the first buffer includes: The third video frame sequence is stored in the first buffer.

6. The method according to claim 5, characterized in that, The step of generating the target video based on the cached video frame sequence includes: The cached third video frame sequence and the second video frame sequence are concatenated to generate the target video.

7. A video recording device, characterized in that, The device includes: The acquisition module is used to acquire the detection data of the inertial measurement unit of the electronic device when it receives a user's instruction to start the shooting function of the electronic device. The determination module is used to determine the user's target status information based on the detection data; The control module is used to control the camera module of the electronic device to shoot in the corresponding target power consumption mode according to the target state information, and to cache the obtained video frame sequence; The generation module is used to generate the target video based on the cached video frame sequence when a user's shooting termination command is received.

8. The apparatus according to claim 7, characterized in that, The detection data includes the angular velocity of the electronic device; The determining module is specifically used for: Calculate the variance of the angular velocity; Based on the relationship between the variance and the preset variance threshold, the user's target state information is determined.

9. The apparatus according to claim 8, characterized in that, The determining module is specifically used for: If the variance is greater than a first preset variance threshold and the variance remains greater than the first preset variance threshold for a first preset duration, the user's target state information is determined to be a motion state. If the variance is less than a second preset variance threshold and the variance remains less than the second preset variance threshold for a second preset duration, the user's target state information is determined to be a static state. Wherein, the first preset variance threshold is greater than or equal to the second preset variance threshold.

10. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the video capture method as described in any one of claims 1-6.