A video processing method and terminal
By copying one audio data stream into N parallel encoding streams during the recording process, the problem of not being able to record the main character video and the original video simultaneously in automatic focus tracking mode was solved, achieving synchronous recording of the main character video and the original video while preserving information about the surrounding environment of the main character.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2022-07-30
- Publication Date
- 2026-04-17
AI Technical Summary
In autofocus shooting mode, the device can only record the video of the main subject and cannot record the original video at the same time, resulting in the loss of environmental information.
By copying one audio data stream into N audio data streams during the recording process and encoding them in parallel to generate N videos, the simultaneous recording of the main video and the original video can be achieved.
It enables the simultaneous generation of the main character's video and the original video during a single recording process, while preserving the environmental information surrounding the main character.
Smart Images

Figure 1
Abstract
Description
[0001] This application is a divisional application. The original application is entitled "A Video Processing Method and Terminal". The original application number is 202210912445.1 and the original application date is July 30, 2022. The entire contents of the original application are incorporated herein by reference. Technical Field
[0002] This application relates to the field of terminals, and more particularly to a video processing method and a terminal. Background Technology
[0003] With the development of smart terminals, mobile phones that support video recording can now achieve automatic focus tracking shooting modes. During video recording, the terminal can receive a user-selected subject. Then, the terminal can continuously track the focus of this subject throughout the subsequent recording process, resulting in a video that includes the selected subject in the image. This type of video can also be called a focus-tracking video or a close-up video, etc.
[0004] However, in autofocus shooting mode, the device only acquires the main subject's video; it cannot record both the main subject's video and the original video simultaneously during a single recording session. The original video consists of the original images captured by the camera. Summary of the Invention
[0005] This application provides a video processing method and terminal. During video recording, one channel of audio data is copied into N channels of audio data. Based on the N channels of audio data, they can be encoded simultaneously to generate N videos in one recording process.
[0006] In a first aspect, this application provides an audio processing method applied to a terminal, the terminal including a camera. The method includes: displaying a recording interface, the recording interface including a first control, a first preview window, and a second preview window; the first preview window being used to display an image captured by the camera; at a first moment, the first preview window displays a first image; the first image is captured by the camera at the first moment; when a first object is detected at a first position in the first image, the terminal displays a second image in the second preview window; the second image is generated based on the first image, and the second image includes the first object; at a second moment, the first preview window displays a third image; the third image is captured by the camera at the second moment; when the first object is detected at a second position in the third image, the terminal displays a third image in the second preview window. The preview window displays a fourth image; this fourth image is generated based on the third image and includes the first object; first audio data is acquired, including a first audio signal and information about the first audio signal; the first audio signal is an audio signal generated based on an audio signal acquired at the first moment; a first operation on the first control is detected, and in response to the first operation, video recording is stopped, and a first video recorded based on the image displayed in the first preview window and a second video recorded based on the image displayed in the second preview window are saved, wherein the first video includes the first image and first target audio data; the second video includes the second image and second target audio data; both the first target audio data and the second target audio data are obtained based on the first audio data.
[0007] In the above embodiments, the terminal can capture two video streams during a single recording session. One video stream can be the original video (first video) mentioned in the embodiments, and the other video stream can be the main subject video (second video) mentioned in the embodiments. The main subject video always includes a focus-tracking object (i.e., the first object). During video recording, the terminal's microphone can collect audio signals, and the terminal can perform real-time processing based on these collected audio signals to obtain the corresponding audio signals from the main subject video and the corresponding audio signals from the original video. It should be understood that because the video recording process is real-time, the audio processing should also be real-time. This ensures that the terminal can obtain two audio signals in a single recording session.
[0008] In conjunction with the first aspect, in some embodiments, before displaying the recording interface, the method further includes: displaying a preview interface, the preview interface including a second control; detecting a second operation on the second control, and in response to the second operation, the camera application entering a first mode; when the first preview window displays a first image, detecting a third operation on the first object; in response to the third operation; displaying a second preview window.
[0009] In the above embodiments, the second control is the protagonist mode control in the embodiments. The protagonist mode control is the entry point provided by this application embodiment for the function of shooting two video streams. It is independent of other shooting modes and is easy to manage.
[0010] In conjunction with the first aspect, after acquiring the first audio data and before stopping video recording, the method further includes: the terminal determining the number of stream configurations N; copying the first audio data into N identical second audio data streams based on the number of stream configurations N, wherein each second audio data stream includes the first audio signal and some information of the first audio signal; obtaining the first target audio data based on one second audio data stream; and obtaining the second target audio data based on another second audio data stream.
[0011] In the above embodiments, the first audio data includes the audio signals related to the generated protagonist video and the original video. However, there is only one first audio data stream (related to hardware configuration), and it has not yet been encoded (the audio needs to be encoded to obtain an audio stream when generating the video, which is then mixed with the image stream to obtain the video). In order to obtain two video streams in real time, it is necessary to ensure that the audio signal corresponding to the protagonist video and the audio signal corresponding to the original video can be encoded simultaneously. Therefore, it is necessary to copy one first audio data stream into two second audio data streams, and then encode these two second audio data streams concurrently to obtain the audio stream corresponding to the generated protagonist video and the audio stream related to the generated original video.
[0012] In conjunction with the first aspect, in some embodiments, the first target audio data is obtained based on one second audio data stream; and the second target audio data is obtained based on another second audio data stream, specifically including: the terminal encodes the second audio data stream and the second second audio data stream to obtain two encoded audio data streams; wherein one encoded audio data stream is used as the first target audio data stream, and the other encoded audio data stream is used as the second target audio data stream.
[0013] In the above embodiments, the first audio data includes the audio signals related to the generated protagonist video and the original video. However, there is only one first audio data stream (related to hardware configuration), and it has not yet been encoded (the audio needs to be encoded to obtain an audio stream when generating the video, which is then mixed with the image stream to obtain the video). In order to obtain two video streams in real time, it is necessary to ensure that the audio signal corresponding to the protagonist video and the audio signal corresponding to the original video can be encoded simultaneously. Therefore, it is necessary to copy one first audio data stream into two second audio data streams, and then encode these two second audio data streams concurrently to obtain the audio stream corresponding to the generated protagonist video and the audio stream related to the generated original video.
[0014] In conjunction with the first aspect, in some embodiments, the terminal determines the number of stream configurations N, specifically including: the terminal detecting the second operation on the second control, and in response to the operation, the terminal determining the number of stream configurations N; after displaying the recording interface, detecting the operation on the first object; and in response to the operation on the first object, the terminal updating the number of stream configurations N.
[0015] In the above embodiment, the second control is a protagonist mode control. When the terminal detects an operation on the protagonist mode control, N can be determined. Subsequently, the terminal can also update N. The timing of updating N is when an operation to determine the focus object is detected (i.e., an operation on the first object), which ensures that the changed N can be obtained even if N changes.
[0016] In conjunction with the first aspect, in some embodiments, the terminal determines the number of stream configurations N, specifically including: the terminal detecting the second operation on the second control, and in response to the operation, the terminal determining the number of stream configurations N; after displaying the recording interface, detecting an operation to stop recording the second video; and in response to the operation to stop recording the second video, the terminal updating the number of stream configurations N.
[0017] In the above embodiment, the second control is the protagonist mode control. When the terminal detects an operation on the protagonist mode control, N can be determined. Subsequently, the terminal can also update N. The timing for updating N is when the operation of stopping the recording of the protagonist video (second video) is detected, which ensures that the changed N can be obtained even if N changes.
[0018] In conjunction with the first aspect, in some embodiments, the terminal includes a first audio encoding module and a second audio encoding module. The terminal encodes one channel of second audio data and another channel of second audio data. Specifically, the first audio encoding module encodes based on the one channel of second audio data, while the second audio encoding module encodes based on the other channel of second audio data.
[0019] In the above embodiments, the first audio data includes the audio signals related to the generated protagonist video and the original video. However, there is only one first audio data stream (related to hardware configuration), and it has not yet been encoded (the audio needs to be encoded to obtain an audio stream when generating the video, which is then mixed with the image stream to obtain the video). In order to obtain two video streams in real time, it is necessary to ensure that the audio signal corresponding to the protagonist video and the audio signal corresponding to the original video can be encoded simultaneously. Therefore, it is necessary to copy one first audio data stream into two second audio data streams, and then encode these two second audio data streams concurrently to obtain the audio stream corresponding to the generated protagonist video and the audio stream related to the generated original video.
[0020] In conjunction with the first aspect, in some embodiments, the first audio data includes the first audio signal as well as the timestamp, data length, sampling rate, and data format of the first audio signal.
[0021] In conjunction with the first aspect, in some embodiments, the second audio data includes the first audio signal as well as the timestamp and data length of the first audio signal.
[0022] Secondly, this application provides an electronic device comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, and the one or more processors calling the computer instructions to cause the electronic device to execute:
[0023] In the above embodiments, the terminal can capture two video streams during a single recording process. One video stream can be the original video (first video) mentioned in the embodiments, and the other video stream can be the main character video (second video) mentioned in the embodiments. The main character video always includes a focus-tracking object (i.e., the first object). During video recording, the terminal's microphone can collect audio signals, and the terminal can perform real-time processing based on these collected audio signals to obtain the corresponding audio signals from the main character video and the original video. It should be understood that because the video recording process is real-time, the audio processing should also be real-time. This ensures that the terminal can obtain two audio signals in a single recording. It should be understood that the first audio data here includes the audio signals involved in generating the main character video and the original video, but this first audio data is only one stream (related to hardware configuration), and at this time, the first audio data has not yet been encoded (when generating video, the audio needs to be encoded to obtain an audio stream, which is then mixed with the image stream to obtain the video). In order to obtain two video streams in real time, it is necessary to ensure that the audio signal corresponding to the main video and the audio signal corresponding to the original video can be encoded simultaneously. Therefore, it is necessary to copy one first audio data stream into two second audio data streams, and then encode them concurrently based on these two second audio data streams to obtain the audio stream corresponding to the main video and the audio stream involved in generating the original video.
[0024] Thirdly, this application provides a chip system applied to an electronic device, the chip system including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform the methods described in the first aspect or any embodiment of the first aspect.
[0025] In the above embodiments, the terminal can capture two video streams during a single recording process. One video stream can be the original video (first video) mentioned in the embodiments, and the other video stream can be the main character video (second video) mentioned in the embodiments. The main character video always includes a focus-tracking object (i.e., the first object). During video recording, the terminal's microphone can collect audio signals, and the terminal can perform real-time processing based on these collected audio signals to obtain the corresponding audio signals from the main character video and the original video. It should be understood that because the video recording process is real-time, the audio processing should also be real-time. This ensures that the terminal can obtain two audio signals in a single recording. It should be understood that the first audio data here includes the audio signals involved in generating the main character video and the original video, but this first audio data is only one stream (related to hardware configuration), and at this time, the first audio data has not yet been encoded (when generating video, the audio needs to be encoded to obtain an audio stream, which is then mixed with the image stream to obtain the video). In order to obtain two video streams in real time, it is necessary to ensure that the audio signal corresponding to the main video and the audio signal corresponding to the original video can be encoded simultaneously. Therefore, it is necessary to copy one first audio data stream into two second audio data streams, and then encode them concurrently based on these two second audio data streams to obtain the audio stream corresponding to the main video and the audio stream involved in generating the original video.
[0026] Fourthly, this application provides a computer program product containing instructions that, when run on an electronic device, cause the electronic device to perform the method described in the first aspect or any embodiment of the first aspect.
[0027] In the above embodiments, the terminal can capture two video streams during a single recording process. One video stream can be the original video (first video) mentioned in the embodiments, and the other video stream can be the main character video (second video) mentioned in the embodiments. The main character video always includes a focus-tracking object (i.e., the first object). During video recording, the terminal's microphone can collect audio signals, and the terminal can perform real-time processing based on these collected audio signals to obtain the corresponding audio signals from the main character video and the original video. It should be understood that because the video recording process is real-time, the audio processing should also be real-time. This ensures that the terminal can obtain two audio signals in a single recording. It should be understood that the first audio data here includes the audio signals involved in generating the main character video and the original video, but this first audio data is only one stream (related to hardware configuration), and at this time, the first audio data has not yet been encoded (when generating video, the audio needs to be encoded to obtain an audio stream, which is then mixed with the image stream to obtain the video). In order to obtain two video streams in real time, it is necessary to ensure that the audio signal corresponding to the main video and the audio signal corresponding to the original video can be encoded simultaneously. Therefore, it is necessary to copy one first audio data stream into two second audio data streams, and then encode them concurrently based on these two second audio data streams to obtain the audio stream corresponding to the main video and the audio stream involved in generating the original video.
[0028] Fifthly, this application provides a computer-readable storage medium including instructions, characterized in that, when the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method described in the first aspect or any embodiment of the first aspect.
[0029] In the above embodiments, the terminal can capture two video streams during a single recording process. One video stream can be the original video (first video) mentioned in the embodiments, and the other video stream can be the main character video (second video) mentioned in the embodiments. The main character video always includes a focus-tracking object (i.e., the first object). During video recording, the terminal's microphone can collect audio signals, and the terminal can perform real-time processing based on these collected audio signals to obtain the corresponding audio signals from the main character video and the original video. It should be understood that because the video recording process is real-time, the audio processing should also be real-time. This ensures that the terminal can obtain two audio signals in a single recording. It should be understood that the first audio data here includes the audio signals involved in generating the main character video and the original video, but this first audio data is only one stream (related to hardware configuration), and at this time, the first audio data has not yet been encoded (when generating video, the audio needs to be encoded to obtain an audio stream, which is then mixed with the image stream to obtain the video). In order to obtain two video streams in real time, it is necessary to ensure that the audio signal corresponding to the main video and the audio signal corresponding to the original video can be encoded simultaneously. Therefore, it is necessary to copy one first audio data stream into two second audio data streams, and then encode them concurrently based on these two second audio data streams to obtain the audio stream corresponding to the main video and the audio stream involved in generating the original video. Attached Figure Description
[0030] Figure 1 A schematic diagram providing supplementary explanation of the protagonist mode involved in the embodiments of this application;
[0031] Figure 2 This is a schematic diagram illustrating one method of entering protagonist mode in an embodiment of this application;
[0032] Figure 3 The diagram shows the interface for selecting the focus object in the preview mode of the main character mode;
[0033] Figure 4A and Figure 4B This is a scene illustration in the recording mode of the protagonist mode;
[0034] Figure 5A This is a schematic diagram illustrating the scenario of exiting the protagonist mode provided in the embodiments of this application;
[0035] Figure 5B This is an illustration of how to save and view videos recorded in protagonist mode;
[0036] Figure 6 An exemplary flowchart illustrating the recording of the original video and the main character's video in an embodiment of this application is shown;
[0037] Figure 7The diagram shows a schematic software architecture block diagram of a terminal that copies one audio data stream into N streams and performs concurrent encoding.
[0038] Figure 8 The diagram illustrates the interactive flow between modules when the terminal copies one audio data stream into N streams and performs concurrent encoding.
[0039] Figure 9 The diagram illustrates the interactive flow between modules when the terminal generates the original video and the main character's video, copies one audio data stream into N streams, and performs concurrent encoding.
[0040] Figures 10A-10D This is an example diagram showing how a terminal obtains the first audio data based on audio signals collected from M microphones.
[0041] Figure 11 This is a schematic diagram illustrating how to extract the audio data corresponding to the main character's video from one second audio data stream and how to extract the audio data corresponding to the original video from another second audio data stream.
[0042] Figure 12 The diagram illustrates how a terminal copies one audio data stream into N streams and performs concurrent encoding to generate N video streams.
[0043] Figure 13 This is a diagram illustrating how the terminal copies the first audio data into two channels of second audio data.
[0044] Figure 14 This is a schematic diagram of the terminal structure provided in the embodiments of this application. Detailed Implementation
[0045] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.
[0046] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0047] In one approach, when the terminal records video in autofocus shooting mode, after identifying the main subject and ending recording, a final video is obtained. This video is a "main subject video" that includes the main subject. However, the surrounding image content in the main subject video is often incomplete. Thus, other image content besides the main subject is ignored during the recording process in the final recorded video (main subject video). Therefore, although a main subject video is obtained after a recording is completed, the environment in which the main subject was located during recording (the state and actions of objects around the main subject, etc.) cannot be obtained.
[0048] This application provides a video processing method that can be applied to terminals capable of recording video, such as mobile phones and tablets.
[0049] In this embodiment, after determining the main subject to be recorded, the terminal can simultaneously generate two videos: one is the main subject video, and the other is the original video. The original video includes the original image captured by the camera. The main subject video includes images obtained by identifying the main subject in the original image and then cropping out the portion of the image containing the main subject; this image can be referred to as the main subject image. During video recording, the terminal can simultaneously display both the original video and the main subject video for user preview.
[0050] In this way, after the main subject is identified, the terminal can either record a video of the main subject that includes the main subject in the image, or obtain the original video composed of the original images captured by the original camera.
[0051] The following describes in detail an exemplary user interface for a terminal implementing the video processing method provided in the embodiments of this application.
[0052] This application provides a video processing method in which a terminal can provide a recording function in protagonist mode, generating both the original video and the protagonist video during a single recording session. The following is in conjunction with... Figure 1 The main character mode involved in the embodiments of this application is further explained.
[0053] The definitions of terms used in the embodiments of this application will be explained below.
[0054] The "Protagonist Mode" can be understood as a mode that generates an additional tracking video while recording video on the device. The person in this tracking video can be considered the "protagonist" the user is interested in. The video corresponding to the "protagonist" is generated by cropping the video content corresponding to the "protagonist" from the video recorded by the device normally. It's understood that the device's Protagonist Mode offers both preview and recording modes. In preview mode, the preview interface is displayed on the device's screen. In recording mode, the recording interface is displayed on the device's screen.
[0055] It should be noted that the interface displayed on the terminal in both preview mode (before recording) and recording mode (during recording) can be referred to as the preview interface; the screen displayed in the preview interface in preview mode (before recording) will not generate video and be saved; the screen displayed in the preview interface in recording mode (during recording) will generate video and be saved. For ease of distinction, in the following text, the preview interface in preview mode (before recording) will be referred to as the preview interface; and the preview interface in recording mode (during recording) will be referred to as the recording interface.
[0056] The preview interface can include a large window and a small window. The large window can be a window with dimensions equal to or slightly smaller than the display screen, displaying the image captured by the camera. The image displayed in the large window in preview mode can be defined as the preview screen of the large window. The small window can be a window with dimensions smaller than the large window, displaying the image of the user-selected focus target (the aforementioned main subject). The terminal can select the focus target based on the tracking identifier associated with it. The image displayed in the small window in preview mode can be defined as the preview screen of the small window. It is understood that in preview mode, the terminal can display the image captured by the camera in the large window and the image of the focus target in the small window, but the terminal may not generate video or save the content displayed in the large and small windows.
[0057] The recording interface may include a large window and a small window. The large window can be a window with dimensions equal to or slightly smaller than the display screen, displaying the image captured by the camera. The image displayed in the large window during recording mode can be defined as the recording screen of the large window. The small window can be a window with dimensions smaller than the large window, displaying the image of the user-selected focus target (the aforementioned main character). The image displayed in the small window during recording mode can be defined as the recording screen of the small window. It is understood that in recording mode, the terminal can not only display the recording screens of the large and small windows, but also generate the video corresponding to the large window (which can also be the original video) and the video corresponding to the small window (which can also be the main character's video) recorded after starting recording mode. The original video can be saved when recording in the large window ends, and the main character's video can be saved when recording in the small window ends. This application embodiment does not limit the naming of the preview mode and the recording mode.
[0058] In the preview and recording interfaces, the image captured by the camera displayed in the large window can also be called the original image, while the image of the user-selected focus object displayed in the small window can also be called the main image.
[0059] It should be noted that the preview interface described in this embodiment can be understood as the preview mode of the terminal's camera application in the main mode; the recording interface can be understood as the recording mode of the terminal's camera application in the main mode. Further details will not be provided hereafter.
[0060] For example, the aforementioned protagonist mode functionality can be implemented in a camera application (also known as a camera or camera app). For instance, in a preview scenario, the protagonist mode preview interface on the terminal could be as follows: Figure 1 As shown in Figure a, the preview interface may include a large window 301, a small window 302, and multiple controls. These controls may include a start recording control 303, a first landscape / portrait switching control 304, a second landscape / portrait switching control 305, a small window close control 306, and an exit protagonist mode control 307. Optionally, the controls may also include a recording settings control 308, a flash control 309, and a zoom control 310, etc.
[0061] The terminal can display a preview image (the original image mentioned above) in a large window 301, which may include multiple people. When the terminal detects the presence of people in the preview image in the large window, a tracking icon associated with the person can be displayed in the preview image. For example, the tracking icon can be a tracking box displayed at the corresponding position of the person (e.g., tracking box 311 and tracking box 312). For example, a male person in the preview image can correspond to tracking box 311, and a female person can correspond to tracking box 312. The tracking box can prompt the user that the corresponding person can be set as the focus object or can be switched to the focus object (the protagonist mentioned above). When the terminal recognizes N people, M (M≤N) tracking boxes can be displayed in the large window. The terminal can set any person as the focus object to generate video content for that focus object. This application embodiment does not limit the "protagonist", where the "protagonist" can be a living being such as a person or an animal, or a non-living being such as a vehicle. It is understood that any item that can be identified based on an algorithm model can be used as the "protagonist" in this application embodiment. In this embodiment, the "protagonist" can be defined as the focus tracking object. The focus tracking object can also be called the protagonist object, the tracking target, the tracking object and the focus tracking target, etc. This embodiment uses a person as the "protagonist" for illustrative purposes, but this embodiment does not limit the concept of "protagonist".
[0062] In some embodiments, the tracking identifier can also be other forms of tracking identifier. For example, when the terminal identifies multiple focusable objects, a large window displays a tracking identifier corresponding to the focusable object near the focusable object. The tracking identifier can be numbers, letters, graphics, etc. When the user clicks on a tracking identifier, the terminal responds to the click operation and selects the focusable object. As another example, multiple focusable objects in the large window are marked with numbers, graphics, user images, or other tracking identifiers. The terminal can arrange multiple tracking identifiers at the edge or other locations of the large window display area, and the user can click on the tracking identifiers in the large window to select the focusable object. This application embodiment uses a tracking frame as an example to illustrate the recording method, but this application embodiment does not limit the form of the tracking identifier.
[0063] It should be noted that the terminal in this application embodiment can mark the corresponding tracking box for the person when it recognizes two or more people; the terminal can also mark the corresponding tracking box for the person when it recognizes a single person, or it can choose not to mark the tracking box, and there is no limitation here.
[0064] Optionally, the N figures displayed in the large window can be trackable objects, with the selected "main character" being the trackable object and the unselected characters being other objects. The tracking frame of the trackable object (e.g., tracking frame 311) can display different styles from the tracking frames of other objects (e.g., tracking frame 312). This makes it easier for users to distinguish the tracked characters.
[0065] In some embodiments, the shape, color, size, and position of the tracking frame are adjustable. For example, the tracking frame 311 for the object being tracked can be a dashed frame. The tracking frame 312 for other objects can be a combination of a dashed frame and a "+" sign. Besides different shapes, embodiments of this application can also set the color of the tracking frame; for example, tracking frame 311 and tracking frame 312 can be different colors. This allows for a direct distinction between the object being tracked and other objects. It is understood that the tracking frame can also have other display forms, as long as it satisfies the function of being triggered by the user to track the focusable object.
[0066] The tracking frame can be marked at any position on the trackable object; this application embodiment does not impose specific limitations. In a possible implementation, to avoid visual interference with the preview of the trackable object in a large window, the tracking frame can avoid the face of the trackable object; for example, the tracking frame can be marked at a relatively centered position on the body of the trackable object. The terminal can perform face recognition and body recognition. When the terminal recognizes a face, it can display the tracking frame. The terminal can determine the display position of the tracking frame based on face recognition and body recognition, and the tracking frame is displayed at the center of the body.
[0067] It should be noted that in some embodiments, the following scenario may occur: the preview screen in the large window includes N people, of which M (M≤N) are focusable objects with marked tracking boxes and NM people are not recognized by the terminal. In actual shooting, the terminal can display tracking boxes based on face recognition technology. When the terminal cannot capture a person's face (e.g., a person's back view), the terminal will not mark a tracking box for that person. This application does not limit the method for implementing the tracking box display.
[0068] In the preview scene, the small window 302 displays a preview image of the object being tracked. The preview image in the small window can be a portion of the preview image in the large window. In a possible implementation, the preview image in the small window is obtained by cropping the preview image in the large window based on the object being tracked. The terminal can crop the image in the large window according to an algorithm, and the small window obtains a portion of the image from the large window. In some embodiments, when the cropping calculation takes a long time, the image displayed in the small window in real time may be a cropped image of the first few frames displayed in real time in the large window. This application embodiment does not specifically limit the image displayed in the small window.
[0069] When the focus target changes, the person in the preview displayed in small window 302 changes accordingly. For example, if the focus target changes from a male to a female person, the preview displayed in small window 302 will also change accordingly. This will be discussed later. Figure 3 The scenarios for selecting or switching the focus target on the terminal will be described in detail, but will not be repeated here.
[0070] In some embodiments, the size, position, and landscape / portrait display mode of the small window are adjustable, and users can adjust the style of the small window according to their recording habits.
[0071] The preview interface also includes several controls, and the functions of each control are explained below.
[0072] Start Recording Control 303 is used to control the terminal to start recording in a large window or to start recording in both a large and small window.
[0073] The first landscape / portrait switching control 304 can be displayed in the large window and is used to adjust the landscape and portrait display of the small window.
[0074] The second landscape / portrait switching control 305 can be displayed in a small window and is also used to adjust the landscape and portrait display of the small window.
[0075] The small window close control 306 is used to close small windows in the terminal.
[0076] The Exit Protagonist Mode control 307 is used to exit protagonist mode and enter regular recording mode in the terminal.
[0077] Understandably, in a preview scenario, the preview interface may include a large window and a small window. The large window's preview image includes the trackable object. When the device selects a trackable object, the small window's preview image can display the trackable object in the center. In some scenarios, the trackable object may be moving; when the trackable object moves but does not leave the lens, the small window's preview image can continue to display the trackable object in the center. For example, if the trackable objects in the preview interface include male and female figures, and the device responds to the user's click on the tracking box for the male figure, the device selects the male figure as the trackable object and enters a preview window as follows: Figure 1 The interface shown in Figure a. Figure 1 In interface A, a small preview window displays a male character centered to the right of the female character. As the male character moves, the device continuously tracks and centers him within the small window. When the male character moves to the left of the female character, the device's interface... Figure 1 As shown in b. Figure 1 In the B interface, the preview window still displays the male character in the center, with the male character to the left of the female character.
[0078] For example, in a recording scenario, the recording interface in the protagonist mode on the terminal can be as follows: Figure 1 As shown in Figure c, the recording interface may include a large window 301, a small window 302, multiple controls, and a recording time. The controls may include a pause recording control 313, an end recording control 314, and an end small window recording control 315.
[0079] Unlike the preview scenario, in the recording scenario, the small window 302 displays the recorded image of the object being tracked. During recording, the terminal can generate an additional video stream recorded in a small window on top of the video recorded in the large window. Similar to the preview process, the recorded image in the small window can be a portion of the recorded image in the large window. In a possible implementation, the recorded image in the small window is obtained by cropping the recorded image of the large window based on the object being tracked. The two video streams are saved independently on the terminal. In this way, the video corresponding to the object being tracked can be obtained without subsequent manual editing of the entire video, making the operation simple and convenient and improving the user experience.
[0080] The recording interface may include multiple controls, and the functions of the controls are explained below.
[0081] The pause recording control 313 is used to pause video recording. Recording in both the large and small windows can be paused simultaneously. When the recording interface does not include the small window, the pause recording control 313 can pause only the recording in the large window.
[0082] The End Recording control 314 is used to end video recording. Recording of the large window and the small window can be ended simultaneously. When the recording interface does not include the small window, the End Recording control 314 can end recording only of the large window.
[0083] The "End Small Window Recording" control 315 is used to end the recording of a small window of video. The terminal can end the recording of a small window using the "End Small Window Recording" control 315, without affecting the recording of a large window.
[0084] Recording time indicates the duration of the currently recorded video. The recording time for the large window can be the same as or different from that for the small window.
[0085] Understandably, in a recording scenario, the recording interface may include a large window and a small window. The large window displays the trackable object. When the device selects a trackable object, the small window displays the trackable object in the center. In some scenarios, the trackable object may be moving. When the trackable object moves but does not leave the lens, the focus moves with the object, and the small window continues to display the trackable object in the center. For example, if the trackable objects in the recording interface include male and female figures, and the device responds to the user's click on the tracking box for the male figure, the device selects the male figure as the trackable object and enters a similar state. Figure 1 The interface shown in c. Figure 1 In the C interface, the recording screen in the small window displays the male figure centered, positioned to the right of the female figure. The focus is on the male figure's face, located slightly to the right of the center of the frame. As the male figure moves, the device continuously tracks and records him, displaying him centered in the small window. When the male figure moves to the left of the female figure, the device's interface... Figure 1 As shown in d. Figure 1 In the d interface, the recording screen in the small window still displays the male figure in the center, with the male figure to the left of the female figure. At this time, the focus is on the male figure's face area, which is located in the middle left part of the screen.
[0086] In this application embodiment, the shooting mode that generates an additional tracking video based on the tracking object is defined as the main character mode. This shooting mode can also be called tracking mode, etc., and this application embodiment does not limit it.
[0087] There are several ways to enter protagonist mode when recording.
[0088] For example, in combination Figure 2 This application provides a detailed description of one method for entering the protagonist mode in its embodiments.
[0089] In one possible implementation, the terminal is in Figure 2 The main interface shown in Figure a allows the terminal to enter [the following interface] when it detects that the user has opened the camera application 401. Figure 2 The image preview interface shown in b is shown below. This preview interface may include a preview screen and shooting mode selection controls. The preview screen can display the scene captured by the terminal's camera in real time. The shooting mode selection controls include, but are not limited to: "Portrait" control, "Photo" control, "Video" control 402, "Pro" control, and "More" control 403.
[0090] When the terminal detects that the user clicks the "Record" control 402, the terminal switches from the photo preview interface to... Figure 2The video preview interface shown in c; the video preview interface may include, but is not limited to: a protagonist mode control 404 for receiving triggers to enter protagonist mode, a recording settings control for receiving triggers to enter settings, a filter control for receiving triggers to enable filter effects, and a flash control for setting flash effects.
[0091] The terminal can enter protagonist mode based on the protagonist mode control 404 in the video preview interface. For example, when a user clicks the protagonist mode control 404 in the interface, the terminal responds to the click operation and enters the protagonist mode. Figure 2 The preview interface shown in d; in the preview interface, there can be multiple shooting objects in the large window. The terminal can identify the multiple shooting objects based on the image content of the large window. The multiple shooting objects can be used as trackable objects. The terminal's preview interface can mark tracking boxes for each trackable object.
[0092] In another possible implementation, the terminal can enter protagonist mode without using the "Record" control. Instead, protagonist mode can be treated as a new mode. For example, the terminal can display protagonist mode through the "More" control 403, and then select protagonist mode to enter it. Further details after entering protagonist mode can be found in the foregoing content and will not be repeated here.
[0093] In this application, the embodiments are combined with Figure 3 This document provides a detailed description of the scenes in the preview mode of the protagonist mode, and combines this application's embodiments with... Figure 4A as well as Figure 4B This section provides a detailed explanation of the scenes in the recording mode within the protagonist mode.
[0094] First, let's introduce the scenarios in preview mode.
[0095] For example, Figure 3 The diagram shows an interface for selecting the focus object in the preview mode of the main character mode.
[0096] like Figure 3 As shown: The terminal enters the preview mode of the main character mode, as follows. Figure 3 As shown in Figure a, the terminal can display a preview interface for the main character mode. The preview interface includes multiple focusable objects, each of which can be marked with its own tracking box (for example, tracking box 311 for male characters and tracking box 312 for female characters).
[0097] The terminal can determine the user-selected focus object based on the user's operation on the tracking frame (e.g., a click). For example, if a user wants to preview the focus image of a male figure in a small window on the terminal, they can click the tracking frame 311 corresponding to the male figure. The terminal responds to this click operation and enters a window such as... Figure 3 The interface shown in b.
[0098] like Figure 3 As shown in Figure b, when the terminal selects a male figure as the focus target, a small window floats in the large window of the preview interface. The small window displays the image corresponding to the position of the focus target in the large window. The focus target can be centered in the small window, highlighting its "protagonist" status. Optionally, after the tracking frame of the focus target is triggered, the color of the tracking frame can change, such as becoming lighter, darker, or changing to another color. The shape of the tracking frame can also change; for example, the tracking frame 311 for the male figure is a dashed frame, and the tracking frame 312 for the female figure is a combination of a dashed frame and a "+". In this embodiment, the tracking frame styles of the focus target and other objects can be any combination of different colors, sizes, and shapes to facilitate user differentiation between the focus target and other objects in the large window. Optionally, after the tracking frame of the focus target is triggered, the corresponding tracking frame can disappear, preventing the user from repeatedly selecting the already selected focus target.
[0099] Understandably, in the preview mode of the main character mode, users can manually change the focus object after selecting it, such as... Figure 3 As shown in the interface (b), when the terminal receives the user's click on the tracking box 312 of the female figure, it enters the following... Figure 3 The interface shown in Figure c. At this point, the focus target in the small window switches from the male character to the female character. The tracking frame state of the character changes; for example, the color and shape of the female character's tracking frame 312 change, while the male character's tracking frame 311 reverts to its unselected style. The changes in the tracking frame style can be seen in [reference needed]. Figure 3 The relevant descriptions in the interface shown in b are not repeated here.
[0100] Optionally, the terminal switches the focus target in preview mode, and the object displayed in the preview window changes from the original focus target to the new focus target. To make the switching process smoother, this application embodiment also provides a dynamic effect for switching the focus target. For example, the design of the dynamic effect is described below using a male figure as the original focus target and a female figure as the new focus target.
[0101] In one possible implementation, the large window of the preview interface includes both male and female figures, while the small window displays the male figure as the focus target. When the terminal detects a click on the tracking box for the female figure, the preview in the small window can switch from focusing on the male figure to a panoramic view, and then back to focusing on the female figure. For example, if the small window initially displays the male figure centered, after the user clicks on the tracking box for the female figure, the cropping ratio between the small and large window previews increases. The small window preview can include more content from the large window preview, which can be represented by the male figure and his background gradually shrinking in the small window until the small window can simultaneously display a panoramic view of both the male and female figures. Subsequently, the small window centers and enlarges the female figure within the panoramic view. Optionally, the panoramic view can be a proportionally scaled-down preview of the large window, or it can be an image cropped from the shared area of the male and female figures within the large window preview.
[0102] In another possible implementation, the large preview window includes both male and female figures, while the smaller window displays the male figure as the focus target. When the terminal detects a click on the tracking box targeting the female figure, the focus in the smaller window's preview gradually shifts from the male figure to the female figure. For example, if the smaller window initially displays the male figure in the center, after the user clicks on the female figure's tracking box, the cropping ratio between the smaller and larger preview windows remains unchanged, but the smaller window's preview will be cropped closer to the female figure, maintaining the original cropping ratio. For instance, if the female figure is to the left of the male figure, during the switching of focus targets, the male figure and his background in the smaller window shift to the right until the female figure is centered in the smaller window.
[0103] In this way, when the terminal switches the focus target, the transition from the original focus target to the new focus target in the small window is smoother, improving the user's recording experience.
[0104] The above embodiments have described the preview mode of the main character mode. The recording mode of the main character mode will now be described with reference to the accompanying drawings. In the recording mode of the main character mode, the terminal can start a small window to record video of the focused object and save the video.
[0105] The following is combined Figure 4A and Figure 4B This section provides a detailed explanation of the scenes in the recording mode within the protagonist mode.
[0106] In a recording mode of the main character mode, videos in both the large and small windows can start recording simultaneously. Figure 4AIn interface A, the large window previews the object being tracked (e.g., a male person), while the small window displays a preview of the object being tracked. When the terminal detects a click on the start recording control 303 in the large window, the terminal enters a state similar to... Figure 4A The interface shown in b. The terminal simultaneously starts recording in both the large window and the small window. The small window can display the object being tracked in the large window in real time. Figure 4A In the interface (b), a small window displays the recording screen and recording time. For example, the small window also displays a recording end control 315 for recording mode, while the recording start control 303 in the large window is converted into a recording pause control 313 and a recording end control 314 for recording mode. The large and small windows can each display their respective recording times, which can be consistent across the large and small windows. To improve the recording interface and reduce obstruction of the tracking object, the display position of the recording time in this embodiment can be as follows: Figure 4A As shown in b, the recording time can also be set at other locations that do not affect the recording.
[0107] Optionally, in some embodiments, when the terminal enters recording mode from preview mode, the first landscape / portrait switching control, the second landscape / portrait switching control, the zoom control, and the small window closing control may disappear, such as... Figure 4A Figure b. Some embodiments may also retain these controls, and this application embodiment does not limit this.
[0108] It should be understood that, Figure 4A In the scenario shown, the timing for the terminal to trigger the recording of both the original video and the main character's video is as follows: Before entering the recording mode in the main character mode, the terminal first detects an operation on the main character mode control. In response to this operation, the terminal prepares to start recording the main character's video (but has not yet started recording). Then, it detects the user's operation to determine the main character in the original image. So, when the terminal detects the user's operation on the start recording control, it can enter the recording process in response to this operation, and can record both the original video and the main character's video simultaneously.
[0109] In another recording mode scenario, the large window and the small window can be recorded sequentially. Figure 4B In interface A, the preview screen in the large window includes the object to be tracked; the terminal did not select the object to be tracked, causing the small window not to open. Responding to the user's click on the start recording control 303, the terminal starts recording in the large window and enters the following... Figure 4B The interface shown in b. Figure 4B In interface b, the large window displays the recording screen and recording time; the small window is not open on the terminal. During recording in the large window, when the terminal detects a click operation by the user selecting tracking box 311, the terminal displays... Figure 4B The interface shown in c. Figure 4BIn the C interface, the terminal can maintain recording in the large window and start recording in the small window.
[0110] The terminal can initiate video recording in a small window based on the above scenarios and obtain multiple video streams. It should be noted that the small window can display the image of the object being tracked in the large window, but the video recorded in the small window and the video recorded in the large window are multiple independent videos, not a picture-in-picture composite video in which a small window is nested within a large window's recording.
[0111] It should be noted that if the terminal does not enable recording in a small window, it will receive one video stream recorded in the large window. If the terminal enables recording in a small window, it will receive one video stream recorded in the large window and multiple video streams recorded in the small windows. For example, while recording video in the large window, the terminal can open small window recording multiple times. When the terminal detects a click operation on the control to end small window recording, it can end the small window recording and receive one video stream. When the small window resumes recording, the terminal will receive a new video stream. The number of videos received in the small window is related to the number of times the small window recording is opened.
[0112] It should be understood that the timing of the terminal triggering the recording of the original video and the main character video is as follows: Before entering the main character mode recording mode, the terminal first detects an operation on the main character mode control. In response to this operation, the terminal prepares to start recording the main character video (but has not yet started recording). Then, before entering the main character mode recording mode, no operation by the user to determine the main character in the original image is detected. Therefore, when the terminal detects the user's operation on the start recording control, it can enter the main character mode recording mode in response to this operation. At this time, the terminal only records the original video, not the main character video. During the recording of the original video, the terminal can detect the user's operation to determine the main character in the original image. In response to this operation, the terminal can trigger the recording of the main character video.
[0113] It should be understood that the terminal can determine the main character not only when it detects a user's selection of the main character in the original image, but also at other times. For example, after detecting an operation on the main character mode control (e.g., a click operation), before detecting an operation on the start recording control, and if no user selection of the main character in the original image is detected within a first time threshold, the terminal can automatically determine a main character based on the original image. In this case, the method of determining the main character includes: identifying a moving object in the central region of the original image and determining that object as the main character. Here, the center of the original image is the geometric center of the original image.
[0114] Users can choose to exit protagonist mode and revert to the regular recording mode when they do not need to use protagonist mode.
[0115] The following is combined Figure 5AThe scenario of exiting the protagonist mode provided in the embodiments of this application will be described.
[0116] like Figure 5A As shown: For example, when the terminal receives a request for... Figure 5A When the recording control 314 is clicked in interface a to end recording, the terminal can simultaneously end recording in both the large and small windows. Furthermore, upon entering... Figure 5A The terminal can save both the video recorded in the large window (original video) and the video recorded in the small window (original video) simultaneously when recording ends. The terminal can save the two videos to the same path or to different paths. For example, the video from the large window and the video from the small window can be saved to a folder in the album, or the video from the large window can be saved to a regular path, and the video from the small window can be saved to a folder in the main character mode of the album. This embodiment does not limit the save path for the two videos.
[0117] Figure 5A In interface b, both the large and small windows have finished recording and returned to preview mode. When a click operation is received on the exit protagonist mode control 307, the terminal enters the following state: Figure 5A The interface shown in Figure c. The terminal resumes normal recording mode. Of course, the terminal can also... Figure 5A After detecting a click on the end recording control 314 in interface A, the main character mode is exited directly, and the following is displayed: Figure 5A The interface shown in c is an example. Alternatively, the user can trigger the terminal to exit the main character mode through gestures or other methods; this embodiment of the application does not impose any restrictions on this.
[0118] After entering protagonist mode, the terminal can acquire at least one video stream and save and view it when recording ends.
[0119] The following is combined Figure 5B The process of saving and viewing videos recorded in protagonist mode is described.
[0120] Optionally, users can browse videos recorded in a large window and multiple videos recorded in a small window based on the camera application's album. The display order of the multiple videos can be the recording order, that is, the terminal can sort them according to the end time or start time of the recording. The display order of the multiple videos can also be the reverse recording order, that is, the terminal can arrange the videos in reverse order according to the end time or start time of the recording.
[0121] Optionally, videos recorded in the large window and videos recorded in the small window can be displayed as video thumbnails on the same album interface. To easily distinguish between videos recorded in the large window and videos recorded in the small window, the terminal can set an identifier for the video recorded in the small window. For example, the terminal can add an outer border, font, and graphics to the video recorded in the small window, and the terminal can also set the size of the thumbnail of the video recorded in the small window, so that there is a size difference between the thumbnails of the videos recorded in the small window and the large window. It is understood that the embodiments of this application do not limit the form of video thumbnails in the album.
[0122] For example, the arrangement order of video thumbnails can be as follows: Figure 5B As shown. Users can base their decisions on... Figure 5B The terminal browses the recorded video on the interface shown in Figure a. When the terminal detects a click on the video icon 1601, the terminal enters... Figure 5B The interface shown in b. Figure 5B The b interface displays thumbnails of currently recorded videos. Among them, videos 1602 and 1603 can be multiple videos obtained from a single recording using the protagonist mode. The video order will be explained below in conjunction with specific recording scenarios.
[0123] For example, the terminal records in protagonist mode. The recording interface includes a large window and a small window. The large window displays male and female characters, while the small window displays male characters. When the terminal detects a click on the start recording control, the large window records video 1602, which includes both male and female characters, while the small window records video 1603, which is focused on the male character. After 40 seconds, the terminal detects a click on the end recording control, and ends recording both videos 1602 and 1603, saving both videos.
[0124] In the above recording scenario, the terminal used the main character mode to record a single video, resulting in two video streams.
[0125] In some instances, the terminal can save multiple video streams based on their sequential end times, with the first saved video stream arranged in order. Figure 5B It's located towards the back of the B interface.
[0126] It is understood that the embodiments of this application exemplarily illustrate the arrangement order of video thumbnails and the saving order of videos, and the embodiments of this application do not impose any limitations on this.
[0127] It is understandable that the video recorded in the large window (the original video) can include both images and sound, and the video recorded in the small window (the main character's video) can also include both images and sound. For example, when the terminal crops the recorded screen of the small window from the image in the large window to obtain the video of the small window, the terminal can also synchronize the sound to the video of the small window.
[0128] It should be understood that in the preview mode of the main character mode, in order to enhance the user's recording experience, the terminal can provide other functions, including but not limited to the following functions.
[0129] In some embodiments, the terminal may experience a loss of focus object in focus tracking mode. That is, after focus tracking is established and the focus object is initially detected, it may then become undetectable. In this case, the preview image in the small window can be the last frame of the preview image displayed before the focus object is lost, and the small window is in a masked state. After the focus object is lost, if it is detected again, tracking of that object can continue, the masked state of the small window is removed, and a preview image including the focus object can be displayed. After the focus loss time reaches a certain threshold (e.g., 5 seconds), the small window can be closed. If the device is in focus tracking recording mode at this time, the video (focus tracking video) in the small window can be saved after closing the small window.
[0130] In other embodiments, the size of the small window can be customized, allowing users to adjust it to an appropriate size and view the preview of the tracked object more clearly. The display position of the small window within the interface can also be adjusted.
[0131] Figure 6 The diagram shows an exemplary flowchart of recording the original video and the main character's video in an embodiment of this application.
[0132] During the recording process, the terminal processes the audio data collected by the microphone and the images collected by the camera in real time to obtain the original video and the main character's video. Figure 6 The process described herein is to process a single frame of image captured at the same time and its corresponding audio data in real time to obtain video. The processing of other frames of image and other frames of video is the same, and you can refer to the relevant description.
[0133] For a detailed description of the process, please refer to the following description of steps S101-S108.
[0134] S101. The terminal collects audio signals through M microphones to obtain the first audio data.
[0135] At the first moment, the terminal collects audio signals through M microphones. During the recording of the original video and the main character's video, the terminal can enable some or all of the microphones to collect audio signals. The terminal has S microphones, and M of them can be selected for audio signal collection, where M is an integer greater than or equal to 1 and S is an integer greater than or equal to M.
[0136] The terminal can process the collected audio signal to obtain one channel of first audio data.
[0137] The first audio data includes audio signals collected by M microphones of the terminal, along with information about these audio signals. This information may include timestamps, data length, sampling rate, and data format. The timestamps represent the time when the M microphones collected the audio signals. The data length represents the length of the audio signals included in the first audio data. The data format can represent the format of the first audio data. The audio signals collected by the M microphones can be considered as the audio signals involved in generating N videos.
[0138] In some instances, the audio signals collected by the M microphones of the terminal can be encapsulated into a single audio frame, and the first audio data can be obtained based on this audio frame.
[0139] In other instances, the audio signals collected by the M microphones of the terminal can be encapsulated into a single audio frame to obtain M audio frames, and the first audio data can be obtained based on these M audio frames.
[0140] S102. The terminal copies the first audio data into two second audio data streams based on the number of streams configured.
[0141] In step S102, the terminal generates the original video and the main character video, with a total of 2 videos.
[0142] The second audio data may include some or all of the content from the first audio data. Typically, the second audio data may include a portion of the content from the first audio data; this portion can be used to generate the video, while other content that cannot be used to generate the video does not need to be copied.
[0143] In some instances, the terminal can copy the audio signal, the timestamp of the audio signal, and the data length of the audio signal from the first audio data to obtain two identical second audio data streams. One of the second audio data streams includes the audio signal from the first audio data, as well as the timestamp of the audio signal and the data length of the audio signal.
[0144] In other instances, the second audio data may include, in addition to the audio signal from the first audio data, the timestamp of the audio signal, and the data length of the audio signal, a data format that represents the format (audiodata) of the second audio data. Specifically, the format (audiodata) of the second audio data can be: {buffer, buffersize, pts}, where buffer represents the audio signal and can be in array form, buffersize represents the data length of the audio signal, and pts represents the timestamp of the audio signal.
[0145] In subsequent steps S103, S107, and S108, the terminal can generate the original video and the main character video based on the two second audio data streams.
[0146] S103. The terminal concurrently encodes the two channels of second audio data to obtain two channels of encoded audio data.
[0147] In some instances, the terminal can concurrently encode two channels of second audio data to obtain two encoded audio data channels. One channel is designated as the first target audio data channel, and the other as the second target audio data channel.
[0148] S104. The terminal acquires the original image, tracks the main character based on the original image, and obtains the main character tracking information.
[0149] At the first moment, the terminal captures the original image through the camera. Each original image corresponds to a timestamp, which indicates the time when the camera captured the original image.
[0150] The terminal identifies the main character in the original image and obtains the main character tracking information, which may include at least one of the following: the main character's face region, the main character's body region, the main character's center coordinates, and the main character's corresponding tracking status.
[0151] The focus tracking state refers to whether the subject appears in the original image. In some examples, this focus tracking state can include whether the subject appears in the original image or not.
[0152] S105. The terminal processes the protagonist's tracking information to obtain the protagonist's coordinates.
[0153] When the terminal determines that the subject is in the original image, it can process the subject tracking information to obtain the subject's coordinates.
[0154] In some examples, the protagonist's coordinates can be the protagonist's center coordinates.
[0155] In other examples, the protagonist's coordinates can be a coordinate system determined based on the protagonist's face region, the protagonist's body region, and the protagonist's center coordinates. For example, face coordinates.
[0156] In some instances, when the subject is not visible in the original image during focus tracking, the subject tracking information can be omitted from determining the new subject coordinates. The subject coordinates calculated based on the previous frame of the original image can be used as the subject coordinates in step S105. If the subject remains not visible in the original image for a third preset time period, the terminal can stop recording the subject video and save the already recorded subject video.
[0157] In some cases, the subject's coordinates can be in the camera sensor (i.e., image sensor) coordinate system (a two-dimensional coordinate system), used to represent the subject's position in the image. In other cases, it can also be in the image coordinate system.
[0158] S106. The terminal generates a main character image based on the main character's coordinates and the original image.
[0159] The terminal can determine the image region where the main character's coordinates are located in the original image, crop the original image to obtain the content of that image region, and generate the main character image based on the content of that image region. In some instances, the image region can be centered on the main character's coordinates.
[0160] Each main character image corresponds to a timestamp, which is the same as the timestamp of the original image.
[0161] S107. The terminal generates the original video based on the first target audio data and the original image.
[0162] The terminal encodes the original image to obtain the encoded original image. The timing of this process may be the same as or different from the timing of the terminal encoding the second audio data.
[0163] When the terminal determines that the timestamp of the audio signal included in the first target audio data is the same as the timestamp of the original image, the terminal can perform mixing based on the first target audio data and the encoded original image to generate the original video.
[0164] S108. The terminal generates a protagonist video based on the second target audio data and the protagonist image.
[0165] The terminal encodes the main character image to obtain the encoded main character image. The timing of this process may be the same as or different from the timing of the terminal encoding the second audio data.
[0166] When the terminal determines that the timestamp of the audio signal included in the second target audio data is the same as the timestamp of the protagonist image, the terminal performs mixing based on the second target audio data and the encoded protagonist image to generate the protagonist video.
[0167] In particular, the encoding of the first target audio data in step S107 and the encoding of the second target audio data in step S108 are performed concurrently, which ensures that the original video and the main character video are obtained in real time.
[0168] There is no specific order in which steps S101 and S104 are executed; they can be executed either after the terminal starts recording.
[0169] It should be understood that the aforementioned steps S101-S108 are examples using the video recorded by the terminal as the original video and the main video, with a stream configuration of 2. In reality, the stream configuration is N, where N is an integer greater than or equal to 1. If the stream configuration is N, the terminal can copy the first audio data into N channels of second audio data. Then, based on these N channels of audio data, they are concurrently encoded to generate N videos, where one channel of second audio data is used to generate one video.
[0170] In existing technologies, the terminal acquires only one channel of first audio data. In this application, to obtain N videos in real time, the first audio data channel needs to be copied to obtain N channels of second audio data. These N channels of second audio data are then concurrently encoded to obtain audio streams corresponding to N videos. The audio stream corresponding to each video is then mixed with the corresponding image stream to generate N videos in real time. A detailed description of this process can be found below.
[0171] Figure 7 The diagram shows a schematic software architecture block diagram of a terminal that copies one audio data stream into N streams and performs concurrent encoding.
[0172] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the system is divided into four layers, from top to bottom: the application layer, the application framework layer, the hardware abstraction layer, and the hardware layer.
[0173] The application layer can include a series of application packages. This application package may include a camera application. Besides a camera application, it can also include other applications, such as a gallery application, etc. (also referred to as applications).
[0174] This camera application may include a protagonist mode, an audio recording management module, and an encoder control module.
[0175] Among them, the protagonist mode is used to provide the function of recording protagonist videos when the terminal enters the mode.
[0176] The audio recording management module can be used to start or stop the terminal's microphone from acquiring audio signals, and can also be used to copy the acquired first audio data to obtain N channels of second audio data.
[0177] The encoder control module can be used to concurrently encode N channels of second audio data.
[0178] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0179] The application framework layer may include an audio framework layer, a camera framework layer, and an encoding framework layer.
[0180] The audio framework layer includes an audio recording module and an audio streaming module. The audio recording module pre-stores the microphones used by the terminal when recording the original video and the main character's video. For example, the terminal may have S microphones, and the audio recording module can select M of them to acquire audio signals, where M is an integer greater than or equal to 1 and S is an integer greater than or equal to M.
[0181] The encoding framework layer provides encoding capabilities to the upper-level encoder control module. Specifically, the audio encoder in this encoding framework layer provides the audio encoding module (including a first audio encoding module and a second audio encoding module) within the encoder control module with the ability to encode audio data. Similarly, the image encoder in this encoding framework layer provides the image encoding module (including a first image encoding module and a second image encoding module) within the encoder control module with the ability to encode images.
[0182] The Hardware Abstraction Layer (HAL) is an interface layer located between the kernel layer (not shown) and the hardware layer. Its purpose is to abstract the hardware and provide a virtual hardware platform for the operating system.
[0183] The hardware abstraction layer can include abstraction layers for different hardware components. For example, there could be a camera abstraction layer and an audio abstraction layer. The camera abstraction layer can be used to abstract the image sensor (the sensor in a camera), performing various calculations based on the images captured by the image sensor to obtain the output. The audio abstraction layer can be used to abstract the microphone, performing various calculations based on the audio signals captured by the microphone to obtain the output.
[0184] The audio abstraction layer includes an instruction processing module, an audio input stream module, and an audio algorithm module.
[0185] The instruction processing module can be used to receive and process instructions from the upper-level audio framework layer. For example, it can receive an instruction from the audio abstraction layer to start the microphone to collect audio signals, and after receiving the instruction, it can start the microphone to collect audio signals.
[0186] The audio input stream module processes the audio signal captured by the microphone to obtain a first audio data stream. This first audio data stream includes information about the audio signal captured by the microphone, such as timestamp, data length, sampling rate, and data format. This first audio data stream is then sent to the audio recording module.
[0187] The audio algorithm module can interact with the audio input stream module to process the audio signals captured by the microphone.
[0188] The hardware layer may include the terminal's microphone and the terminal's image sensor (the sensor in the camera), the microphone being used to acquire audio signals and the image sensor being used to acquire images.
[0189] In this process, the original image captured by the image sensor is transmitted to the camera frame layer through the camera abstraction layer. The camera frame layer can then generate a main character image based on the original image. Both the original image and the main character image are then transmitted to the first image encoding module and the second image encoding module, respectively, for encoding, resulting in the encoded original image and the encoded main character image. Subsequently, the first mixing module can mix the encoded original image and one channel of encoded second audio data to generate the original video. Later, the first mixing module can also mix the encoded main character image and another channel of encoded second audio data to generate the main character video.
[0190] The following embodiments of this application will be combined with Figure 7 This section provides an illustrative processing flow of the terminal copying the first audio data to obtain N channels of second audio data, and then concurrently encoding these N channels of second audio data.
[0191] Step 1: Upon confirming entry into protagonist mode, the audio recording management module detects an operation on the start recording control. It can then use the audio framework layer and audio abstraction layer to call the terminal's M microphones to begin capturing audio signals. Furthermore, after detecting the operation on the start recording control, it determines the number of audio copies, N. Step 1 refers to... Figure 7 The gray circle in the middle is marked ①.
[0192] In some possible cases, refer to the aforementioned... Figure 2 As described in section c, when an operation targeting the protagonist mode control 404 is detected, the audio recording management module can determine that it has entered protagonist mode. After entering protagonist mode, the terminal can display the information described above. Figure 2The user interface shown in d is as follows.
[0193] For example, when the audio recording management module determines that it has entered the main character mode, it detects an operation on the start recording control and can determine the number of streams, N. Furthermore, the audio recording management module can send a first instruction to the audio recording module in the audio framework layer, which is used to start the microphones to collect audio signals. After receiving the instruction, the audio recording module can determine the M microphones for collecting audio signals, and then send a second instruction to the audio stream module, which in turn starts the M microphones to collect audio signals. The audio stream module can then send the second instruction to the instruction processing module, which, after parsing the second instruction, can start the M microphones in the terminal to collect audio signals.
[0194] Step 2: The audio signals collected by the M microphones can be transmitted to the audio abstraction layer. This layer can process the audio signals to generate the first audio data. Then, the first audio data is transmitted to the audio recording management module through the audio framework layer. Step 2 refers to… Figure 7 The gray circle in the middle is marked ②.
[0195] For example, the audio signals collected by M microphones can be transmitted to the audio input stream module in the audio abstraction layer. This audio input stream module can encapsulate the audio signals collected by the microphones into frames to obtain first audio data. Then, the audio input stream module can send the first audio data to the audio stream module in the audio framework layer, and then the audio stream module sends the first audio data to the audio recording module. The audio recording module then sends the first audio data to the audio recording management module.
[0196] Step 3: The audio recording management module can copy the first audio data into N identical second audio data streams based on the configured number N. Here, we take N equal to 2 as an example. Then, these two second audio data streams are sent to the encoder control module respectively. Step 3 is... Figure 7 The gray circle in the middle is marked ③.
[0197] In some possible cases, refer to the aforementioned provisions. Figure 4A The content shown in b and Figure 4B The content shown in c indicates that when the number of streams N equals 2, the terminal records the main character's video while simultaneously recording the original video.
[0198] In some possible cases, the number of stream configurations N can be greater than 2. For example, if the terminal records two main character videos plus one original video, the terminal can identify the two main characters and track them separately to obtain the two main character videos.
[0199] In some other possible cases, the number of stream configurations N can also be equal to 1. Refer to the aforementioned... Figure 4B As shown in b, the terminal is not recording the main character's video at this time, but only recording the original video, so N=1.
[0200] For example, the audio recording management module then sends one second audio data stream to the first audio encoding module in the encoder control module and another second audio data stream to the second audio encoding module in the encoder control module.
[0201] Step 4: The encoder control module can concurrently encode the N channels of second audio data using the encoder. Step 4 refers to... Figure 7 The gray circle in the diagram is marked ④.
[0202] For example, the first audio encoding module can call the audio encoder to encode one channel of second audio data to obtain the first target audio data. Furthermore, the second audio encoder can call the audio encoder to encode another channel of second audio data to obtain the second target audio data.
[0203] In this way, when N videos are generated during a single recording process, although only one first audio data is obtained at the audio abstraction layer and only that first audio data is transmitted to the camera application at the application layer, the audio data related to the generation of the video can be copied from the first audio data in the camera application to obtain N second audio data involved in generating N videos.
[0204] Figure 8 The diagram illustrates the interactive flow between modules when the terminal copies one audio data stream into N streams and performs concurrent encoding.
[0205] The following is combined Figure 7 as well as Figure 8 This document describes in detail the process by which each module of the terminal copies the first audio data to obtain N channels of second audio data during video recording, and then concurrently encodes these N channels of second audio data.
[0206] The process can be described in detail below with reference to steps S201-S216.
[0207] S201. The terminal enters main mode and detects an operation on the start recording control.
[0208] As mentioned above Figure 2 The interface shown in 'c' indicates that when an operation (such as a click) is detected on the main character mode control, the terminal determines to enter main character mode.
[0209] In some instances, after the terminal determines it has entered protagonist mode but before detecting an operation on the start recording control, if a protagonist has been identified, then in response to the detected start recording control operation, the terminal can simultaneously begin recording both the original video and the protagonist's video. For example, as mentioned above... Figure 4A As shown, once an operation (such as a click) on the start recording control is detected, the terminal can begin recording both the original video and the main character's video.
[0210] In other embodiments, after the terminal determines that it has entered the main character mode, but before detecting an operation on the start recording control, if no main character has been determined, then in response to the detected start recording control operation, the terminal may start recording the original video but not the main character video. For example, the aforementioned... Figure 4B a and Figure 4B As shown in b, if an operation (such as a click) is detected on the start recording control, the terminal can start recording the original video but not the main character's video.
[0211] S202. The number of stream configurations N obtained in protagonist mode.
[0212] After the terminal starts recording video, the number of streams configured, N, can be determined. N is the number of video streams configured. The number of streams configured, N, is the number of video streams involved when the terminal generates video. It should be understood that different videos recorded by the terminal correspond to different video streams. If the terminal records N videos, then the number of video streams is N.
[0213] The timing for the protagonist mode to acquire the number N of flow configurations includes, but is not limited to, the following timings.
[0214] (1) Upon detecting an operation on the start recording control, the terminal begins recording video in response to the operation, and the terminal can determine the number of streams configured, N. After the terminal starts recording video, the main character mode can re-determine the number of streams configured, N, in the following situations, including but not limited to:
[0215] Case 1: If the operation of determining the main character is detected, the terminal can redetermine the number of stream configurations N and update the number of stream configurations N.
[0216] Scenario 2: In the event that the main character is lost, the terminal can re-determine the number of stream configurations N and update the number of stream configurations N.
[0217] Scenario 3: If an operation to stop recording the main character's video is detected, such as the aforementioned operation to end the small window recording control, the terminal can redetermine the number of stream configurations N and update the number of stream configurations N.
[0218] (2) After the terminal starts recording video, the terminal can determine the number of streams N according to a certain time frequency.
[0219] If the video recorded by the terminal only contains the original video and does not include the main character's video, then the number of streams configured, N, equals 1. If the video recorded by the terminal includes both the original video and the main character's video, then the number of streams configured, N, equals 2.
[0220] In other possible scenarios, the terminal may generate more videos during a single recording session, in which case N may be greater than or equal to 3.
[0221] In some possible scenarios, the protagonist mode can determine the number of stream configurations N based on the stream management module (not shown in the diagram). This stream management module can be used to configure video streams, and N is the number of video streams configured.
[0222] S203. In protagonist mode, the number N of the stream configuration is sent to the audio recording management module.
[0223] After receiving the number of streams N, the audio recording management module can perform initialization settings and execute the following steps S204 and S205 to obtain the content related to the subsequent acquisition of N channels of second audio data.
[0224] S204. The audio recording management module determines the number of audio copies N based on the number of stream configurations N.
[0225] The audio copy number N is used to determine the number of copies of the first audio data in step S213 below. This audio copy number N is equal to the stream configuration number N. This allows N videos to be generated based on the copied N audio data during video recording.
[0226] In some possible cases, after determining the number of audio copies N, the audio recording management module can also configure channel information. In this case, when N equals 2, stereo can be configured. Subsequently, the audio signal captured by the microphone can be divided into stereo channels to obtain a stereo audio signal as the first audio data.
[0227] S205. The audio recording management module sends the first instruction to the audio framework layer to start the microphone to collect audio signals.
[0228] The first instruction is used to notify the start of microphone acquisition of audio signals. At this time, the terminal may not be sure which microphone is involved in the acquisition of audio signals, and can determine it in step S206 below.
[0229] The first instruction carries parameters related to the terminal's microphone acquiring audio signals, including sampling rate and audio format. The sampling rate represents the frequency at which the microphone acquires the audio signal. The audio format determines the sampling depth of each microphone's audio signal, i.e., the number of bytes occupied by each sample point in the audio signal acquired by a microphone, such as 8 bits or 16 bits. In addition to these parameters, other parameters may be included, such as channel configuration, which will not be described in detail here.
[0230] After receiving the first instruction, the audio framework layer can generate the second instruction involved in step S206 below, and execute step S206 below.
[0231] In some instances, such as reference Figure 7 Specifically, the audio recording management module can send this first instruction to the audio recording module in the audio framework layer.
[0232] S206. The audio framework layer determines the M microphones involved in acquiring audio signals and sends a second instruction to the audio abstraction layer to start acquiring audio signals using the M microphones.
[0233] The second instruction is used to notify the activation of M microphones to collect audio signals, that is, to determine which microphones are involved in collecting audio signals. For example, if the terminal has three microphones, it can determine to use the first and second microphones to collect audio signals. The specific number of microphones to collect audio signals is preset in the audio frame layer.
[0234] The second instruction still carries parameters related to the microphone's acquisition of audio signals, including sampling rate and audio format. For an explanation of these parameters, please refer to the aforementioned description of step S205; they will not be repeated here.
[0235] In some instances, such as reference Figure 7 The audio recording module in the audio framework layer can be used to send the second instruction to the audio stream module, which in turn sends the second instruction to the instruction processing module in the audio abstraction layer.
[0236] S207. The audio abstraction layer starts M microphones to collect audio signals.
[0237] After receiving the second instruction, the audio abstraction layer can determine the M microphones involved in acquiring the audio signal and activate those M microphones to acquire the audio signal. That is, the terminal has S microphones, and M of them can be selected for audio signal acquisition, where M is an integer greater than or equal to 1 and S is an integer greater than or equal to M.
[0238] S208. The M microphones collect audio signals.
[0239] The audio signals collected by the M microphones may include the sound of the subject being filmed (e.g., the aforementioned subject 1) as well as ambient sounds.
[0240] S209. The M microphones send the collected audio signals to the audio abstraction layer.
[0241] In some instances, refer to Figure 7 The M microphones can send the collected audio signals to the audio input stream module in the audio abstraction layer.
[0242] S210. The audio abstraction layer processes the audio signal to obtain a first audio data, which includes audio signals collected by M microphones and information about the audio signals.
[0243] The first audio data includes audio signals collected by M microphones and information about those audio signals.
[0244] This audio abstraction layer encapsulates the audio signals captured by M microphones into frames and determines the information of these audio signals. This information may include timestamps, data length, sampling rate, and data format. The timestamp represents the time when the M microphones captured the audio signals. The data length represents the length of the audio signal included in the first audio data. The data format can represent the format of the first audio data.
[0245] The audio abstraction layer encapsulates the audio signals collected by M microphones into frames, which may include encapsulating the audio signals collected by the M microphones of the terminal into a frame of audio signal, and obtaining the first audio data based on the frame of audio signal.
[0246] In some instances, refer to Figure 7 The process involved in step S210 can be completed by the audio input stream module included in the audio abstraction layer.
[0247] Subsequently, through the following steps S211 and S212, the terminal can transmit the first audio data from the audio abstraction layer to the audio recording management module in the application layer.
[0248] S211. The audio abstraction layer sends the first audio data to the audio framework layer.
[0249] In some instances, refer to Figure 7 Specifically, the first audio data can be sent from the audio input stream module in the audio abstraction layer to the audio stream module in the audio framework layer.
[0250] S212. The audio framework layer sends the first audio data to the audio recording management module.
[0251] In some instances, refer to Figure 7 The first audio data can be sent from the audio stream module in the audio framework layer to the audio recording module, and then the audio recording module can send the first audio data to the audio recording management module.
[0252] S213. The audio recording management module copies the first audio data according to the number of audio copies N, to obtain N channels of second audio data. The second audio data includes audio signals collected by M microphones and some information of the audio signals.
[0253] The N-channel second audio data is the same audio data.
[0254] Generally speaking, the second audio data may include some content from the first audio data. This part of the content can be used to generate a video, while other content that cannot be used to generate a video can be left uncopied.
[0255] In some instances, the audio recording management module can copy the audio signal, timestamp, and data length of the first audio data to obtain N identical second audio data streams. One of the second audio data streams includes the audio signal from the first audio data, along with its timestamp and data length.
[0256] The timestamp of the audio signal can be used to match the timestamp of the image. Only audio signals and images with the same timestamp can be mixed after encoding to generate video.
[0257] The data length of the audio signal can be used for verification in step S216 below. The purpose of the verification is to determine the validity of the second audio data. In step S216, valid second audio data can be encoded, while invalid second audio data can be left unencoded.
[0258] In some instances, the audio recording management module can only copy the first audio data if it determines that the first audio data is correct. One way to determine the correctness of the first audio data includes determining that the data length of the audio signal included in the first audio data is equal to the minimum data length. This minimum data length is determined based on the aforementioned parameters such as the sampling rate and data format.
[0259] Subsequently, the terminal can execute the following steps S214-S216 to encode the N channels of the second audio data concurrently, thereby obtaining N channels of encoded audio data to generate N videos in real time.
[0260] S214. The audio recording management module sends N channels of second audio data to the encoder control module.
[0261] In some instances, refer to Figure 7 The audio recording management module can send the N channels of second audio data to different audio encoding modules in the encoder control module. Taking N=2 as an example, the audio recording management module can send one channel of second audio data to the first audio encoding module in the encoder control module, and send the other channel of second audio data to the second audio encoding module in the encoder control module.
[0262] S215. The encoder control module instructs the audio encoder to encode the N channels of second audio data.
[0263] In some instances, refer to Figure 7 Different audio encoding modules (including the first audio encoding module and the second audio encoding module) in the encoder control module can call the audio encoder to encode the received second audio data.
[0264] S216. The audio encoder encodes based on N channels of second audio data.
[0265] The audio encoder can encode the N channels of the second audio data concurrently to obtain N channels of encoded audio data.
[0266] Figure 9 The diagram illustrates the interactive flow between modules when the terminal generates the original video and the main character's video, copies one audio data channel into N channels, and performs concurrent encoding.
[0267] Figure 9 The following example illustrates the process of generating the original video and the main character's video on the terminal, with a streaming configuration of N=2.
[0268] The terminal can process the audio signal collected by the microphone to obtain the target audio data corresponding to the original video (denoted as the first target audio data), and process the audio signal collected by the microphone to obtain the target audio data corresponding to the main character video (denoted as the second target audio data).
[0269] In some possible cases, the first target audio data and the second target audio data may differ. The sound information in the first target audio data can be the sound information included in the audio signal captured by the microphone. The sound information in the second target audio data can be the sound information after focusing the audio signal captured by the microphone. In this case, the sound information from the direction of the main character in the audio signal captured by the microphone is highlighted, while other sound information is suppressed.
[0270] The audio tracking processing may include: during the recording of the protagonist's video, determining the protagonist's direction, enhancing the audio signal in the protagonist's direction from the audio signal collected by the microphone, and suppressing the sound signals in other directions to obtain the second target audio data.
[0271] In some possible cases, the first target audio data corresponding to the original video may include two-channel audio data, one channel denoted as left channel audio signal 1 and the other as right channel audio signal 1. The second target audio data corresponding to the main character video may also include two-channel audio data, one channel denoted as left channel audio signal 2 and the other as right channel audio signal 2. This allows for stereo sound when playing both the original video and the target audio data corresponding to the main character video.
[0272] To ensure that both the first target audio data corresponding to the original video and the second target audio data corresponding to the main character's video are stereo audio, channel configuration is required before the microphone captures the audio signal. Here, we configure it to have 4 channels. Of these 4 channels, two channels correspond to the original video, and the other two channels correspond to the main character's video.
[0273] After the microphone acquires the audio signal, the terminal can copy the audio signal to obtain two audio signals. Then, the terminal performs mixing processing on one of the audio signals to obtain a two-channel target audio signal 1 (corresponding to the original video), and performs focus tracking processing on the other audio signal to obtain a two-channel target audio signal 2 (corresponding to the main character's video). Based on a 4-channel configuration, 4-channel audio data is obtained using target audio signal 1 and target audio signal 2. Detailed information about this process can be found in the description of step S311 below, and will not be repeated here.
[0274] Subsequently, the terminal can obtain the first audio data based on the 4-channel audio data. This first audio data is then copied to obtain two channels of second audio data. These two channels of second audio data are identical, each containing 4 channels of audio data. Then, the terminal can extract the first channel audio data (the left channel audio data corresponding to the original video) and the second channel audio data (the right channel audio data corresponding to the original video) from one of the second audio data channels to obtain the audio data corresponding to the original video. This audio data is then encoded to obtain the target audio data (first target audio data) corresponding to the original video. Furthermore, the terminal can extract the third channel audio data (the left channel audio data corresponding to the main character's video) and the fourth channel audio data (the right channel audio data corresponding to the main character's video) from the other second audio data channel to obtain the audio data corresponding to the main character's video. This audio data is then encoded to obtain the target audio data (second target audio data) corresponding to the main character's video.
[0275] During this process, the interaction flow of each module can be referred to in the following description of steps S301-S317.
[0276] S301. The terminal enters main mode and detects an operation on the start recording control.
[0277] This step S301 is the same as the aforementioned step S201, and can be referred to the aforementioned description of step S201, which will not be repeated here.
[0278] S302. The number of stream configurations N is obtained in the protagonist mode.
[0279] This step S302 is the same as the aforementioned step S202, and can be referred to the aforementioned description of step S202, which will not be repeated here.
[0280] S303. In protagonist mode, the number N of the stream configuration is sent to the audio recording management module.
[0281] This step S303 is the same as the aforementioned step S203, and can be referred to the aforementioned description of step S203, which will not be repeated here.
[0282] S304. The audio recording management module determines the number of audio copies N based on the number of stream configurations N.
[0283] Step S304 is the same as step S204 mentioned above, and can be referred to the description of step S204 mentioned above, which will not be repeated here.
[0284] The S305 audio recording management module can configure channel information based on the number of video copies N. When N equals 2, it can configure 4 channels.
[0285] The audio channel information can include the number of audio channels and the video corresponding to each channel. For example, when N equals 2, configuring 4 channels means that two of the channels correspond to the main character's video, and the other two channels correspond to the original video.
[0286] S306. The audio recording management module sends the first instruction to the audio framework layer to start the microphone to collect audio signals.
[0287] Step S306 is the same as step S205 mentioned above, and can be referred to the description of step S205 above, which will not be repeated here.
[0288] S307. The audio framework layer determines the M microphones involved in acquiring audio signals and sends a second instruction to the audio abstraction layer to start acquiring audio signals using the M microphones.
[0289] Step S307 is the same as step S206 mentioned above, and can be referred to the description of step S206 above, which will not be repeated here.
[0290] S308. The audio abstraction layer starts M microphones to collect audio signals.
[0291] Step S308 is the same as step S207 mentioned above, and can be referred to the description of step S207 above, which will not be repeated here.
[0292] S309. The M microphones collect audio signals.
[0293] Step S309 is the same as step S208 mentioned above, and can be referred to the description of step S208 above, which will not be repeated here.
[0294] S310. The M microphones send the collected audio signals to the audio abstraction layer.
[0295] Step S310 is the same as step S209 mentioned above, and can be referred to the description of step S209 above, which will not be repeated here.
[0296] S311. The audio abstraction layer copies the audio signal into two audio signals. One audio signal is used to generate the target audio signal 1 corresponding to the original video, and the other audio signal is used to generate the target audio signal 2 corresponding to the main video. Based on the configuration of 4 channels, the target audio signal 1 and the target audio signal 2 are used to obtain 4-channel audio data. Then, based on the 4-channel audio data, the first audio data is obtained. The first audio data includes the audio signals corresponding to the two videos and the information of the audio signals.
[0297] The first audio data includes audio signals corresponding to the two videos, i.e., 4-channel audio data, which is obtained based on audio signals captured by M microphones. The first audio data also includes information about the audio signals, including timestamps, data length, sampling rate, and data format. The timestamp represents the time when the M microphones captured the audio signals. The data length represents the length of the 4-channel audio signals included in the first audio data. The data format can represent the format of the first audio data.
[0298] In some possible cases, step S311 can be completed collaboratively by the audio input stream module and the audio algorithm module included in the audio abstraction layer.
[0299] The following is based on Figures 10A-10D An exemplary description is given of the process by which a terminal obtains first audio data based on audio signals collected from M microphones.
[0300] like Figure 10A The following explanation uses M=3 as an example. The microphones used to collect audio signals include a top microphone, a back microphone, and a bottom microphone. The audio signals collected by the three microphones can include: a mono audio signal 1 collected by the back microphone (meaning each frame of audio signal 1 contains only mono data); a mono audio signal 2 collected by the top microphone (meaning each frame of audio signal 2 contains only mono data); and a mono audio signal 3 collected by the bottom microphone (meaning each frame of audio signal 3 contains only mono data). The terminal can then copy the audio signals collected by these three microphones to obtain two identical audio signals. One of these audio signals is then mixed to obtain the target audio signal 1 corresponding to the original video. This process can be referred to the following... Figure 10B The description is as follows. Furthermore, audio tracking processing is performed on another audio signal to obtain the second audio signal corresponding to the main character's video. This process can be referenced in the following description. Figure 10C The description.
[0301] The audio abstraction layer can process one audio signal using mixing algorithms such as aligned mixing, direct addition, and addition with clamping to obtain the target audio signal 1 for two channels. For example... Figure 10B As shown, one audio signal can include sound signal 1 captured by the back microphone, sound signal 2 captured by the top microphone, and sound signal 3 captured by the bottom microphone. The audio abstraction layer can mix all the above sound signals based on mixing algorithms such as aligned mixing, direct addition, and addition with clamping to obtain the two-channel target audio signal 1 (corresponding to the original video). That is to say, each frame of the target audio signal 1 includes two-channel audio data (referred to as left channel audio data and right channel audio data). Figure 10B The 1 in the original video represents the left channel audio data, and the 2 represents the right channel audio data.
[0302] like Figure 10C As shown, the other audio signal may include sound signal 1 captured by the back microphone, sound signal 2 captured by the top microphone, and sound signal 3 captured by the bottom microphone. The audio abstraction layer can perform gain processing on the sound signal in the direction of the main character in sound signal 1 using beamforming technology to obtain a beamformed audio signal. This enhances the sound signal in the direction of the main character while suppressing sound signals from other directions. Then, the audio abstraction layer can use the ambient noise included in sound signals 2 and 3 as a reference and employ noise suppression algorithms (such as anti-reflection algorithms) to filter out the ambient noise from the beamformed audio signal, thus obtaining the target audio signal 2. Each frame of this second audio signal includes dual-channel audio data (referred to as left channel audio data and right channel audio data). Figure 10CIn the text, 1 represents the left channel audio data corresponding to the main character's video, and 2 represents the right channel audio data corresponding to the main character's video.
[0303] Subsequently, the audio abstraction layer can obtain 4-channel audio data using target audio signal 1 and target audio signal 2 based on the 4-channel configuration.
[0304] by Figure 10D The example described above uses placing the 4-channel audio data in a buffer. This buffer can be configured with four channel regions: 1 represents channel 1, 2 represents channel 2, 3 represents channel 3, and 4 represents channel 4. Channel 1 records the left channel audio data from the original video, channel 2 records the right channel audio data, channel 3 records the left channel audio data from the main character's video, and channel 4 records the right channel audio data from the main character's video. The audio abstraction layer obtains the 4-channel audio data by: filling the left channel audio data from target audio signal 1 into channel 1, filling the right channel audio data from target audio signal 1 into channel 2, filling the left channel audio data from target audio signal 2 into channel 3, and filling the right channel audio data from target audio signal 2 into channel 4. Figure 10D The way the 4-channel audio data is arranged has the following functions: when frame loss occurs in the buffer, it can reduce the data loss of target audio signal 1 and target audio signal 2, and ensure the data integrity of target audio signal 1 and target audio signal 2 as much as possible.
[0305] Subsequently, the audio abstraction layer can use this 4-channel audio data as the audio signal in the first audio data. In addition to the 4-channel audio data, the first audio data may also include information about the audio signal. Information about the audio signal can be found in the preceding description and will not be repeated here.
[0306] S312. The audio abstraction layer sends the first audio data to the audio framework layer.
[0307] Step S312 is the same as step S211 mentioned above, and can be referred to the description of step S211 above, which will not be repeated here.
[0308] S313. The audio framework layer sends the first audio data to the audio recording management module.
[0309] Step S313 is the same as step S212 mentioned above, and can be referred to the description of step S212 mentioned above, which will not be repeated here.
[0310] S314. The audio recording management module copies the first audio data according to the number of audio copies N to obtain two channels of second audio data. The two channels of second audio data include the audio signals corresponding to the two videos and some information of the audio signals.
[0311] The two second audio data streams are the same audio data.
[0312] The second audio data includes the audio signals corresponding to the two videos, i.e., 4-channel audio data.
[0313] Generally speaking, the second audio data may include some content from the first audio data. This part of the content can be used to generate a video, while other content that cannot be used to generate a video can be left uncopied.
[0314] In some instances, the audio recording management module can copy the 4-channel audio signal and some information (such as timestamps and data length) from the first audio data to obtain two identical second audio data streams. That is, either second audio data stream includes a 4-channel audio signal and some information, such as timestamps and data length.
[0315] The timestamp can be used to match the timestamp of the image. Only audio signals and images with the same timestamp can be mixed after encoding to generate video.
[0316] S315. The audio recording management module sends two channels of second audio data to the encoder control module.
[0317] Step S315 is the same as step S214 mentioned above, and can be found in the description of step S214 above, so it will not be repeated here. Simply change N to 2.
[0318] S316. The encoder control module extracts the audio data corresponding to the main video from one second audio data channel and extracts the audio data corresponding to the original video from another second audio data channel.
[0319] The encoder control module can extract the audio data corresponding to the main character video from the 4-channel audio data included in one audio data stream, and extract the audio data corresponding to the original video from the 4-channel audio data included in another audio data stream. Subsequently, the encoder control module can instruct the audio encoder to encode the audio data corresponding to the original video and the audio data corresponding to the main character video. This allows the audio encoder to perform the following step S317.
[0320] Figure 11 This is a schematic diagram illustrating how to extract the audio data corresponding to the main character's video from one second audio data stream and how to extract the audio data corresponding to the original video from another second audio data stream.
[0321] like Figure 11 As mentioned above, channel 1 represents the left channel audio signal corresponding to the original video, and channel 2 represents the right channel audio signal corresponding to the original video; channel 3 represents the left channel audio signal corresponding to the main character's video, and channel 4 represents the right channel audio signal corresponding to the main character's video.
[0322] like Figure 11 As shown in (a), the audio data corresponding to the original video can be obtained by separating the audio data of channel 1 and channel 2 from the 4-channel audio data included in the second audio data.
[0323] like Figure 11 As shown in (b), the audio data corresponding to the main character's video can be obtained by separating the audio data of channel 3 and channel 4 from the 4-channel audio data included in the other second audio data.
[0324] S317. The audio encoder encodes audio data based on the audio data corresponding to the main character video and the audio data corresponding to the original video.
[0325] An audio encoder can encode the audio data corresponding to the main video to obtain encoded audio data (which can be called the first target audio data), and encode the audio data corresponding to the original video to obtain encoded audio data (which can be called the second target audio data).
[0326] Figure 12 The diagram illustrates how a terminal copies one audio data stream into N streams and performs concurrent encoding to generate N video streams.
[0327] The following is combined Figure 12 The document describes in detail the process by which the terminal copies the first audio data to obtain N channels of second audio data during video recording, and then concurrently encodes these N channels of second audio data to generate N videos.
[0328] For a detailed description of the process, please refer to the following description of steps S401-S407.
[0329] S401. The terminal enters main mode, detects the operation on the start recording control, and begins recording video.
[0330] As mentioned above Figure 2 The interface shown in 'c' indicates that when an operation (such as a click) is detected on the main character mode control, the terminal determines to enter main character mode.
[0331] After detecting an action on the start recording control, the terminal begins recording video.
[0332] In some instances, after the terminal determines it has entered protagonist mode but before detecting an operation on the start recording control, if a protagonist has been identified, then in response to the detected start recording control operation, the terminal can simultaneously begin recording both the original video and the protagonist's video. For example, as mentioned above... Figure 4A As shown, once an operation (such as a click) on the start recording control is detected, the terminal can begin recording both the original video and the main character's video.
[0333] In other embodiments, after the terminal determines that it has entered the main character mode, but before detecting an operation on the start recording control, if no main character has been determined, then in response to the detected start recording control operation, the terminal may start recording the original video but not the main character video. For example, the aforementioned... Figure 4B a and Figure 4B As shown in b, if an operation (such as a click) is detected on the start recording control, the terminal can start recording the original video but not the main character's video.
[0334] It should be understood that in step S401, in addition to entering the main character mode through the main character mode control, the terminal can also enter the main character mode in other ways, such as setting the main character mode as a shooting mode option, and then selecting the main character mode after entering the camera application.
[0335] S402. The terminal determines the number of stream configurations N, and determines the number of audio copies N based on the number of stream configurations N.
[0336] After the terminal starts recording video, the number of stream configurations N can be determined. The number of audio copies N is equal to the number of stream configurations N.
[0337] If the video recorded by the terminal only contains the original video excluding the main character's video, then the number of streams configured, N, equals 1. If the video recorded by the terminal includes both the original video and the main character's video, then the number of streams configured, N, equals 2.
[0338] In other possible scenarios, the terminal may generate more videos during a single recording session, in which case N may be greater than or equal to 3.
[0339] The timing when the terminal triggers the determination of the number N of flow configurations includes, but is not limited to, the following timings.
[0340] (1) Upon detecting an operation on the start recording control, the terminal starts recording video in response to the operation, and the terminal can determine the number of streams configured, N. After the terminal starts recording video, the terminal can re-determine the number of streams configured, N, under the following circumstances, including:
[0341] Case 1: If the operation of determining the main character is detected, the terminal can redetermine the number of stream configurations N and update the number of stream configurations N.
[0342] Scenario 2: In the event that the main character is lost, the terminal can re-determine the number of stream configurations N and update the number of stream configurations N.
[0343] Scenario 3: If an operation to stop recording the main character's video is detected, such as the aforementioned operation to end the small window recording control, the terminal can redetermine the number of stream configurations N and update the number of stream configurations N.
[0344] (2) After the terminal starts recording video, the terminal can determine the number of streams N according to a certain time frequency.
[0345] S403. Start M microphones to collect audio signals, and obtain one first audio data based on the audio signals.
[0346] Once the terminal detects the start recording control, it can activate M microphones to capture audio signals.
[0347] The terminal can enable some or all of the microphones to collect audio signals. The terminal has S microphones, and M microphones can be selected from them for audio signal collection, where M is an integer greater than or equal to 1 and S is an integer greater than or equal to M.
[0348] The description of the first audio data can be found in the aforementioned content, and will not be repeated here.
[0349] S404. Copy the first audio data into N channels of second audio data according to the number of audio copies N.
[0350] The N-channel second audio data is the same audio data.
[0351] Generally speaking, the second audio data may include some content from the first audio data. This part of the content can be used to generate a video, while other content that cannot be used to generate a video can be left uncopied.
[0352] In some instances, the terminal can copy the audio signal, the timestamp of the audio signal, and the data length of the audio signal from the first audio data to obtain N identical second audio data streams. One of the second audio data streams includes the audio signal from the first audio data, the timestamp of the audio signal, and the data length of the audio signal.
[0353] The relevant descriptions of the audio signal information can be found in the preceding content and will not be repeated here.
[0354] In some instances, the first audio data can only be copied after the terminal confirms its correctness. One way to determine the correctness of the first audio data includes determining that the data length of the audio signal included in the first audio data is equal to a minimum data length. This minimum data length is determined based on the aforementioned parameters such as the sampling rate and data format.
[0355] For example, such as Figure 13 The diagram shown illustrates how the terminal copies the first audio data into two channels of second audio data.
[0356] like Figure 13 As shown, the first audio data includes an audio signal, as well as information about the audio signal such as data length, timestamp, and sampling rate. The terminal copies the first audio data into two second audio data streams, one of which may include the audio signal, as well as information about the audio signal such as data length and timestamp.
[0357] In other instances, the second audio data may include, in addition to the audio signal from the first audio data, the timestamp of the audio signal, and the data length of the audio signal, a data format that represents the format (audiodata) of the second audio data. Specifically, the format (audiodata) of the second audio data can be: {buffer, buffersize, pts}, where buffer represents the audio signal and can be in array form, buffersize represents the data length of the audio signal, and pts represents the timestamp of the audio signal.
[0358] The timestamp of the audio signal can be used to match the timestamp of the image. Only audio signals and images with the same timestamp can be mixed after encoding to generate video.
[0359] The data length of the audio signal can be used for verification in step S405 below. The purpose of the verification is to determine the validity of the second audio data. In step S405, valid second audio data can be encoded, while invalid second audio data can be left unencoded.
[0360] S405. Encode N channels of second audio data to obtain N channels of encoded audio data.
[0361] The terminal can concurrently encode N channels of second audio data to obtain N channels of encoded audio data.
[0362] In some instances, before determining whether to encode a second audio data stream, the terminal first checks whether the length of the second audio data stream is greater than 0.
[0363] If the length of the second audio data is greater than 0, the terminal determines that the second audio data is valid and encodes it to obtain a channel of encoded audio data.
[0364] If the length of the data included in the second audio data is equal to 0, the terminal determines that the second audio data is invalid and may not encode it.
[0365] S406. Generate N videos based on N encoded audio data, with different encoded audio data used to generate different videos.
[0366] For example, in the case of generating the original video and the main character video (N=2), the terminal can generate the original video based on one encoded audio data (first target audio data) and generate the main character video based on another encoded data (second target audio data).
[0367] The process of generating the original video includes: the terminal encodes the original image to obtain the encoded original image. When the terminal determines that the timestamp of the audio signal included in the first target audio data is the same as the timestamp of the original image, the terminal can perform mixing based on the first target audio data and the encoded original image to generate the original video.
[0368] The process of generating the original video includes: the terminal encodes the main character image to obtain the encoded main character image. When the terminal determines that the timestamp of the audio signal included in the second target audio data is the same as the timestamp of the original image, the terminal can perform mixing based on the second target audio data and the encoded main character image to generate the main character video.
[0369] S407. An operation targeting the end recording control was detected. Video recording was stopped, and N videos were obtained.
[0370] The terminal can obtain N videos during a single recording process and can view all N videos. For example, during a single recording process, it can obtain the original video and the video of the main character, as mentioned above. Figure 5B This is an example interface for displaying the original video and the main character's video.
[0371] In this embodiment of the application, the large window can be referred to as the first preview window, the small window can be referred to as the second preview window, the protagonist mode control can be referred to as the second control, the protagonist mode can be referred to as the first mode, the end recording control can be referred to as the first control, the audio signal included in the first audio data can be referred to as the first audio signal, and the information of the audio signal included in the first audio data can be referred to as the information of the first audio signal.
[0372] The following describes an exemplary terminal provided in the embodiments of this application.
[0373] Figure 14 This is a schematic diagram of the terminal structure provided in the embodiments of this application.
[0374] The following description uses a terminal as an example to illustrate the embodiments. It should be understood that a terminal may have more or fewer components than shown in the figures, may combine two or more components, or may have different component configurations. The various components shown in the figures can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0375] The terminal may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0376] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal. In other embodiments of this application, the terminal may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0377] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0378] The controller can serve as the nerve center and command center of the terminal. Based on the instruction opcode and timing signals, the controller generates operation control signals to control the fetching and execution of instructions.
[0379] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0380] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include inter-integrated circuit (I2C) interfaces, inter-integrated circuit sound (I2S) interfaces, pulse code modulation (PCM) interfaces, etc.
[0381] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a limitation on the structure of the terminal. In other embodiments of this application, the terminal may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.
[0382] The charging management module 140 is used to receive charging input from the charger.
[0383] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110.
[0384] The terminal's wireless communication function can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor.
[0385] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the terminal can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.
[0386] The mobile communication module 150 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G in terminals. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc.
[0387] A modem processor may include a modulator and a demodulator.
[0388] The wireless communication module 160 can provide solutions for wireless communication applications on terminals, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), and other wireless communication solutions.
[0389] The terminal implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0390] Display screen 194 is used to display images, videos, etc.
[0391] The terminal can achieve shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0392] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and color. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0393] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the terminal may include one or N cameras 193, where N is a positive integer greater than 1.
[0394] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when a terminal selects a frequency, a DSP can perform a Fourier transform on the frequency energy.
[0395] Video codecs are used to compress or decompress digital video. A terminal can support one or more video codecs. This allows the terminal to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0396] An NPU (Neural Processing Unit) is a neural network (NN) computing processor that borrows from the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, to rapidly process input information and continuously learn. NPUs enable intelligent cognitive applications in terminals, such as image recognition, facial recognition, speech recognition, and text understanding.
[0397] Internal memory 121 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM).
[0398] Random access memory can include static random-access memory (SRAM), dynamic random access memory (DRAM), etc.
[0399] Non-volatile memory can include disk storage devices and flash memory.
[0400] Flash memory can be classified according to its operating principle, including NOR FLASH, NAND FLASH, 3D NAND FLASH, etc., and according to the level of the storage cell, it can be classified according to the level of the storage cell, including single-level cell (SLC) and multi-level cell (MLC), etc.
[0401] The random access memory can be directly read and written by the processor 110. It can be used to store executable programs (such as machine instructions) of the operating system or other running programs, as well as user and application data.
[0402] Non-volatile memory can also store executable programs and user and application data, and can be pre-loaded into random access memory for direct reading and writing by the processor 110.
[0403] The external memory interface 120 can be used to connect to external non-volatile memory to expand the terminal's storage capacity.
[0404] The terminal can implement audio functions, such as music playback and recording, through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, and an application processor.
[0405] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0406] The loudspeaker 170A, also known as a "loudspeaker", is used to convert audio electrical signals into sound signals.
[0407] The receiver 170B, also known as the "earpiece", is used to convert audio electrical signals into sound signals.
[0408] The microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals.
[0409] Touch sensor 180K, also known as "touch panel". Touch sensor 180K can be set on display screen 194. Touch sensor 180K and display screen 194 together form touch screen, also known as "touch screen".
[0410] Button 190 includes the power button, volume buttons, etc. Button 190 can be a mechanical button.
[0411] In this embodiment of the application, the processor 110 can call computer instructions stored in the internal memory 121 to cause the terminal to execute the video processing method in this embodiment of the application.
[0412] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0413] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0414] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0415] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method of video processing, the method comprising: Applied to a terminal, the terminal including a camera, the method includes: The recording interface includes a first control, a first preview window, and a second preview window; the first preview window is used to display the image captured by the camera. Acquire first audio data, which includes a first audio signal; the first audio signal is an audio signal collected by the microphone of the terminal at a first moment. Upon detecting a first operation on the first control, in response to the first operation, video recording is stopped, and a first video recorded based on the image displayed in the first preview window and a second video recorded based on the image displayed in the second preview window are saved; wherein... The method further includes: The first audio data is copied into N identical second audio data streams, each of which includes the first audio signal. N is the number of video streams recorded by the camera at the first moment. The video streams recorded at the first moment include the first video stream corresponding to the first preview window and the second video stream corresponding to the second preview window. The first video is obtained by mixing one second audio stream with the first video stream, and the second video is obtained by mixing another second audio stream with the second video stream.
2. The method according to claim 1, characterized in that, The first video is obtained by mixing a second audio stream with the first video stream, including: Encode the second audio data to obtain the first target audio data; The first video stream and the first audio data are mixed to obtain the first video; The second video is obtained by mixing the second audio data and the second video stream, including: The second audio data is encoded to obtain the second target audio data. The second video stream and the second audio data are mixed to obtain the second video.
3. The method of claim 1, wherein, The method further includes: At the first moment, the first preview window displays the first image; the first image is captured by the camera at the first moment. When a first object is detected at a first position in the first image, the terminal displays a second image in the second preview window; the second image is generated based on the first image, and the second image includes the first object. At the second moment, the first preview window displays the third image; the third image is captured by the camera at the second moment. When the first object is detected at the second position of the third image, the terminal displays a fourth image in the second preview window; the fourth image is generated based on the third image, and the fourth image includes the first object. Acquire first audio data, which includes a first audio signal and information about the first audio signal; the first audio signal is an audio signal collected by the microphone of the terminal at the first moment and recorded in the audio abstraction layer of the terminal.
4. The method according to claim 3, characterized in that, Before displaying the recording interface, the method further includes: Display a preview interface, the preview interface including a second control; Upon detecting a second operation on the second control, the camera application enters a first mode in response to the second operation; When the first preview window displays the first image, a third operation targeting the first object is detected; In response to the third operation, a second preview window is displayed.
5. The method according to claim 4, characterized in that, After acquiring the first audio data and before stopping video recording, the method further includes: The terminal determines the number N of stream configurations.
6. The method according to claim 5, characterized in that, The terminal determines the number N of stream configurations, specifically including: The terminal detects the second operation on the second control, and in response to the operation, the terminal determines the number of stream configurations N; After the recording interface is displayed, the operation targeting the first object is detected; In response to the operation on the first object, the terminal updates the number of stream configurations N.
7. The method according to claim 5, characterized in that, The terminal determines the number N of stream configurations, specifically including: The terminal detects the second operation on the second control, and in response to the operation, the terminal determines the number of stream configurations N; After the recording interface is displayed, an operation to stop recording the second video is detected; In response to the operation of stopping recording the second video, the terminal updates the number of stream configurations N.
8. The method according to claim 2, characterized in that, The terminal includes a first audio encoding module and a second audio encoding module, which encodes one channel of second audio data to obtain the first target audio data; and encodes another channel of second audio data to obtain the second target audio data, specifically including: The first audio encoding module encodes the second audio data based on one channel to obtain the first target audio data, while the second audio encoding module encodes the second audio data based on the other channel to obtain the second target audio data.
9. The method according to any one of claims 1-8, characterized in that, The first audio data includes the first audio signal, as well as the timestamp, data length, sampling rate, and data format of the first audio signal.
10. The method according to any one of claims 1-9, characterized in that, The second audio data includes the first audio signal, as well as the timestamp and data length of the first audio signal.
11. An electronic device, characterized in that, It includes one or more processors and one or more memories; wherein the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, the computer program code including computer instructions, which, when executed by the one or more processors, cause the method of any one of claims 1-10 to be performed.
12. A chip system, characterized in that, The chip system is applied to a terminal, and the chip system includes one or more processors, the processors being used to invoke computer instructions to cause the terminal to perform the method as described in any one of claims 1-10.
13. A computer program product containing instructions, characterized in that, When the instructions are executed on the electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-10.
14. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on an electronic device, they cause the method described in any one of claims 1-10 to be performed.
Citation Information
Patent Citations
Screen image recording method, terminal and computer readable storage medium
CN107277607A
Audio processing method and equipment
CN113365013A