A method, device, and storage medium for realizing real-time mixing of screen recording

By mixing and recording system sound and microphone sound in real time in the smart terminal of the Android system, the problem of low screen recording efficiency and out-of-synchronization of audio and video is solved, and the screen recording effect is realized.

CN115767170BActive Publication Date: 2025-08-05LANGYUAN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211393343.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-08
Publication Date
2025-08-05
Estimated Expiration
2042-11-08

AI Technical Summary

Technical Problem

In the process of recording screens, the recording system sound and microphone sound are stored separately, resulting in low efficiency, and the screen recording file cannot be generated instantly, and the audio and video are prone to problems that are not synchronized.

Method used

By creating multimedia processing tools in the Android system's smart terminal, mixing the system sound and microphone sound in real time, and outputting the mixing source directly to the video file, realizing recording and stopping.

Benefits of technology

It realizes the effect of recording and stopping, improves the stability and user experience of screen recording, and avoids the problem of waiting time and audio and video out of later synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115767170B_ABST
    Figure CN115767170B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, and storage medium for implementing real-time audio mixing of screen recording, which is applied to smart terminals of Android systems. The method includes the following steps: Step S1, creating a multimedia processing tool, and recording the video data of the smart terminal through the multimedia processing tool; Step S2, initializing the first sound source and the second sound source, and keeping the attributes of the first sound source and the second sound source consistent; Step S3, extracting the first sound source and the second sound source through a preset function, mixing the first sound source and the second sound source through the multimedia processing tool, obtaining a mixed sound source, and outputting the mixed sound source to a buffer; Step S4, upon receiving a call instruction, adding the mixed sound source located in the buffer to the video data, and outputting a screen recording file. The present application directly outputs the real-time mixing of media sound and microphone sound to a video file, achieving the effect of recording and stopping immediately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, and storage medium for implementing real-time audio mixing for screen recording. Background Art

[0002] When recording screens in educational or conference settings, it's important to capture both the entire device's multimedia audio and the speaker's speech. Typically, the upper-layer application simultaneously records system audio and microphone audio, mixes them, and saves them to an audio file. After recording is complete, the audio and video data are combined into a single screen recording file.

[0003] However, the above method is inefficient for synthesizing screen recording files. Screen recording files cannot be generated immediately upon stopping, requiring a waiting period. Synthesis time is proportional to the recording time, making this wait time lengthy in large or long meetings. Furthermore, because audio and video files are stored separately, audio and video synchronization issues can easily occur. Summary of the Invention

[0004] In order to solve the above problems, the embodiments of the present application provide a method, device, and storage medium for implementing real-time mixing of screen recording, which mixes media sound and microphone sound in real time and outputs them directly to a video file, achieving the effect of recording and stopping immediately.

[0005] To this end, one aspect of the present application provides a method for implementing real-time audio mixing of screen recording, which is applied to an Android smart terminal. The method includes the following steps:

[0006] Step S1: Create a multimedia processing tool and record video data of the smart terminal through the multimedia processing tool;

[0007] Step S2: Initialize the first sound source and the second sound source, and keep the attributes of the first sound source and the second sound source consistent;

[0008] Step S3: extracting the first sound source and the second sound source through a preset function, mixing the first sound source and the second sound source through the multimedia processing tool to obtain a mixed sound source, and outputting the mixed sound source to a buffer;

[0009] Step S4: When a call instruction is received, the audio mixing source in the buffer is added to the video data, and a screen recording file is output.

[0010] Optionally, in combination with any of the above aspects, in another implementation of this aspect, the calling instruction includes a first sound source calling instruction, a second sound source calling instruction, and a mixing source calling instruction; after receiving the first sound source calling instruction or the second sound source calling instruction, the first sound source or the second sound source is called, and the first sound source or the second sound source is mixed with the video data; when the mixing source calling instruction is received, the mixing source is obtained from the buffer through the mixing source calling function, and the mixing source is mixed with the video data.

[0011] Optionally, in combination with any of the above aspects, in another implementation of this aspect, mixing the first sound source and the second sound source through the multimedia processing tool includes mixing the first sound source and the second sound source, or mixing the first sound source and the second sound source based on the video data.

[0012] Optionally, in combination with any of the above aspects, in another implementation of this aspect, the attributes of the first sound source and the second sound source include audio parameters sampling rate, sample size, number of channels, and bit rate.

[0013] Optionally, in combination with any of the above aspects, in another implementation of this aspect, the smart terminal includes a first smart terminal and a second smart terminal, the second smart terminal is used to record the user interface of the first smart terminal and the first sound source and the second sound source; the first sound source is emitted by the first smart terminal.

[0014] Optionally, in combination with any of the above aspects, in another implementation of this aspect, the multimedia processing tool is based on the Android underlying framework.

[0015] Optionally, in combination with any of the above aspects, in another implementation of this aspect, step S3, extracting the first sound source and the second sound source through a preset function is extracting the first sound source and the second sound source through a read() function; the mixed sound source calling function is a start() function.

[0016] Optionally, in combination with any of the above aspects, in another implementation of this aspect, the first sound source is a microphone sound source, and the second sound source is a system media sound source.

[0017] In another aspect of the present application, an electronic device is provided, which includes a processor, a memory, and a computer program stored in the memory and runnable on the processor, and when the processor executes the computer program, it implements a method for implementing real-time screen recording and mixing as described above.

[0018] In another aspect of the present application, a storage medium is provided, on which a computer program is stored. When the computer program is executed, it implements any of the above-described methods for implementing real-time screen recording and mixing.

[0019] As described above, the present application provides a method, device, and storage medium for implementing real-time mixing of screen recording. By creating a multimedia processing tool, the system media sound and microphone sound are mixed in real time and directly output to the video file, achieving the effect of recording and stopping at the same time without the need for post-synthesis, greatly improving the stability of screen recording and user experience.

[0020] The above summary is provided to introduce some concepts in a simplified form, which are further described in detail in the detailed description below. The above summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. The claimed subject matter is not limited to implementations that solve any or all of the disadvantages identified in the background. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can also be obtained based on these drawings without paying creative labor. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application for those skilled in the art by reference to specific embodiments.

[0022] Figure 1 This is a flowchart of a method for implementing real-time mixing of screen recording provided by this application. DETAILED DESCRIPTION

[0023] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0024] It should be noted that, in this document, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, components, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined by their explanation in the specific embodiment or further combined with the context of the specific embodiment.

[0025] It should be understood that although the terms "first," "second," "third," etc. may be used herein to describe various information, such information should not be limited to these terms. These terms are used solely to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the term "if," as used herein, may be interpreted as "upon," "when," or "in response to a determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context indicates otherwise. It should be further understood that the terms "comprising" and "including" indicate the presence of the recited features, steps, operations, elements, components, items, types, and / or groups, but do not preclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, types, and / or groups. The terms "or," "and / or," "including at least one of the following," etc., as used herein, may be interpreted as inclusive, meaning any one or any combination. An exception to this definition occurs only when a combination of elements, functions, steps, or operations are inherently mutually exclusive in some manner.

[0026] It should be understood that, although the various steps in the flowchart in the embodiment of the present application are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence in the order indicated by the arrows. Unless clearly stated herein, the execution of these steps is not strictly limited in order, and they can be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and their execution order is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0027] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0028] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0029] See also Figure 1 This application provides a method for implementing real-time mixing of screen recording, which is applied to smart terminals such as conference terminals and mobile phones. It mixes media sound and microphone sound in real time and outputs it directly to the video file, achieving the effect of recording and stopping immediately.

[0030] This method comprises the following steps:

[0031] Step S1: Create a multimedia processing tool and use it to record a video from a smart terminal to obtain a video file. Create a multimedia processing tool in the smart terminal based on the Android underlying framework. The multimedia processing tool is a MixAudioSource, which provides audio source data for mixing. It provides an audio data source to an MPEG4Writer, enabling automatic mixing of audio sources. When the mixed audio source is subsequently needed, it can be directly provided through the multimedia processing tool.

[0032] Step S2: Initialize the first sound source and the second sound source, and keep the attributes of the first sound source and the second sound source consistent.

[0033] The first and second audio sources can originate from different smart terminals or the same smart terminal. During a conference, multiple participants are participating remotely via conference terminals, while at least one participant is also present at the smart terminal. During screen recording, both the system sound from the smart terminal and the sound received by the microphone need to be recorded. The methods for receiving and recording these two sounds are different.

[0034] Set the microphone audio source, mMicAudioSource, as the first audio source and the system media audio source, mSysAudioSource, as the second audio source. Initialize the received audio source and the second audio source in the multimedia processing tool to keep the properties of the first audio source and the second audio source consistent. The properties of the first audio source and the second audio source include audio parameter sampling rate, sample size, number of channels, and bit rate. By keeping the audio parameter sampling rate and sample size consistent, the synchronization of the two received audio sources can be directly synchronized and mixed later without readjustment.

[0035] Step S3: extracting the first sound source and the second sound source through a preset function, mixing the first sound source and the second sound source through the multimedia processing tool to obtain a mixed sound source, and outputting the mixed sound source to a buffer.

[0036] The first and second audio sources are read simultaneously by the read() function, and the first and second audio sources are mixed by a multimedia processing tool. The mixing of the first and second audio sources can be based on a video file or can be performed directly.

[0037] Step S4: When a call instruction is received, the audio mixing source in the buffer is added to the video file, and a screen recording file is output.

[0038] During the screen recording process, if there is a need to adjust the sound source, the user can select different call instructions at any time, and determine the sound source required for screen recording through the call instructions, so that the corresponding screen recording file can be quickly output. The call instructions include a first sound source call instruction, a second sound source call instruction, and a mixed sound source call instruction; after receiving the first sound source call instruction or the second sound source call instruction, the first sound source or the second sound source is called, and the first sound source or the second sound source is mixed with the video data; when the mixed sound source call instruction is received, the mixed sound source is obtained from the buffer through the mixed sound source call function, and the mixed sound source is mixed with the video data.

[0039] Add the audio source to MPEG4Writer (a tool for writing MP4 video streams provided by the native Android system). When the start() function is called, a separate thread is started to call the read() function in step S3 to extract the audio source. The video file and audio source are encapsulated by MPEG4Writer to output the screen recording file.

[0040] Furthermore, the smart terminal of the present application includes a first smart terminal and a second smart terminal, wherein the second smart terminal is used to record the user interface of the first smart terminal and the first sound source and the second sound source; the first sound source is emitted by the second smart terminal, and the first sound source is the microphone sound source received by the second smart terminal. The first smart terminal can be a video conferencing terminal or an electronic whiteboard, a slide, etc., and the second smart terminal is a smart terminal provided with this method. The interface of the first smart terminal is recorded by the second smart terminal, and the first sound source and the second sound source received by the second smart terminal are mixed to output the recorded screen image in real time. Through the recording of the second smart terminal, the mixing of the first sound source, the second sound source and the video data of other smart terminals is realized, thereby solving the problem of low efficiency in recording and mixing of video data of non-smart terminals.

[0041] This application provides a method for implementing real-time mixing of screen recording. By creating a multimedia processing tool, the system media sound and microphone sound are mixed in real time and directly output to a video file, achieving the effect of recording and stopping at the same time without the need for post-synthesis, greatly improving the stability of screen recording and user experience.

[0042] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0043] In this application, the same or similar terminology, technical solutions and / or application scenario descriptions are generally only described in detail the first time they appear. When they appear again later, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, for the same or similar terminology, technical solutions and / or application scenario descriptions that are not described in detail later, you can refer to the previous relevant detailed descriptions.

[0044] In this application, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0045] The various technical features of the technical solution of this application can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0046] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as above, and includes a number of instructions for enabling a terminal device (which can be an electrical device or a network device, etc.) to execute the method of each embodiment of the present application.

[0047] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for implementing real-time audio mixing for screen recording, characterized in that: Applied to an Android-based smart terminal, this method includes the following steps: Step S1: Create a multimedia processing tool and record video data of a smart terminal through the multimedia processing tool; the multimedia processing tool is based on the Android underlying framework; Step S2: Initializing the first sound source and the second sound source, and maintaining the properties of the first sound source and the second sound source consistent, wherein the properties of the first sound source and the second sound source include audio parameters sampling rate, sample size, number of channels, and bit rate; Step S3: extracting the first sound source and the second sound source through a preset function, mixing the first sound source and the second sound source through the multimedia processing tool to obtain a mixed sound source, and outputting the mixed sound source to a buffer; mixing the first sound source and the second sound source through the multimedia processing tool includes mixing the first sound source and the second sound source, or mixing the first sound source and the second sound source based on the video data; extracting the first sound source and the second sound source through the preset function is extracting the first sound source and the second sound source through a read() function; the mixed sound source calling function is a start() function; Step S4: upon receiving the call instruction, adding the audio mixing source in the buffer to the video data and outputting a screen recording file; The calling instruction includes a first sound source calling instruction, a second sound source calling instruction, and a mixed sound source calling instruction; after receiving the first sound source calling instruction or the second sound source calling instruction, the first sound source or the second sound source is called, and the first sound source or the second sound source is mixed with the video data; when the mixed sound source calling instruction is received, the mixed sound source is obtained from the buffer through the mixed sound source calling function, and the mixed sound source is mixed with the video data, so as to directly output the mixed sound source to the video file.

2. A method for implementing real-time audio mixing for screen recording according to claim 1, characterized in that: The smart terminal includes a first smart terminal and a second smart terminal. The second smart terminal is used to record the user interface of the first smart terminal and a first sound source and a second sound source. The first sound source is emitted by the first smart terminal.

3. A method for implementing real-time audio mixing for screen recording according to claim 1, characterized in that: The first sound source is a microphone sound source, and the second sound source is a system media sound source.

4. An electronic device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, it implements a method for implementing real-time screen recording and mixing as described in any one of claims 1 to 3.

5. A storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed, a method for implementing real-time screen recording and mixing as described in any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Screen recording method and apparatus, and computer-readable storage medium

    WO2021163879A1