Multimedia recording method and device and electronic equipment
By dynamically scheduling multiple media acquisition devices and synthesizing media streams during multimedia recording, the problems of fragmented and discontinuous recording files are solved, generating high-quality, structured recording files and improving the intelligence and efficiency of recording.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies suffer from problems such as fragmented and discontinuous recording files, as well as heavy post-processing burdens, due to limitations of a single device or manual switching of devices.
By dynamically scheduling multiple media acquisition devices based on changes in the acquisition scene information during the multimedia recording process, automatically dividing the acquisition stages, and synthesizing media streams from different devices, a continuous and complete recording file is generated.
It optimizes the entire recording process, generating high-quality, structured, and easily manageable recording files. This reduces the manual workload for users in file management and post-production compositing, and improves the intelligence and quality of the recording process.
Smart Images

Figure CN121750918A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and more particularly to a multimedia recording method, apparatus, and electronic device. Background Technology
[0002] With the increasing use of multimedia recording technology in scenarios such as remote collaboration and online meetings, users have placed higher demands on the completeness, quality, and post-processing efficiency of recorded content.
[0003] Current common recording solutions typically rely on a single, fixed audio / video capture device for continuous recording. While this approach is suitable for stable environments, it reveals significant shortcomings in dynamic and complex real-world scenarios. When the primary device experiences a decline in capture quality due to movement, obstruction, network fluctuations, or changes in the acoustic environment, manually switching to another available device often results in multiple independent and discrete recording file fragments. These scattered fragments disrupt the spatiotemporal continuity of the recorded events and impose a burden of subsequent tedious file alignment, merging, and organization on the user, creating a "file fragmentation" problem that ultimately impacts the overall management efficiency and application value of the recording results. Summary of the Invention
[0004] This application provides a multimedia recording method, apparatus, and electronic device to solve the technical problems of fragmented and discontinuous recorded files and heavy post-processing burden caused by the limitations of a single device or manual switching of devices in the prior art.
[0005] In a first aspect, this application provides a multimedia recording method, the method comprising: During the execution of a multimedia recording task, the media stream acquisition process is divided into multiple acquisition stages according to changes in the acquisition scene information, and at least one of the multiple media acquisition devices is scheduled to participate in the media stream acquisition of the multimedia recording task. Based on the multiple acquisition stages, the media streams from the scheduled media acquisition devices are synthesized to generate the recording file for the multimedia recording task.
[0006] In one possible implementation, scheduling at least one of a plurality of media acquisition devices to participate in the media stream acquisition of the multimedia recording task includes: In response to a media acquisition source adjustment event, the second media acquisition device indicated by the media acquisition source adjustment event is scheduled to participate in the media stream acquisition of the multimedia recording task, and the media stream from the second media acquisition device is integrated into the media stream space maintained for the multimedia recording task according to the acquisition time; The process of synthesizing the media streams from the scheduled media acquisition devices based on the multiple acquisition stages to generate the recording file for the multimedia recording task includes: Based on the multiple acquisition stages, the media streams integrated in the media stream space are uniformly synthesized to generate the recording file for the multimedia recording task.
[0007] In one possible implementation, scheduling the second media acquisition device indicated by the media acquisition source adjustment event to participate in the media stream acquisition of the multimedia recording task includes: Switch the media acquisition device currently participating in the media stream acquisition from the first media acquisition device to the second media acquisition device indicated by the media acquisition source adjustment event; or, The first media acquisition device currently participating in the media stream acquisition and the second media acquisition device indicated by the media acquisition source adjustment event are scheduled to jointly participate in the media stream acquisition of the multimedia recording task.
[0008] In one possible implementation, scheduling the second media acquisition device indicated by the media acquisition source adjustment event to participate in the media stream acquisition of the multimedia recording task includes: The access information used to access the media stream space is shared to the second media acquisition device; In response to the joining request initiated by the second media acquisition device based on the access information, the second media acquisition device is scheduled to participate in the media stream acquisition of the multimedia recording task.
[0009] In one possible implementation, the method further includes: Establish a media stream transmission channel between the second media acquisition device and the selected device; The media stream acquired by the second media acquisition device is transmitted to the selected device through the media stream transmission channel, so as to display the text content obtained by speech-to-text transcription of the media stream acquired by the second media acquisition device on the selected device.
[0010] In one possible implementation, the method further includes: During the process of obtaining text content through speech-to-text transcription, in response to the user's adjustment of the selected text in the text, the adjusted text is used as the unified identification content, and in subsequent speech-to-text transcription processes, the unified identification content is used for speech-to-text transcription.
[0011] In one possible implementation, the method further includes: In response to triggering a pause command on the first media acquisition device, a control command is sent to the second media acquisition device to cause the second media acquisition device to pause media stream acquisition.
[0012] In one possible implementation, the process of combining the media stream from the scheduled media acquisition device to generate the recording file for the multimedia recording task includes: Time stamp alignment is performed on media streams from different scheduled media acquisition devices; The recording file for the multimedia recording task is generated based on the media stream that has been time-stamped and aligned.
[0013] In one possible implementation, the media stream acquisition process is divided into multiple acquisition stages based on changes in the acquisition scenario information, including: Obtain the geographic location information of the media acquisition devices currently participating in the media stream acquisition; Based on the changes in the geographical location information, the media stream acquisition process is divided into stages, generating a new acquisition stage.
[0014] In one possible implementation, the method further includes: The recorded file is subjected to content analysis and processing; Based on the results of the content analysis and processing, a structured summary is generated.
[0015] In one possible implementation, the plurality of media acquisition devices include: a virtual media source provided by a designated online conferencing application running on an electronic device.
[0016] Secondly, this application provides a multimedia recording apparatus, the apparatus comprising: The intelligent segmentation module is used to divide the media stream acquisition process into multiple acquisition stages based on changes in the acquisition scene information during the execution of a multimedia recording task. The device scheduling module is used to schedule at least one of multiple media acquisition devices to participate in the media stream acquisition of the multimedia recording task; The integration module is used to synthesize the media streams from the scheduled media acquisition devices based on the multiple acquisition stages, and generate the recording file for the multimedia recording task.
[0017] In one possible implementation, the device scheduling module is specifically used for: In response to a media acquisition source adjustment event, the second media acquisition device indicated by the media acquisition source adjustment event is scheduled to participate in the media stream acquisition of the multimedia recording task, and the media stream from the second media acquisition device is integrated into the media stream space maintained for the multimedia recording task according to the acquisition time; The integration module is specifically used for: Based on the multiple acquisition stages, the media streams integrated in the media stream space are uniformly synthesized to generate the recording file for the multimedia recording task.
[0018] In one possible implementation, the device scheduling module schedules the second media acquisition device indicated by the media acquisition source adjustment event to participate in the media stream acquisition of the multimedia recording task, including: Switch the media acquisition device currently participating in the media stream acquisition from the first media acquisition device to the second media acquisition device indicated by the media acquisition source adjustment event; or, The first media acquisition device currently participating in the media stream acquisition and the second media acquisition device indicated by the media acquisition source adjustment event are scheduled to jointly participate in the media stream acquisition of the multimedia recording task.
[0019] In one possible implementation, the device scheduling module schedules the second media acquisition device indicated by the media acquisition source adjustment event to participate in the media stream acquisition of the multimedia recording task, including: The access information used to access the media stream space is shared to the second media acquisition device; In response to the joining request initiated by the second media acquisition device based on the access information, the second media acquisition device is scheduled to participate in the media stream acquisition of the multimedia recording task.
[0020] In one possible implementation, the device further includes: The media streaming module is used to establish a media streaming channel between the second media acquisition device and the selected device; The media stream acquired by the second media acquisition device is transmitted to the selected device through the media stream transmission channel, so as to display the text content obtained by speech-to-text transcription of the media stream acquired by the second media acquisition device on the selected device.
[0021] In one possible implementation, the device further includes: The adaptive adjustment module is used to respond to the user's adjustment of selected text in the text during the process of obtaining text content through speech transcription, and to use the adjusted text as the unified identification content. In subsequent speech transcription processes, the unified identification content is used for speech transcription.
[0022] In one possible implementation, the device further includes: The acquisition control module is used to send a control command to the second media acquisition device in response to a pause command triggered on the first media acquisition device, so as to cause the second media acquisition device to pause media stream acquisition.
[0023] In one possible implementation, the integration module includes: The time alignment unit is used to perform time stamp alignment processing on media streams from different scheduled media acquisition devices; The generation unit is used to generate the recording file for the multimedia recording task based on the media stream that has been time-stamped and aligned.
[0024] In one possible implementation, the intelligent sharding module is specifically used for: Obtain the geographic location information of the media acquisition devices currently participating in the media stream acquisition; Based on the changes in the geographical location information, the media stream acquisition process is divided into stages, generating a new acquisition stage.
[0025] In one possible implementation, the device further includes: The minutes generation module is used to perform content analysis and processing on the recorded files; Based on the results of the content analysis and processing, a structured summary is generated.
[0026] In one possible implementation, the plurality of media acquisition devices include: a virtual media source provided by a designated online conferencing application running on an electronic device.
[0027] Thirdly, this application provides an electronic device, including: a processor and a memory, wherein the processor is configured to execute a multimedia recording program stored in the memory to implement the multimedia recording method described in any one of the first aspects.
[0028] Fourthly, this application provides a storage medium storing one or more programs that can be executed by one or more processors to implement the multimedia recording method described in any one aspect.
[0029] Compared with the prior art, the technical solution provided in this application has the following advantages: First, the method provided in this application automatically divides the acquisition stage by responding to changes in the acquisition scene information, transforming the originally lengthy and unstructured continuous recording process into a series of logically independent recording segments. This provides a structural foundation for intelligent content organization and accurate retrieval in the later stages, solving the problem of difficulty in reviewing and organizing traditional recordings due to the lack of automatic segmentation. Second, multiple media acquisition devices are dynamically scheduled to participate in the recording process, enabling the system to flexibly select or combine the optimal audio and video sources according to real-time conditions (such as sound quality and device availability). This overcomes the limitations of the sound reception quality of a single fixed device in mobile scenarios or complex acoustic environments, ensuring the quality and continuity of the media stream from the source. Finally, based on the divided acquisition stages, media streams from different acquisition time periods and different media acquisition devices are synthesized, aligning, integrating, and encapsulating the scattered and heterogeneous multiple media streams on the timeline into a unified recording file. This fundamentally solves the "file fragmentation" problem caused by multi-device relay recording, ensuring the integrity and coherence of the recording results.
[0030] In summary, the technical solution provided in this application deconstructs the multimedia recording process into three core components: "dynamic stage division based on changes in the acquisition scenario," "coordinated scheduling and acquisition of multiple devices," and "unified media stream synthesis across stages." This optimizes the entire recording task process from acquisition and organization to generation, automates the scheduling, alignment, and encapsulation of media from multiple devices, and ultimately produces a high-quality, structured, and easily manageable complete recording file. This significantly reduces the manual workload for users in file management and post-production synthesis, and improves the intelligence, content quality, and efficiency of multimedia recording. Attached Figure Description
[0031] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0032] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0033] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0034] Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this application; Figure 2 A flowchart illustrating an embodiment of a multimedia recording method provided in this application; Figure 3 A flowchart illustrating an embodiment of another multimedia recording method provided in this application; Figure 4 This is an example diagram of an application interface involved in an embodiment of this application; Figure 5 This is an example diagram of another application interface involved in the embodiments of this application; Figure 6 A block diagram illustrating an embodiment of a multimedia recording device provided in this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0036] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0037] To address the technical problems of fragmented and discontinuous recorded files and heavy post-processing burdens caused by limitations of single devices or manual device switching in existing technologies, this application provides a multimedia recording method, apparatus, and electronic device. By deconstructing the multimedia recording process into three core links—"dynamic stage division based on changes in the acquisition scenario," "cooperative scheduling and acquisition of multiple devices," and "unified media stream synthesis across stages"—the entire recording task is optimized from acquisition and organization to generation. It automates the scheduling, alignment, and encapsulation of media from multiple devices, ultimately producing a high-quality, structured, and easily manageable complete recording file. This significantly reduces the manual operation burden on users in file management and post-processing, and improves the intelligence, content quality, and efficiency of multimedia recording.
[0038] Figure 1 This is a schematic diagram of an application scenario involved in an embodiment of this application. Figure 1 The scenario shown is a meeting. In this scenario, to fully record the meeting content, participants can use electronic devices (such as tablets, laptops, smartphones, etc.) to record the meeting. Recording can be in the form of pure audio or audio-visual recording. In traditional recording solutions, the media acquisition device (i.e., the aforementioned electronic device) is usually fixed. When the user moves the device to a location far from the speaker, the device's audio reception is poor, or other problems occur, it is often necessary to manually stop the current recording and restart it on a new device. For example, in... Figure 1 In the scenario shown, participant Xiao Wang first used his smartphone to record the meeting. Later, because he needed to answer a phone call, the recording task was transferred to participant Xiao Zhang using his smartphone. As a result, the meeting recording files were scattered across different devices, making it difficult to organize and unify the content later.
[0039] exist Figure 1 In the example application scenario, if the technical solution provided in the embodiments of this application is applied, it is possible to dynamically schedule multiple media acquisition devices during the recording process, and automatically divide the acquisition stage according to the changes in the acquisition scene information (such as switching acquisition devices, changing device positions, changing sound environment, etc.). Finally, the media streams that may come from different media acquisition devices in multiple acquisition stages are intelligently synthesized to generate a continuous and complete recording file, thereby effectively solving problems such as recording interruption, file dispersion, and tedious post-production content unification.
[0040] also, Figure 1The examples provided in this application are merely illustrative application scenarios of the technical solutions. In practical applications, this solution can also be applied to various other scenarios, such as: course recording and note generation in the education and training industry, scenarios requiring mobile recording such as interviews / research, intelligent assistants in recruitment and interview scenarios, and scenarios requiring precise recording of dialogues such as legal / medical scenarios. These scenarios share the commonality of involving continuous and complete multimedia recording in a dynamic environment, thus all of which can utilize the methods provided in this application to achieve multi-device collaboration, intelligent segmentation, and unified synthesis, improving recording continuity and post-processing efficiency. This application does not limit the specific application scenarios.
[0041] Figure 2 This is a flowchart illustrating an embodiment of a multimedia recording method provided in this application. Figure 2 As shown, the method includes the following steps: Step 201: During the execution of the multimedia recording task, the media stream acquisition process of the multimedia recording task is divided into multiple acquisition stages according to the changes in the acquisition scene information, and at least one of the multiple media acquisition devices is scheduled to participate in the media stream acquisition of the multimedia recording task.
[0042] Among them, multimedia recording tasks are, for example, Figure 1 The meeting recording task in the scenario shown.
[0043] Media acquisition devices include various electronic devices that can be used to participate in multimedia recording tasks, such as smartphones, tablets, laptops, conference room audio systems, or virtual media sources provided by online conferencing applications. This application does not limit the specific form of the media acquisition device.
[0044] The technical solution provided in this application embodiment monitors one or more "collection scene information" in real time during the execution of a multimedia recording task. When it is determined that the "collection scene information" has undergone a "change" that meets preset conditions, it automatically triggers the logical segmentation of the current continuous recording process and starts a new "collection stage". Each "collection stage" corresponds to a logically independent recording segment, providing a structural basis for subsequent intelligent processing and content organization.
[0045] In one embodiment, the "change in acquisition scenario information" includes switching media acquisition devices. For example, switching the media acquisition device currently participating in media stream acquisition from a first media acquisition device (e.g., a user's mobile phone) to a second media acquisition device (e.g., a conference room audio system). The media acquisition device switching event signifies a fundamental change in the acquisition scenario, at which point the division of the acquisition phase can be automatically triggered, initiating a new "acquisition phase".
[0046] The media streams acquired in each acquisition stage are logically independent and have corresponding metadata, including but not limited to: stage number (in chronological order), identification of the device used, start and end time markers, geographical location information, and storage path of the media stream in that stage.
[0047] Simultaneously, during the execution of the multimedia recording task, at least one of multiple media acquisition devices is scheduled to participate in the media stream acquisition. Here, "scheduling" refers to the execution entity of this application embodiment acting as a scheduling center to coordinate and control the one or more media acquisition devices to acquire audio, video, or a mixed audio-video "media stream." It should be noted that this scheduling mechanism is a continuous action throughout the multimedia recording task execution process, working in conjunction with the acquisition phase division mechanism to jointly ensure the continuity and high quality of the recording task. Specifically: In one aspect, the scheduling mechanism acts as the triggering and execution engine for the division of acquisition phases: the scheduling operation itself is a crucial "acquisition scenario information" that triggers the division of acquisition phases. When "scheduling" causes a change in the media acquisition devices participating in the acquisition (i.e., "device switching"), regardless of whether the switching occurs between devices of the same or different types, this device switching event is recognized as a fundamental change in the acquisition scenario, thus automatically ending the current acquisition phase and starting a new one. Subsequently, the scheduling mechanism will identify and start at least one media acquisition device for this new acquisition phase. Therefore, each switching of media acquisition devices naturally defines an independent acquisition phase.
[0048] In another aspect, the scheduling mechanism facilitates continuous optimization within the acquisition phase: after an acquisition phase is initiated and the dominant media acquisition device is determined, the scheduling center can still perform optimized scheduling based on audio quality, network conditions, or user commands during the duration of that phase, without triggering a switch to the dominant media acquisition device. For example, this could involve parameter tuning of the current device or adding an auxiliary acquisition device within the current acquisition phase (whose media stream is mixed with the dominant device's stream but does not replace its dominant position). This type of scheduling aims to finely improve the recording quality of the current segment without inducing a phase re-division.
[0049] Furthermore, to address the issues of media stream fragmentation and post-processing difficulties caused by "media acquisition source adjustment events" (which, in the scheduling mechanism, can be considered a unified "media acquisition source adjustment event," whether it's triggering a "device switch" during the acquisition phase or adding equipment for phase optimization), this embodiment creates and maintains a unified recording space for multimedia recording tasks. This recording space is a logical environment created and maintained for multimedia recording tasks, used to uniformly host and manage all media streams and metadata. Its physical implementation can be one or more combinations of local storage space, cloud storage space, or distributed storage space.
[0050] Accordingly, in one embodiment, scheduling at least one of a plurality of media acquisition devices to participate in media stream acquisition for a multimedia recording task includes: in response to a media acquisition source adjustment event, scheduling a second media acquisition device indicated by the media acquisition source adjustment event to participate in media stream acquisition for a multimedia recording task, and integrating the media stream from the second media acquisition device into a media stream space maintained for the multimedia recording task according to the acquisition time.
[0051] In this context, integrating the media stream from the second media acquisition device into the media stream space maintained for the multimedia recording task according to the acquisition time means: based on the time stamp information carried or associated with the media stream, aligning and merging it with the existing media streams in the recording space on a unified time axis, ensuring that the multi-source media streams are continuously and consistently recorded in the recording space in the actual time sequence, thereby logically forming a coherent and complete recording content, laying the foundation for the subsequent generation of a unified recording file.
[0052] In addition, in practical applications, "changes in the scene information being collected" can also include, but are not limited to, the following situations: Changes in the geographical location of the acquisition device: For example, if a user moves the media stream acquisition device from "the conference room on the 3rd floor of Company Building A" to "the exhibition hall on the 1st floor of Company Building A," and the distance moved exceeds a set threshold (e.g., 50 meters), logical segmentation of the current continuous recording process is automatically triggered to distinguish the recording content corresponding to different physical spaces. Accordingly, in one embodiment, based on changes in the acquisition scenario information, the media stream acquisition process is divided into multiple acquisition stages, including: obtaining the geographical location information of the media acquisition device currently participating in the media stream acquisition; and triggering the stage division of the media stream acquisition process based on changes in the geographical location information, generating a new acquisition stage.
[0053] Agenda or topic transition: For example, based on a preset meeting agenda, a new data collection phase can be automatically triggered at 10:30 when the "Project Discussion (10:00-10:30)" session ends, thus distinguishing it from the subsequent "Break" session. Another example is the automatic triggering of a new data collection phase when predefined key phrases such as "Next, we will discuss the next topic" or "End of Chapter One" are recognized through real-time speech transcription.
[0054] Significant changes in acoustic scene: By analyzing environmental acoustic characteristics (such as background noise, number of voices, and reverberation), abrupt changes in the scene can be identified. For example, when a scene changes from "intense debate among multiple people" to "single-person presentation," or from "indoor meeting environment" to "outdoor meeting environment," a new data acquisition phase can be automatically triggered.
[0055] Continuous silence detection: When continuous silence is detected for more than a preset time (such as 2 minutes), it may mean that the topic has ended or the meeting has been paused. At this time, a new data collection phase can be automatically triggered.
[0056] User-initiated interactive marking: Users can manually add segment markers by clicking the marking button during the recording process. In response to user-initiated interactive marking, a new collection stage will be automatically triggered.
[0057] Therefore, the technical solution provided in this application, by supporting the aforementioned diverse scenario change triggering conditions, can flexibly adapt to various complex recording scenarios, from formal meetings and mobile presentations to offline seminars. Whether it's movement in physical space, the progression of the meeting agenda, switching of the acoustic environment, or silent intervals and user-initiated markings, all can be automatically identified and transformed into structured acquisition stages. This achieves precise alignment between the recording process and the dynamics of the real scene, ensuring that the generated content maintains both continuity and clear logical segmentation, significantly improving the organization of the recording results and the efficiency of post-processing.
[0058] The specific method for scheduling at least one of the multiple media acquisition devices to participate in the multimedia recording task will be explained in detail below through specific embodiments, and will not be elaborated here.
[0059] Step 202: Based on multiple acquisition stages, the media streams from the scheduled media acquisition devices are synthesized to generate recording files for the multimedia recording task.
[0060] In step 202, based on the "acquisition phase" defined in step 201, the original "media streams" from potentially different media acquisition devices in multiple acquisition phases are integrated and post-processed to generate recording files for the multimedia recording task.
[0061] In one embodiment, the process of synthesizing media streams from scheduled media acquisition devices to generate recording files for a multimedia recording task includes: performing time-stamp alignment processing on media streams from different scheduled media acquisition devices; and generating recording files for a multimedia recording task based on the time-stamp aligned media streams.
[0062] This embodiment aims to solve the problem of media stream asynchrony caused by differences in device clocks, different acquisition start times, or network transmission delays in multi-device collaborative recording. Specifically, by identifying and correcting the embedded or associated time stamp information of each media stream, all media streams are unified onto a single reference timeline, achieving precise timing alignment. After alignment, depending on the recording requirements and scenario, overlapping or continuous media stream segments on the timeline can be further processed using methods such as mixing (for overlapping parts) or seamless splicing (for continuous parts), laying the foundation for generating a high-quality, coherent final recording file.
[0063] Based on this, and using media streams that have undergone time-stamped alignment, the recording files for multimedia recording tasks are generated, including but not limited to the following technical processes: Stream synthesis and encoding: This involves applying optimization processing such as mixing, noise reduction, and gain equalization to multiple audio and video streams that are within the same acquisition stage and have completed time synchronization, and compressing them according to a preset encoding format (such as H.264) to generate high-quality independent sub-media files corresponding to that acquisition stage. Then, through file integration and encapsulation, all sub-media files generated in all acquisition stages are seamlessly connected and uniformly encapsulated according to chronological order, ultimately outputting a complete, continuous, and clearly structured recording file. This file not only contains coherent media content but can also embed or associate metadata from each acquisition stage, such as start and end times, and triggered scene information (such as device switching, geographical location movement, keyword recognition, etc.), thereby solidifying the structured information of intelligent segmentation at the file level. This complete synthesis process realizes the end-to-end processing of multi-source media streams from synchronization and optimization to structured encapsulation. This allows the technical advantages brought by dynamic scheduling and intelligent segmentation to ultimately translate into high-quality recording results that users can directly use and that are easy to locate, retrieve, and manage later.
[0064] In another embodiment, based on multiple acquisition stages, media streams from the scheduled media acquisition devices are synthesized to generate recording files for a multimedia recording task, including: based on multiple acquisition stages, performing unified synthesis processing on media streams integrated in the media stream space to generate recording files for a multimedia recording task.
[0065] The purpose of the "uniform compositing process" here is to transform heterogeneous media streams from different acquisition stages and different media acquisition devices into a coherent file that maintains a high degree of consistency in timing, format, and semantic content. This process includes not only basic timestamp alignment, encoding format unification, and stream encapsulation, but more importantly, the introduction of a context-based intelligent correction and fusion mechanism.
[0066] Specifically, to address semantic noise introduced during actual recording due to differences in equipment, recording environment, or speaker's dialect accent, the system utilizes complete contextual information integrated in the media stream space (including preceding and following speech content, historical pronunciation characteristics of the same speaker, background of the topic at each stage, etc.) to perform in-depth analysis and intelligent correction on the initial text generated by speech transcription.
[0067] For example, for the same technical term, device A might be identified as "model" due to distance and ambient noise, while device B might be identified as "module" due to clear sound quality. Traditional solutions would retain this inconsistency. However, in the embodiments of this application, the "uniformity synthesis processing" analyzes the context in which the term appears (e.g., discussing "machine learning model training") and, combined with the speaker's pronunciation habits during this stage, automatically corrects "module" to "model," ensuring that the reference to the same entity in the entire recording file remains consistent and accurate.
[0068] This process effectively corrects recognition biases caused by differences in device performance or accents by fusing recognition results from multiple devices and applying acoustic and language models. As a result, the final recorded file and its associated text summary achieve semantic coherence and consistent word choice.
[0069] The technical solution provided in this application firstly automatically divides the recording process into stages by responding to changes in the scene information. This transforms the originally lengthy and unstructured continuous recording process into a series of logically independent recording segments. This provides a structural foundation for intelligent content organization and accurate retrieval later, solving the problem of difficulty in reviewing and organizing traditional recordings due to the lack of automatic segmentation. Secondly, it dynamically schedules multiple media acquisition devices during the recording process, enabling the system to flexibly select or combine the optimal audio and video sources based on real-time conditions (such as sound quality and device availability). This overcomes the limitations of single fixed devices in mobile scenarios or complex acoustic environments, ensuring the quality and continuity of the media stream from the source. Finally, based on the divided multiple recording stages, the media streams from different recording times and different media acquisition devices are synthesized. The scattered and heterogeneous multiple media streams are aligned, integrated, and packaged into a unified recording file on the timeline. This fundamentally solves the "file fragmentation" problem caused by relay recording from multiple devices, ensuring the integrity and coherence of the recording results.
[0070] In summary, the technical solution provided in this application deconstructs the multimedia recording process into three core components: "dynamic stage division based on changes in the acquisition scenario," "coordinated scheduling and acquisition of multiple devices," and "unified media stream synthesis across stages." This optimizes the entire recording task process from acquisition and organization to generation, automates the scheduling, alignment, and encapsulation of media from multiple devices, and ultimately produces a high-quality, structured, and easily manageable complete recording file. This significantly reduces the manual workload for users in file management and post-production synthesis, and improves the intelligence, content quality, and efficiency of multimedia recording.
[0071] The above provides an overall explanation of the multimedia recording method provided in the embodiments of this application. The following describes an exemplary implementation of the above-mentioned "scheduling mechanism" through specific embodiments.
[0072] See Figure 3 This is a flowchart of an embodiment of another multimedia recording method provided in this application. Figure 3 The process shown is in Figure 2 Based on the illustrated process, the following steps are included: Step 301: When the multimedia recording task starts, schedule the first media acquisition device to participate in the media stream acquisition of the multimedia recording task.
[0073] The first media acquisition device refers to the media acquisition device that is enabled by default or preferred when recording begins.
[0074] In an exemplary application scenario, when a user launches a recording application on an electronic device (such as a smartphone, tablet, or laptop) and clicks the "Start Recording" button on the application interface, the execution entity of this application embodiment defaults to using the electronic device as the first media acquisition device and schedules the first media acquisition device to immediately start acquiring media streams.
[0075] Another exemplary application scenario involves recording third-party online conferencing applications. In this scenario, users can start the recording task in two ways: one is automatic detection triggering, where the user enables the "automatic meeting detection" function in the relevant settings of the online conferencing application, and a recording prompt will automatically pop up on the interface when a meeting is detected; the other is user-initiated triggering, where the user actively selects to record the online meeting process in the online conferencing application without automatic prompting. Regardless of the method, once the user confirms "start recording," the execution entity of this application embodiment will initiate the cloud recording process for the online meeting. At this time, the first media acquisition device is the virtual audio source provided by the online conferencing application.
[0076] Step 302: During the execution of the multimedia recording task, in response to the media acquisition source adjustment event, the access information used to access the media stream space is shared to the second media acquisition device.
[0077] Step 303: In response to the join request initiated by the second media acquisition device based on the access information, schedule the second media acquisition device to participate in the media stream acquisition of the multimedia recording task.
[0078] Steps 302 and 303 together describe the dynamic adjustment and device scheduling process of the media acquisition source during the recording task. Step 302 corresponds to the "Media Acquisition Source Adjustment Event Triggering and Access Information Sharing" stage, which involves sharing the access information required to access the current media stream recording space with the target device after detecting a media acquisition source adjustment event. Step 303 corresponds to the "Device Joining and Acquisition Scheduling" stage, which is the process of allocating and scheduling the media acquisition role of the target device after it initiates a join request based on the access information.
[0079] Specifically, the "media source adjustment incidents" here can be divided into two types: One method is device switching: this involves switching the media capture device currently participating in media stream acquisition from the first media capture device to the second media capture device indicated by the media capture source adjustment event. In this scenario, the second media capture device takes over the multimedia recording task, and the first media capture device stops capturing.
[0080] Another approach is multi-device joint acquisition: that is, scheduling the first media acquisition device currently participating in media stream acquisition and the second media acquisition device indicated by the media acquisition source adjustment event to jointly participate in the media stream acquisition of the multimedia recording task.
[0081] For adjustment events such as "device switching", they can be triggered in the following three ways: The first triggering method: responding to the selection of a media acquisition source through the interface of the first media acquisition device. For details, see... Figure 4 For example, a user can click the "Device Switch" button 41 on the recording interface of the first media acquisition device, and then select a second media acquisition device (such as another mobile device or conference room device) from the pop-up list. At this time, the execution subject of this application embodiment will execute step 302, generating and sharing access information according to the selection operation. For mobile devices, in response to the access information, the user is prompted to launch the same account recording application on the second media acquisition device and select "Synchronous Recording"; for conference room devices, a screen projection code input window pops up for the user to manually enter. Subsequently, when the second media acquisition device initiates a join request based on this information, the execution subject of this application embodiment will execute step 303: scheduling the second media acquisition device to take over from the first media acquisition device as the media acquisition source.
[0082] The second triggering method is in response to a connection request for a multimedia recording task initiated by the second media acquisition device. Specifically, this method is an implementation where the second media acquisition device actively initiates joining the multimedia recording task. For example, the second media acquisition device has the same recording application installed as the first media acquisition device and is logged into the same account. The user can trigger the "Discover Recording Task" button on the application interface of the second media acquisition device. Since the first and second media acquisition devices are logged into the same account on their respective recording applications, the second media acquisition device can detect the multimedia recording task currently being participated in by the first media acquisition device. Next, the user can select the multimedia recording task on the application interface of the second media acquisition device and click "Synchronize Recording." At this time, the execution entity of this embodiment executes step 302: automatically completes access information synchronization based on account trust relationship, and then executes step 303: schedules the second media acquisition device to participate in the acquisition.
[0083] The aforementioned automatic account discovery mechanism is merely an exemplary implementation. In practical applications, this application embodiment supports access via other information media. For example, a user can manually enter the task ID of a multimedia recording task generated and shared by the first media acquisition device on the interface of the second media acquisition device; or, by scanning a QR code containing task access information displayed on the interface of the first media acquisition device, a joining request can be quickly initiated. This application embodiment does not impose any limitations on this.
[0084] The third triggering method: This is triggered when the media stream acquisition quality of the first media acquisition device is detected to be unsatisfactory, failing to meet the set quality conditions. Specifically, when the execution entity of this embodiment detects an anomaly or quality degradation in the current first media acquisition device (such as an external microphone), a switching process can be automatically triggered. First, step 302 is executed: the access information is directed to a system-recommended or preset backup device (such as a mobile device for an account). After the backup device responds and joins, step 303 is executed: it is scheduled as the new media acquisition device, achieving automatic disaster recovery in abnormal situations.
[0085] In the technical solution provided in this application embodiment, regardless of the method by which device switching is triggered, a seamless background scheduling mechanism ensures the continuity of media stream acquisition and recording tasks. For the user, the entire switching process is transparent and imperceptible. Specifically, when the user actively selects to switch through the interface of the first media acquisition device, the first media acquisition device continues to work throughout the entire process of responding to the operation, sharing access information, and adding the new device to the schedule. Only after the new device (the second media acquisition device) is ready and begins acquisition does the system smoothly hand over the acquisition task, thus achieving a seamless hot switch. In the scenario where the second media acquisition device actively initiates a connection, the recording task runs continuously as an independent service. The joining request of the second media acquisition device and the system authorization process are both completed in the background, without affecting the recording of existing media streams, achieving parallel recording and expansion. In the disaster recovery scenario where the system automatically triggers switching based on quality monitoring, anomaly detection and backup device scheduling both run as background guardian services. When the system determines that the quality of the primary device (the first media acquisition device) is substandard, it automatically activates the backup acquisition link. The entire process is fast and discreet, effectively avoiding recording interruptions or quality degradation caused by device failure.
[0086] This demonstrates that regardless of the device switching method, all media source adjustments are completed while the recording task is running continuously, ensuring the continuity of the user experience and the integrity of the recorded content.
[0087] For adjustment events such as "joint device acquisition", the triggering method mainly relies on the first or second implementation mentioned above (usually switching to a media acquisition device or joining as a collaborating end based on the user's choice). The calling logic of steps 302 and 303 is consistent with the above description. The only difference is that the scheduling strategy of step 303 is "joint participation" rather than "complete switching", which will not be elaborated here.
[0088] Furthermore, in the process of achieving multi-device collaboration, embodiments of this application also establish an instruction control mechanism. In one embodiment, in response to a pause instruction triggered on the first media acquisition device (e.g., by a user clicking...), Figure 4 The "End Recording" button on the interface triggers a pause command, which sends a synchronous control command to the second media acquisition device currently participating in the recording, causing the second media acquisition device to also pause its media stream acquisition. This ensures consistent operational status across multiple devices and global controllability of the recording task.
[0089] Step 304: Establish a media stream transmission channel between the second media acquisition device and the selected device. Transmit the media stream acquired by the second media acquisition device to the selected device through the media stream transmission channel so as to display the text content obtained by speech transcription of the media stream acquired by the second media acquisition device on the selected device.
[0090] In step 304, after the second media acquisition device is successfully scheduled, the execution entity of the embodiment of the present application establishes a stable real-time media stream transmission channel between itself and the "selected device" (usually the device for the user's main operation and viewing, such as the first media acquisition device). Exemplarily, for conference room devices, this channel can be established by inputting and verifying a "screen mirroring code"; for devices with the same account, a secure connection can be established through cloud services and the account system.
[0091] The media stream collected by the second media acquisition device is transmitted in real time to the "selected device" through the above dedicated transmission channel. This enables the media stream to be reliably converged to the terminal for centralized processing regardless of the location of the physical acquisition device.
[0092] After receiving the media stream from the second media acquisition device, the "selected device" can call the backend speech-to-text service to process the media stream in real time, and dynamically display the text content obtained by the transcription on its own "recording interface". In one embodiment, during the transcription process, voiceprint analysis and speaker identification can be performed synchronously to distinguish audio segments of different speakers, and the identified speaker identity information can be associated with the corresponding transcribed text. Subsequently, the text content obtained by the transcription, together with the speaker identifier (such as name, role, or representative avatar), is dynamically displayed on its own "recording interface" (for example, see Figure 5 the example), thus realizing speaker differentiation and structured presentation of meeting records.
[0093] The above processing constitutes a complete closed loop centered on speech-to-text, with the main processing end as the hub and multi-terminal collaborative interaction.
[0094] Further, in one embodiment, during the process of obtaining text content by speech-to-text, in response to the user's adjustment of the selected text in the text, the adjusted text is used as the unified recognition content, and in the subsequent speech-to-text process, the unified recognition content is used for speech-to-text.
[0095] During the speech-to-text process, when the user manually corrects the real-time generated text content (for example, changing the misrecognized "Zhang" to "Zhang"), the execution entity of the embodiment of the present application will start a set of unified recognition mechanisms. This mechanism regards the user's correction behavior as a clear context feedback, automatically associates the corrected vocabulary with its audio segment and context features for correlation analysis, and forms a unified recognition record with context attributes. In subsequent speech-to-text, such records will be actively called, and when the same or similar context is recognized, the表述 corrected by the user (such as "Zhang") will be preferentially adopted, thus achieving the continuous optimization effect of "one correction, subsequent automatic alignment".
[0096] This mechanism not only significantly reduces the workload of users repeatedly correcting the same type of errors, but also enables the speech-to-text service to dynamically adapt to users' language habits, professional terminology and pronunciation characteristics, gradually shifting from general recognition to personalized and accurate recognition, ensuring the continuity of transcription while continuously improving the accuracy and consistency of content generation.
[0097] Figure 3 The process shown is in Figure 2 Based on the process shown, a flexible device scheduling mechanism was used to achieve seamless switching and collaborative work between different media acquisition devices during the recording process, ensuring the continuity and stability of the task in complex multi-device scenarios.
[0098] Finally, this application also provides the following embodiments: performing content analysis processing on the recorded file; generating a structured summary based on the results of the content analysis processing.
[0099] After recording is completed and the final recording file is generated, the audio and transcribed text can be analyzed and processed to automatically generate structured meeting minutes. For example, this process includes: first, performing semantic understanding and analysis on the text to identify and extract core topics, key conclusions, to-do items, and responsible persons from the meeting, transforming linear dialogue content into logically clear structured information.
[0100] Furthermore, it can intelligently link the minutes content with existing knowledge bases within the user's authorized scope: when analysis reveals mentions of specific projects, contracts, or topics in the transcribed text, it can automatically retrieve and call relevant documents with high matching degrees (such as past meeting minutes, project plans, or draft contracts), incorporating them as references or attachments into the current minutes. For example, when discussing "project goals," it automatically links to project planning documents. This function not only significantly reduces the time spent on manual data organization and retrieval, but more importantly, it proactively connects the output of a single meeting with existing knowledge assets, making the generated minutes an effective information hub that connects the preceding and following sessions and supports decision-making, achieving an upgrade from "dialogue records" to "knowledge integration."
[0101] Furthermore, in one embodiment, this application also supports generating independent minutes based on the acquisition stage. For example, each acquisition stage corresponds to a media stream slice with a clear geographical location or scene characteristics, and a structured minute can be generated for each slice. This minute not only includes basic information, such as the recording duration and geographical location of the slice, but can also integrate full-text obtained from speech-to-text transcription, automatically extracted intelligent summaries, and extracted keywords, thereby achieving a refined description and knowledge extraction of the recorded content in each independent scene.
[0102] Whether it's a global meeting summary generated for a complete recording task or a segmented summary generated for each independent acquisition stage, both can be uniformly stored in the aforementioned recording space. This space, acting as a multi-dimensional information hub, can organically integrate meeting recordings, various summary texts, related referenced documents, and identified to-do items. It achieves intelligent association and unified management of multi-dimensional information such as audio, text, documents, and events, significantly improving the structuring level, retrieval efficiency, and knowledge reuse value of the recorded content.
[0103] Figure 6 This is a block diagram illustrating an embodiment of a multimedia recording device provided in this application. Figure 6 As shown, the device includes: The intelligent segmentation module 61 is used to divide the media stream acquisition process into multiple acquisition stages according to changes in the acquisition scene information during the execution of the multimedia recording task. Device scheduling module 62 is used to schedule at least one of multiple media acquisition devices to participate in the multimedia recording task for media stream acquisition; The integration module 63 is used to synthesize the media streams from the scheduled media acquisition devices based on the multiple acquisition stages to generate the recording file of the multimedia recording task.
[0104] In one possible implementation, the device scheduling module 62 is specifically used for: In response to a media acquisition source adjustment event, the second media acquisition device indicated by the media acquisition source adjustment event is scheduled to participate in the media stream acquisition of the multimedia recording task, and the media stream from the second media acquisition device is integrated into the media stream space maintained for the multimedia recording task according to the acquisition time; The integration module 63 is specifically used for: Based on the multiple acquisition stages, the media streams integrated in the media stream space are uniformly synthesized to generate the recording file for the multimedia recording task.
[0105] In one possible implementation, the device scheduling module schedules the second media acquisition device indicated by the media acquisition source adjustment event to participate in the media stream acquisition of the multimedia recording task, including: Switch the media acquisition device currently participating in the media stream acquisition from the first media acquisition device to the second media acquisition device indicated by the media acquisition source adjustment event; or, The first media acquisition device currently participating in the media stream acquisition and the second media acquisition device indicated by the media acquisition source adjustment event are scheduled to jointly participate in the media stream acquisition of the multimedia recording task.
[0106] In one possible implementation, the device scheduling module 62 schedules the second media acquisition device indicated by the media acquisition source adjustment event to participate in the media stream acquisition of the multimedia recording task, including: The access information used to access the media stream space is shared to the second media acquisition device; In response to the joining request initiated by the second media acquisition device based on the access information, the second media acquisition device is scheduled to participate in the media stream acquisition of the multimedia recording task.
[0107] In one possible implementation, the device further includes: The media streaming module is used to establish a media streaming channel between the second media acquisition device and the selected device; The media stream acquired by the second media acquisition device is transmitted to the selected device through the media stream transmission channel, so as to display the text content obtained by speech-to-text transcription of the media stream acquired by the second media acquisition device on the selected device.
[0108] In one possible implementation, the device further includes: The adaptive adjustment module is used to respond to the user's adjustment of selected text in the text during the process of obtaining text content through speech transcription, and to use the adjusted text as the unified identification content. In subsequent speech transcription processes, the unified identification content is used for speech transcription.
[0109] In one possible implementation, the device further includes: The acquisition control module is used to send a control command to the second media acquisition device in response to a pause command triggered on the first media acquisition device, so as to cause the second media acquisition device to pause media stream acquisition.
[0110] In one possible implementation, the integration module 63 includes: The time alignment unit is used to perform time stamp alignment processing on media streams from different scheduled media acquisition devices; The generation unit is used to generate the recording file for the multimedia recording task based on the media stream that has been time-stamped and aligned.
[0111] In one possible implementation, the intelligent sharding module 61 is specifically used for: Obtain the geographic location information of the media acquisition devices currently participating in the media stream acquisition; Based on the changes in the geographical location information, the media stream acquisition process is divided into stages, generating a new acquisition stage.
[0112] In one possible implementation, the device further includes: The minutes generation module is used to perform content analysis and processing on the recorded files; Based on the results of the content analysis and processing, a structured summary is generated.
[0113] In one possible implementation, the plurality of media acquisition devices includes a virtual media source provided by a designated online conferencing application running on a first media acquisition device.
[0114] like Figure 7 As shown in the figure, this application provides an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114. Memory 113 is used to store computer programs; In one embodiment of this application, when the processor 111 executes the program stored in the memory 113, it implements the multimedia recording method provided in any of the foregoing method embodiments, including: During the execution of a multimedia recording task, the media stream acquisition process is divided into multiple acquisition stages according to changes in the acquisition scene information, and at least one of the multiple media acquisition devices is scheduled to participate in the media stream acquisition of the multimedia recording task. Based on the multiple acquisition stages, the media streams from the scheduled media acquisition devices are synthesized to generate the recording file for the multimedia recording task.
[0115] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the multimedia recording method provided in any of the foregoing method embodiments.
[0116] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0117] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0118] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0119] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A multimedia recording method, characterized in that, The method includes: During the execution of a multimedia recording task, the media stream acquisition process of the multimedia recording task is divided into multiple acquisition stages according to the changes in the acquisition scene information, and at least one of the multiple media acquisition devices is scheduled to participate in the media stream acquisition of the multimedia recording task. Based on the multiple acquisition stages, the media streams from the scheduled media acquisition devices are synthesized to generate the recording file for the multimedia recording task.
2. The method according to claim 1, characterized in that, The scheduling of at least one of multiple media acquisition devices to participate in the media stream acquisition of the multimedia recording task includes: In response to a media acquisition source adjustment event, the second media acquisition device indicated by the media acquisition source adjustment event is scheduled to participate in the media stream acquisition of the multimedia recording task, and the media stream from the second media acquisition device is integrated into the media stream space maintained for the multimedia recording task according to the acquisition time; The process of synthesizing the media streams from the scheduled media acquisition devices based on the multiple acquisition stages to generate the recording file for the multimedia recording task includes: Based on the multiple acquisition stages, the media streams integrated in the media stream space are uniformly synthesized to generate the recording file for the multimedia recording task.
3. The method according to claim 2, characterized in that, The scheduling of the second media acquisition device indicated by the media acquisition source adjustment event to participate in the media stream acquisition of the multimedia recording task includes: Switch the media acquisition device currently participating in the media stream acquisition from the first media acquisition device to the second media acquisition device indicated by the media acquisition source adjustment event; or, The first media acquisition device currently participating in the media stream acquisition and the second media acquisition device indicated by the media acquisition source adjustment event are scheduled to jointly participate in the media stream acquisition of the multimedia recording task.
4. The method according to claim 2, characterized in that, The scheduling of the second media acquisition device indicated by the media acquisition source adjustment event to participate in the media stream acquisition of the multimedia recording task includes: The access information used to access the media stream space is shared to the second media acquisition device; In response to the joining request initiated by the second media acquisition device based on the access information, the second media acquisition device is scheduled to participate in the media stream acquisition of the multimedia recording task.
5. The method according to claim 2, characterized in that, The method further includes: Establish a media stream transmission channel between the second media acquisition device and the selected device; The media stream acquired by the second media acquisition device is transmitted to the selected device through the media stream transmission channel, so as to display the text content obtained by speech-to-text transcription of the media stream acquired by the second media acquisition device on the selected device.
6. The method according to claim 5, characterized in that, The method further includes: During the process of obtaining text content through speech-to-text transcription, in response to the user's adjustment of the selected text in the text, the adjusted text is used as the unified identification content, and in subsequent speech-to-text transcription processes, the unified identification content is used for speech-to-text transcription.
7. The method according to claim 2, characterized in that, The method further includes: In response to triggering a pause command on the first media acquisition device, a control command is sent to the second media acquisition device to cause the second media acquisition device to pause media stream acquisition.
8. The method according to claim 1, characterized in that, The process of synthesizing the media stream from the scheduled media acquisition device to generate the recording file for the multimedia recording task includes: Time stamp alignment is performed on media streams from different scheduled media acquisition devices; The recording file for the multimedia recording task is generated based on the media stream that has been time-stamped and aligned.
9. The method according to claim 1, characterized in that, The media stream acquisition process is divided into multiple acquisition stages based on changes in the acquisition scenario information, including: Obtain the geographic location information of the media acquisition devices currently participating in the media stream acquisition; Based on the changes in the geographical location information, the media stream acquisition process is divided into stages, generating a new acquisition stage.
10. The method according to claim 1, characterized in that, The method further includes: The recorded file is subjected to content analysis and processing; Based on the results of the content analysis and processing, a structured summary is generated.
11. The method according to any one of claims 1-10, characterized in that, The plurality of media acquisition devices include: virtual media sources provided by a designated online conferencing application running on an electronic device.
12. A multimedia recording device, characterized in that, The device includes: The intelligent segmentation module is used to divide the media stream acquisition process into multiple acquisition stages based on changes in the acquisition scene information during the execution of a multimedia recording task. The device scheduling module is used to schedule at least one of multiple media acquisition devices to participate in the media stream acquisition of the multimedia recording task; The integration module is used to synthesize the media streams from the scheduled media acquisition devices based on the multiple acquisition stages, and generate the recording file for the multimedia recording task.
13. An electronic device, characterized in that, include: A processor and a memory, the processor being configured to execute a multimedia recording program stored in the memory to implement the multimedia recording method according to any one of claims 1-11.