Conference audio and video mixing method and device, electronic equipment and storage medium

By introducing meeting topic recognition and participant role awareness mechanisms into video mixing technology, and dynamically adjusting audio and video weights, the rigidity of mixing strategies and resource waste in existing technologies are solved, thereby improving the efficiency of key information transmission and user experience in video conferencing.

CN122179522APending Publication Date: 2026-06-09CHINA MOBILE COMM LTD RES INST +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE COMM LTD RES INST
Filing Date
2026-01-20
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing video mixing technologies lack the ability to perceive the real-time communication context and cannot dynamically optimize mixing strategies. This leads to problems such as audio mixing and distortion, obscuring of key images, and wasted bandwidth resources in complex network environments or changing usage scenarios, affecting the service quality and user experience of video conferencing systems.

Method used

By introducing a meeting topic recognition and participant role awareness mechanism, the system can acquire audio and video streams and meeting information in real time, identify meeting topics and participant roles, dynamically adjust audio and video weights, generate personalized mixing solutions, and optimize audio and video processing.

Benefits of technology

It significantly improves the efficiency of transmitting key information and user experience in multi-party video conferencing, reduces network bandwidth and system resource consumption, and solves the problems of rigid mixing strategies and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122179522A_ABST
    Figure CN122179522A_ABST
Patent Text Reader

Abstract

The present disclosure provides a conference audio and video mixed flow method and device, electronic equipment and storage medium, comprising: in response to a conference request, real-time acquisition of audio and video stream and conference information; based on the audio and video stream and the conference information, identifying conference theme information and role information; based on the theme information and the role information, adjusting the initial audio weight and the target video weight to generate the target audio weight and the target video weight; based on the target audio weight and the target video weight, determining a target audio and video mixed flow scheme, and based on the target audio and video mixed flow scheme, performing mixed flow processing on the audio and video stream to generate target audio and video data. By introducing conference theme recognition and participant role perception mechanism, the dynamic, personalized and scenario of the audio and video mixed flow strategy are realized, the transmission efficiency of key information and user experience in multi-party video conference are significantly improved, and the network bandwidth and system resource consumption are effectively reduced, solving the problem of resource waste in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of video vertical technology, and in particular to a method, apparatus, electronic device and storage medium for mixing audio and video in a conference. Background Technology

[0002] With the rapid development of video network technology, video conferencing, as one of its core application scenarios, has been widely used in many fields such as remote work, online education, medical consultation, and emergency command. In multi-party video conferencing systems, multiple audio and video streams collected by participating terminals need to be collaboratively processed by a media server to achieve a low-latency, highly synchronized real-time interactive experience. Among these processes, mixing technology, as a key processing step, is responsible for merging the original audio and video streams from multiple participants into a single unified output stream. This not only significantly reduces the decoding and rendering burden on the terminal side but also simplifies the network transmission structure, providing crucial technical support for ensuring the stable operation of large-scale multi-party conferences.

[0003] Currently, mainstream mixing solutions generally adopt a static and uniform processing strategy. Specifically, the system typically mixes all input streams indiscriminately according to preset fixed logic (such as screen layout templates, audio overlay rules, encoding parameters, etc.), and the mixing behavior remains unchanged regardless of changes in the network conditions, device capabilities, speaking activity levels, or business scenarios of the participants. For example, in audio processing, all channels are often simply overlaid or mixed with a fixed gain; in video processing, a fixed picture-in-picture layout or a polling switching strategy is used. Summary of the Invention

[0004] This disclosure aims to at least partially address one of the technical problems in the related art.

[0005] Therefore, one objective of this disclosure is to propose a method for mixing audio and video during conferencing.

[0006] The second objective of this disclosure is to provide a conferencing audio and video mixing device.

[0007] The third objective of this disclosure is to propose an electronic device.

[0008] The fourth objective of this disclosure is to provide a non-transitory computer-readable storage medium.

[0009] The fifth objective of this disclosure is to provide a computer program product.

[0010] To achieve the above objectives, a first aspect of this disclosure provides a method for audio and video mixing in a conference, comprising: in response to receiving a conference request, acquiring in real time the audio and video streams and conference information of each participant in a target conference; identifying the conference topic information of the target conference and the role information corresponding to each participant based on the audio and video streams and the conference information; adjusting the initial audio weight and initial video weight of the participants based on the topic information and the role information to generate target audio weight and target video weight; for any participant, determining a target audio and video mixing scheme for the participant based on the target audio weight and the target video weight, and mixing the audio and video streams corresponding to the participant based on the target audio and video mixing scheme to generate target audio and video data.

[0011] According to one embodiment of this disclosure, determining the meeting topic information and the role information corresponding to each participant based on the audio and video stream and the meeting information includes: obtaining session metadata of the audio and video stream; and determining the meeting topic information and the role information based on the meeting information and the session metadata.

[0012] According to one embodiment of this disclosure, adjusting the initial audio weight and initial video weight of the participant based on the topic information and the role information to generate target audio weight and target video weight includes: determining the weight priority of the participant based on the topic information and the role information; and adjusting the initial audio weight and initial video weight of the participant based on the weight priority to generate target audio weight and target video weight.

[0013] According to one embodiment of this disclosure, the method further includes: determining the key semantics of the participant based on the audio and video stream; and adjusting the target audio weight and the target video weight of the participant based on the key semantics to generate adjusted target audio weight and target video weight.

[0014] According to one embodiment of this disclosure, the target audio weight and the target video weight satisfy weight constraint conditions.

[0015] According to one embodiment of this disclosure, determining the target audio-video mixing scheme for any participant based on the target audio weight and the target video weight includes: determining the target audio mixing scheme based on the target audio weight, and determining the target video mixing scheme based on the target video weight; and using the target audio mixing scheme and the target video mixing scheme as the target audio-video mixing scheme for the participant.

[0016] According to one embodiment of this disclosure, determining a target audio mixing scheme based on the target audio weight includes: comparing the target audio weight with a target audio weight judgment threshold; in response to the target audio weight being greater than or equal to the target audio weight judgment threshold, determining the target audio mixing scheme as performing high-priority mixing processing on the audio stream of the participant; or, in response to the target audio weight being less than the target audio weight judgment threshold, determining the target audio mixing scheme as performing low-priority mixing processing on the audio stream of the participant, wherein the low-priority mixing processing includes at least one of downsampling, bandwidth compression, or silence suppression on the audio stream.

[0017] According to one embodiment of this disclosure, determining the target video mixing scheme based on the target video weight includes: determining a target display area based on the target video weight, and comparing the target video weight with a target video weight judgment threshold; in response to the target video weight being greater than or equal to the target video weight judgment threshold, determining the target video mixing scheme as follows: rendering the participant's video stream to the target display area and performing full frame rate video mixing processing; or, in response to the target video weight being less than the target video weight judgment threshold, determining the target video mixing scheme as follows: intelligently sampling the participant's video stream based on the target video weight to obtain a sampled video stream, and rendering the sampled video stream to the target display area to perform video mixing processing.

[0018] According to one embodiment of this disclosure, determining the target display area based on the target video weight includes: determining the target display position and target display size of the video stream corresponding to the participant based on the target video weight; and determining the target display area based on the target display position and the target display size.

[0019] According to one embodiment of this disclosure, the step of mixing the audio and video streams corresponding to the participants based on the target audio and video mixing scheme to generate target audio and video data includes: performing audio mixing processing on the audio and video streams corresponding to the participants based on the target audio mixing scheme to generate target audio data, and performing video mixing processing on the target video mixing scheme based on the target video mixing scheme to generate target video data; and performing synchronization processing on the target audio data and the target video data to generate the target audio and video data.

[0020] According to one embodiment of this disclosure, the method further includes: receiving a viewing instruction from a participating terminal, the viewing instruction specifying the audio and video of a target participant; and adjusting the target audio weight and target video weight displayed on the participating terminal for the target participant to the upper limit values ​​of the target audio weight and the target video weight, respectively.

[0021] According to one embodiment of this disclosure, determining the target audio-video mixing scheme for any participant based on the target audio weight and the target video weight includes: determining the target audio mixing scheme based on the target audio weight and determining the target lightweight video mixing scheme based on the target audio weight when the participant's terminal coefficient is less than a processing threshold; and using the target audio mixing scheme and the target lightweight video mixing scheme as the participant's target audio-video mixing scheme.

[0022] To achieve the above objectives, a second aspect of this disclosure provides a conference audio mixing device, comprising: an acquisition module, configured to acquire, in real time, the audio and video streams of each participant in a target conference in response to receiving a conference request; a determination module, configured to identify the conference theme information of the target conference and the role information corresponding to each participant based on the audio and video streams; an allocation module, configured to allocate target audio weights and target video weights to the participants based on the theme information and the role information; and a processing module, configured to determine a target audio and video mixing scheme for any participant based on the target audio weights and the target video weights, and to perform mixing processing on the audio and video streams corresponding to the participant based on the target audio and video mixing scheme to generate target audio and video data.

[0023] To achieve the above objectives, a third aspect of this disclosure provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to implement the conference audio and video mixing method as described in the first aspect of this disclosure.

[0024] To achieve the above objectives, a fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to implement the conference audio and video mixing method as described in the first aspect of this disclosure.

[0025] To achieve the above objectives, a fifth aspect of this disclosure provides a computer program product including a computer program that, when executed by a processor, implements the conference audio and video mixing method as described in the first aspect of this disclosure.

[0026] By introducing a meeting topic recognition and participant role awareness mechanism, the audio and video mixing strategy is made dynamic, personalized, and contextualized. This not only significantly improves the efficiency of key information transmission and user experience in multi-party video conferences, but also effectively reduces network bandwidth and system resource consumption, solving problems such as resource waste in existing technologies. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of a conference audio and video mixing method according to one embodiment of the present disclosure; Figure 2 This is a schematic diagram of another method for mixing audio and video during conferencing according to one embodiment of this disclosure; Figure 3 This is a schematic diagram of another method for mixing audio and video during conferencing according to one embodiment of this disclosure; Figure 4 This is a schematic diagram of another method for mixing audio and video during conferencing according to one embodiment of this disclosure; Figure 5 This is a schematic diagram of another method for mixing audio and video during conferencing according to one embodiment of this disclosure; Figure 6 This is a schematic diagram of a conference audio and video mixing device according to one embodiment of the present disclosure; Figure 7 This is a schematic diagram of an electronic device according to one embodiment of the present disclosure. Detailed Implementation

[0028] Embodiments of this disclosure are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.

[0029] The acquisition, storage, use, and processing of data in this disclosed technical solution all comply with the relevant provisions of relevant laws and regulations.

[0030] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0031] Current technologies have significant shortcomings in their "one-size-fits-all" mixing mechanisms: On the one hand, it lacks the ability to perceive the real-time communication context and cannot dynamically optimize the mixing strategy based on factors such as voice activity, image importance, and network bandwidth fluctuations; On the other hand, it is difficult to meet the differentiated needs of diverse business scenarios. For example, high-fidelity voice conferencing requires highlighting the speaker's voice quality, while emergency command scenarios need to prioritize the clear presentation of key images.

[0032] As a result, in complex network environments or variable usage scenarios, problems such as audio distortion and noise, obstruction of key images, and waste of bandwidth resources are likely to occur, which seriously restrict the quality of service (QoS) and user experience (QoE) of video conferencing systems.

[0033] To address the aforementioned issues, this disclosure proposes a method for mixing audio and video during conferences. Figure 1 This is a schematic diagram of a conference audio and video mixing method according to one embodiment of the present disclosure, as shown below. Figure 1 As shown, the audio and video mixing method for this conference includes the following steps: S101, in response to receiving a meeting request, acquires the audio and video streams and meeting information of each participant in the target meeting in real time.

[0034] The audio and video mixing method of this application embodiment can be applied to video conferencing scenarios. The entity executing the audio and video mixing of this application embodiment can be the audio and video mixing device of this application embodiment, which can be installed on an electronic device.

[0035] It should be noted that the target meeting is the specific meeting currently being processed.

[0036] The number of participants in this embodiment is multiple, depending on the actual meeting, and no limitation is made here.

[0037] In a multi-party conference, the server needs to receive N independent audio and video streams simultaneously (N = number of participants) to provide a basis for subsequent multi-stream mixing.

[0038] In this embodiment of the disclosure, meeting information can be synchronized from a meeting database or a scheduling center. The meeting information may include various types of data, such as, but not limited to, the following: Meeting topic, creator, start / expected end time; Meeting modes (such as "presentation mode", "roundtable discussion", "mute all participants"); List of attendees and their roles (facilitator, collaborator, audience); Permission configuration (whether to allow camera access, screen sharing, and chat); Network policies (whether to enable SVC layered encoding, maximum resolution limit); Associated resources (whiteboard ID, shared document URL, recording status).

[0039] S102 identifies the meeting topic information of the target meeting and the corresponding role information of each participant based on audio and video streams and meeting information.

[0040] It should be noted that the meeting topic information represents a structured label for the meeting type, such as "Project Progress Meeting," "Client Negotiation," or "Fault Review," which can be output by keyword clustering, agenda matching, or classification models.

[0041] Role information describes the functional attributes of participants in the current meeting, including but not limited to: moderator, speaker, collaborator, observer, system bot, etc.

[0042] In this embodiment of the disclosure, audio and video streams and meeting information can be analyzed to determine the role information corresponding to each participant. For example, the meeting type or topic can be determined by analyzing the speaking keywords, shared documents, PPT titles, meeting time, and participant identities in the audio and video streams, such as "project weekly meeting," "product review," or "emergency dispatch."

[0043] It should be noted that identifying the role information of each participant means determining the identity or function of each participant in the meeting, such as "moderator", "speaker", "recorder", "ordinary participant", "expert", etc.

[0044] S103, adjust the initial audio weights and initial video weights of participants based on topic information and role information to generate target audio weights and target video weights.

[0045] In this embodiment, the initial audio and video weights of participants are pre-assigned, typically based on a default strategy (such as "high weight upon activation"). However, these initial weights do not consider the semantic context of the meeting or the user's role. Therefore, this solution introduces topic information and role information as adjustment factors to generate more reasonable target audio and video weights.

[0046] It should be noted that the target audio weight and target video weight are numerical priority parameters assigned by the system to each participant, used to control their volume, screen size, encoding quality, and whether they are highlighted.

[0047] In this embodiment of the disclosure, there are various methods for adjusting the initial audio weights and initial video weights of participants based on topic information and role information to generate target audio weights and target video weights, and no limitation is made here.

[0048] In one possible implementation, the target audio weight and target video weight for each participant can be determined by looking up the topic information and role information in a pre-defined mapping table. This mapping relationship can be pre-designed, based on historical video conferencing data, or it can be modified according to actual design needs. For example, it can be shown in the table below:

[0049] In another possible implementation, a weight generation model can be used to process topic information, role information, initial audio weights, and initial video weights to generate target audio weights and target video weights for each participant. This weight generation model is pre-trained and stored in the electronic device's storage space for easy retrieval when needed.

[0050] Adjusting the initial audio and video weights of participants based on topic and role information to generate target audio and video weights allows for more intelligent mixing results that better suit the needs of the meeting scenario. Specifically, it can achieve the following: 1. Enhance the visibility and listenability of key information: During the "product launch event," automatically amplify the product manager's voice and visuals; In "emergency command," ensure that the commander is always on the main screen and that their voice is clear; In the "classroom teaching" mode, the teacher's video is displayed in full screen, while the student's image is minimized or hidden.

[0051] 2. Optimize resource allocation: Reduce video resolution or audio sampling rate for low-weight participants to save bandwidth; To avoid the "noisy reverberation" problem caused by the equal superposition of all sounds.

[0052] 3. Achieve dynamic adaptation: The same person may have different weight in different meetings. For example, A may be the speaker in meeting A and an audience member in meeting B. Roles may switch during the same meeting, such as the host temporarily giving up the microphone, and the weight can be updated in real time.

[0053] S104. For any participant, determine the target audio and video mixing scheme based on the target audio weight and the target video weight, and perform mixing processing on the audio and video streams corresponding to the participant based on the target audio and video mixing scheme to generate target audio and video data.

[0054] The target audio and video data is the final generated single-channel synthesized audio and video stream, which is received and played by all participants. For example, a window displaying "multi-person video + mixed audio" in a multi-person conference.

[0055] In this embodiment of the disclosure, based on the "importance of each person's voice and image" (i.e., target audio weight and target video weight), it is determined how to process their audio and video—for example, increase the volume, enlarge the screen, or mute them; then, according to this rule, all the audio and video of all people are combined into one final output.

[0056] In this embodiment, upon receiving a meeting request, the system first acquires the audio and video streams and meeting information of each participant in the target meeting in real time. Then, based on the audio and video streams and meeting information, it identifies the meeting topic information and the corresponding role information of each participant. Next, based on the topic information and role information, it adjusts the initial audio weight and initial video weight of each participant to generate target audio weight and target video weight. Finally, for any participant, it determines the target audio and video mixing scheme based on the target audio and video weights, and performs mixing processing on the corresponding audio and video streams to generate target audio and video data. Thus, by introducing a meeting topic recognition and participant role awareness mechanism, the system achieves dynamic, personalized, and scenario-based audio and video mixing strategies. This not only significantly improves the transmission efficiency of key information and user experience in multi-party video conferences but also effectively reduces network bandwidth and system resource consumption. It solves the problems of rigid mixing strategies, resource waste, and the easy submersion of key content in existing technologies, providing key technical support for building next-generation intelligent video conferencing systems.

[0057] In the above embodiments, the meeting topic information and the role information of each participant are determined based on the audio and video streams and meeting information. Furthermore, it can be achieved through... Figure 2 To further explain, the method includes: S201, retrieve session metadata of audio and video streams.

[0058] It should be noted that session metadata can be descriptive information about the audio and video streams, or data obtained after parsing the audio and video streams. There are various methods for obtaining session metadata, and no restrictions are placed here.

[0059] In one possible implementation, session metadata can be obtained through a meeting scheduling system, such as obtaining the meeting title, description, agenda, etc.

[0060] It can also be obtained through signaling protocols (such as SIP, WebRTC SDP), for example, by obtaining the initiator ID, participant list, device type, etc.

[0061] It can also be obtained through user account information, such as department, role, and organization affiliation.

[0062] It can also be obtained through the network and contextual information, such as determining the meeting start time, duration, and whether external personnel are included.

[0063] Data can also be obtained through data analysis of audio and video streams, such as identifying attendees and their environment.

[0064] S202, determine meeting topic information and role information based on session metadata and meeting information.

[0065] In this embodiment of the disclosure, metadata and meeting information can be analyzed through preset rules or pre-trained models to determine features and role information. For example, the meeting topic can be inferred from the meeting information, and role information can be determined based on the participants' speech content or built-in tags in the session metadata and meeting information.

[0066] In this embodiment, session metadata of the audio and video streams is first acquired, and then the meeting topic information is determined based on the session metadata and meeting information. This solution determines the meeting topic information by acquiring and analyzing the meeting's session metadata and meeting information, without relying on deep analysis of the audio and video content. Thus, while protecting user privacy, it achieves rapid semantic understanding and intelligent strategy preloading at the beginning of the meeting.

[0067] In the above embodiments, the initial audio weights and initial video weights of participants are adjusted based on topic information and role information to generate target audio weights and target video weights. Furthermore, [the system can be modified / adjusted] through [other methods]. Figure 3 To further explain, the method includes: S301, determine the weight priority of participants based on topic information and role information.

[0068] It should be noted that the weight priority is based on the relative ranking or category labels assigned to the participants. For example: Priority level: High / Medium / Low; Rank number: 1 (highest), 2, 3...; Group labels: Host / Explainer / Regular attendee.

[0069] In this embodiment of the disclosure, the weight priority of participants based on topic information and role information can be determined based on a preset mapping relationship.

[0070] S302, adjust the initial audio weight and initial video weight of the participants based on weight priority to generate target audio weight and target video weight.

[0071] In this embodiment of the disclosure, after obtaining the weight priority, the initial audio weight and initial video weight can be adjusted to generate target audio weight and target video weight. The target audio weight and target video weight can be directly used to characterize data such as "how much the sound should be amplified" and "how big the screen should be" for the corresponding participants.

[0072] In one possible implementation, "priority" can be converted into specific weight values ​​through preset rules or a strategy table. For example, it could be shown in the table below:

[0073] In this embodiment, the weight priority of participants is first determined based on topic information and role information. Then, the initial audio weight and initial video weight of the participants are adjusted based on the weight priority to generate target audio weight and target video weight. Thus, through this solution, video conferencing can no longer be a simple splicing of images and sound, but rather an intelligent interactive experience that "knows who to listen to and who to watch," just like a human.

[0074] In another possible implementation, the key semantics of the participants can be determined based on the audio and video streams, and then the target audio weights and target video weights of the participants can be adjusted based on the key semantics to generate adjusted target audio weights and target video weights.

[0075] In one possible implementation, artificial intelligence (AI) algorithms can be used to identify keywords in speech based on the video theme, video communication content, or user-defined keywords. For example, keywords in a parenting scenario might include "falling down" or "fever," while keywords in a business scenario might include "budget" or "delay." When the media server detects key content spoken by a low-weight user (such as AI identifying "budget overrun"), it automatically increases that user's weight. In this embodiment of the disclosure, in order to avoid imbalance between audio and target video weights, the target audio weight and the target video weight satisfy weight constraint conditions.

[0076] It should be noted that the weight constraints are pre-designed. The target audio weight (such as volume gain coefficient) and target video weight (such as screen ratio or encoding quality) assigned to each participant are not arbitrary, but must conform to a set of preset mathematical or business rules (i.e., "weight constraints") to avoid resource conflicts, output distortion or experience degradation.

[0077] In this embodiment of the disclosure, the weight constraint condition can be that the difference between the audio and target video weights for each user is less than a threshold α.

[0078] In the above embodiments, for any participant, the target audio and video mixing scheme for that participant is determined based on the target audio weight and the target video weight. It can also be achieved through... Figure 4 To further explain, the method includes: S401, determine the target audio mixing scheme based on the target audio weight, and determine the target video mixing scheme based on the target video weight.

[0079] In this embodiment of the disclosure, the target audio mixing scheme is determined based on the target audio weight. First, the target audio weight is compared with the target audio weight judgment threshold. In response to the target audio weight being greater than or equal to the target audio weight judgment threshold, the target audio mixing scheme is determined to perform high-priority mixing processing on the audio stream of the participants. Alternatively, in response to the target audio weight being less than the target audio weight judgment threshold, the target audio mixing scheme is determined to perform low-priority mixing processing on the audio stream of the participants. The low-priority mixing processing includes at least one of downsampling, bandwidth compression, or silence suppression on the audio stream.

[0080] It should be noted that the target audio weight judgment threshold is a preset critical value (e.g., 0.5 or 0.6), which is equivalent to the "importance score line".

[0081] If the target audio weight is greater than or equal to the target audio weight judgment threshold, high-priority mixing processing is performed. High-priority mixing processing can be one or more of the following, without any limitations here: Preserve full sound quality; The volume is relatively high, serving as the main channel output; No compression or noise reduction cropping is performed; Dominates in multi-person mixing.

[0082] High-priority mixed streaming is suitable for speakers, moderators, key decision-makers, etc.

[0083] If the target audio weight is less than the target audio weight judgment threshold, it is processed with low priority. Low-priority processing is to save bandwidth, reduce interference, and reduce computational overhead, and specifically includes at least one of the following methods: Downsampling: Reduces audio quality from high quality (e.g., 48kHz) to low quality (e.g., 8kHz), retaining only the fundamental frequencies of human voices; Bandwidth compression: Encode with a lower bitrate (e.g., compress from 64kbps to 16kbps) to reduce network traffic. Mute suppression: When the sound is weak or unimportant, it is not sent or is muted (common among inactive listeners).

[0084] In this embodiment of the disclosure, the target video mixing scheme is determined based on the target video weight. The target display area can be determined based on the target video weight, and the target video weight can be compared with the target video weight judgment threshold. If the target video weight is greater than or equal to the target video weight judgment threshold, the target video mixing scheme is determined as follows: the participant's video stream is rendered to the target display area and full frame rate video mixing processing is performed. Alternatively, if the target video weight is less than the target video weight judgment threshold, the target video mixing scheme is determined as follows: the participant's video stream is intelligently sampled based on the target video weight to obtain a sampled video stream, and the sampled video stream is rendered to the target display area to perform video mixing processing.

[0085] It should be noted that the target video weight judgment threshold is a preset critical value used to divide the boundary between "important" and "unimportant".

[0086] When the target video weight is greater than or equal to the target video weight judgment threshold, the participant's video stream is rendered to the target display area and full frame rate video mixing processing is performed. It should be noted that full frame rate video mixing processing includes one or more of the following processes: Use the original resolution (e.g., 1080p); Maintain a high frame rate (e.g., 30fps); No compression, no cropping, no bitrate reduction; The entire scene is presented (including details such as background and gestures).

[0087] When the target video weight is less than the target video weight judgment threshold, the participant's video stream needs to be intelligently sampled based on the target video weight.

[0088] Intelligent sampling refers to strategically reducing video quality based on weights, and may include one or more of the following operations: Reduce the resolution, for example, convert the participants' video streams from 080p to 480p or 360p; Reduce the frame rate, for example, by reducing the participant's video stream from 30fps to 10fps; Region cropping, for example, keeping only the face area of ​​the participant's video stream and removing the background; Bitrate compression, for example, transmitting the participants' video streams with less bandwidth.

[0089] The sampled video will still be displayed in the assigned target area (such as a small window), but resource consumption will be significantly reduced.

[0090] In this embodiment of the disclosure, the target display area is determined based on the target video weight. First, the target display position and target display size of the video stream corresponding to the participant are determined based on the target video weight, and then the target display area is determined based on the target display position and target display size.

[0091] In one possible implementation, to avoid cluttering user windows, the meeting page is divided into three levels of areas based on weight, ensuring high visibility of video images for users with high weight.

[0092] Main display area (>60%): Occupied by high-authority users (such as speakers), with core content displayed first; Secondary display area (20%-60%): Occupied by users with medium weight (such as active participants) to ensure basic visibility; Truncation area (<20%): Occupied by low-weight users (such as silent participants), saving space resources.

[0093] For example, in a 3-person meeting, when the target video weights are distributed as (0.6, 0.3, 0.1), the user with a weight of 0.6 occupies the main display area (70% of the area), the user with a weight of 0.3 occupies the secondary display area (20% of the area), and the user with a weight of 0.1 occupies the thumbnail area (10% of the area).

[0094] In one possible implementation, to encourage active participation, under certain conditions, low-weight users with high activity levels are allowed to receive a larger screen display than medium-weight users. This is an intentionally designed "inverted" mechanism as an incentive strategy to reward active participation. It should be noted that this "inverted" mechanism only applies between medium and low weight participants; high-weight participants always receive the largest screen display. The duration of the inverted state can be set, for example, to 60 seconds. After the timeout, the default layout will be restored. The maximum number of people allowed to be inverted at the same time can also be set, for example, a maximum of 2 low-weight users can trigger the inverted state at the same time.

[0095] S402, the target audio mixing scheme and the target video mixing scheme are used as the target audio and video mixing scheme for the participants.

[0096] In this embodiment, a target audio mixing scheme is first determined based on the target audio weight, and a target video mixing scheme is determined based on the target video weight. Then, the target audio mixing scheme and the target video mixing scheme are used as the target audio-video mixing scheme for the participants. Thus, by independently generating the target audio mixing scheme and the target video mixing scheme based on the target audio weight and the target video weight respectively, and combining them into a complete audio-video mixing strategy for the participants, decoupling optimization and collaborative fusion of audio and video processing are achieved.

[0097] In this embodiment of the disclosure, the audio and video streams corresponding to the participants are mixed based on the target audio and video mixing scheme to generate target audio and video data. First, the audio and video streams corresponding to the participants are mixed based on the target audio mixing scheme to generate target audio data, and the video and video streams are mixed based on the target video mixing scheme to generate target video data. Then, the target audio data and target video data are synchronized to generate target audio and video data.

[0098] In one possible implementation, when a user actively selects to watch the video of a participant in a video conference, the mixing scheme can be adjusted individually for that user.

[0099] After receiving the viewing instruction from the participating terminal, which specifies the audio and video of the target participant, the target audio weight and target video weight displayed on the participating terminal can be adjusted to the upper limit of the target audio weight and the upper limit of the target video weight, respectively.

[0100] In another possible implementation, the target audio weight and target video weight displayed on the participant's terminal are adjusted to the upper limit values ​​of the target audio weight and target video weight, respectively, while other participants still use... Figures 1-3 The method for mixing conference audio and video in the embodiments.

[0101] In another possible implementation, for any participant, the target audio and video mixing scheme for that participant is determined based on the target audio weight and the target video weight. This can also be achieved through... Figure 5 To further explain, the method includes: S501, in response to the fact that the terminal coefficient of the participant is less than the processing threshold, determines the target audio mixing scheme based on the target audio weight, and determines the target lightweight video mixing scheme based on the terminal coefficient.

[0102] It should be noted that the terminal coefficient is a quantitative indicator of terminal performance, which comprehensively reflects the terminal's computing power, network bandwidth, screen size, battery status, and other resource conditions.

[0103] The processing threshold is a preset performance boundary line of the system. When the terminal coefficient is lower than this threshold, it means that the device is unable to carry standard audio and video streams and a degradation strategy needs to be activated.

[0104] It should be noted that, since video consumes more resources (encoding / decoding, rendering, bandwidth), a lightweight strategy is selected directly based on the specific values ​​of the terminal coefficients. For example, reduce the resolution, such as from 720p to 360p or 240p.

[0105] Reduce the frame rate, for example, from 30fps to 10~15fps.

[0106] Change the layout, for example, from a nine-square grid to a single screen (only the speaker) or a four-square grid.

[0107] Change the lightweight encoding strategy, for example, H.264 High Profile → Baseline Profile.

[0108] The goal of the target lightweight video mixing solution is to generate a video stream with minimized computational and bandwidth overhead, while still retaining identifiable content.

[0109] S502 specifies the target audio mixing solution and the target lightweight video mixing solution as the target audio and video mixing solutions for the participants.

[0110] In this embodiment, firstly, in response to a participant's terminal coefficient being less than a processing threshold, a target audio mixing scheme is determined based on the target audio weight, and a target lightweight video mixing scheme is determined based on the terminal coefficient. Then, the target audio mixing scheme and the target lightweight video mixing scheme are used as the participant's target audio-video mixing scheme. Thus, by using terminal capabilities as an anchor point and employing a differentiated strategy of "heavy semantics in audio, heavy load in video," optimal audio-video service balance is achieved in resource-constrained scenarios. Users with weak terminals no longer exit the meeting due to buffering, while also reducing invalid traffic and saving server bandwidth and computing resources.

[0111] Corresponding to the conference audio and video mixing methods provided in the above embodiments, one embodiment of this disclosure also provides a conference audio and video mixing device. Since the conference audio and video mixing device provided in this disclosure corresponds to the conference audio and video mixing methods provided in the above embodiments, the implementation methods of the above conference audio and video mixing methods are also applicable to the conference audio and video mixing device provided in this disclosure, and will not be described in detail in the following embodiments.

[0112] Figure 6 Figure 6 is a schematic diagram of a conference audio and video mixing device according to one embodiment of the present disclosure. As shown in Figure 6, the conference audio and video mixing device 600 includes: an acquisition module 610, a determination module 620, an allocation module 630, and a processing module 640.

[0113] The acquisition module 610 is used to acquire the audio and video streams and meeting information of each participant in the target meeting in real time in response to receiving a meeting request.

[0114] The determination module 620 is used to identify the meeting topic information of the target meeting and the corresponding role information of each participant based on the audio and video streams and meeting information.

[0115] The allocation module 630 is used to adjust the initial audio weights and initial video weights of participants based on topic information and role information in order to generate target audio weights and target video weights.

[0116] The processing module 640 is used to determine the target audio and video mixing scheme for any participant based on the target audio weight and the target video weight, and to perform mixing processing on the audio and video streams corresponding to the participant based on the target audio and video mixing scheme to generate target audio and video data.

[0117] According to one embodiment of this disclosure, determining the meeting topic information and the role information corresponding to each participant based on the audio and video stream and the meeting information includes: obtaining session metadata of the audio and video stream; and determining the meeting topic information and the role information based on the meeting information and the session metadata.

[0118] According to one embodiment of this disclosure, adjusting the initial audio weight and initial video weight of the participant based on the topic information and the role information to generate target audio weight and target video weight includes: determining the weight priority of the participant based on the topic information and the role information; and adjusting the initial audio weight and initial video weight of the participant based on the weight priority to generate target audio weight and target video weight.

[0119] According to one embodiment of this disclosure, the method further includes: determining the key semantics of the participant based on the audio and video stream; and adjusting the target audio weight and the target video weight of the participant based on the key semantics to generate adjusted target audio weight and target video weight.

[0120] According to one embodiment of this disclosure, the target audio weight and the target video weight satisfy weight constraint conditions.

[0121] According to one embodiment of this disclosure, determining the target audio-video mixing scheme for any participant based on the target audio weight and the target video weight includes: determining the target audio mixing scheme based on the target audio weight, and determining the target video mixing scheme based on the target video weight; and using the target audio mixing scheme and the target video mixing scheme as the target audio-video mixing scheme for the participant.

[0122] According to one embodiment of this disclosure, determining a target audio mixing scheme based on the target audio weight includes: comparing the target audio weight with a target audio weight judgment threshold; in response to the target audio weight being greater than or equal to the target audio weight judgment threshold, determining the target audio mixing scheme as performing high-priority mixing processing on the audio stream of the participant; or, in response to the target audio weight being less than the target audio weight judgment threshold, determining the target audio mixing scheme as performing low-priority mixing processing on the audio stream of the participant, wherein the low-priority mixing processing includes at least one of downsampling, bandwidth compression, or silence suppression on the audio stream.

[0123] According to one embodiment of this disclosure, determining the target video mixing scheme based on the target video weight includes: determining a target display area based on the target video weight, and comparing the target video weight with a target video weight judgment threshold; in response to the target video weight being greater than or equal to the target video weight judgment threshold, determining the target video mixing scheme as follows: rendering the participant's video stream to the target display area and performing full frame rate video mixing processing; or, in response to the target video weight being less than the target video weight judgment threshold, determining the target video mixing scheme as follows: intelligently sampling the participant's video stream based on the target video weight to obtain a sampled video stream, and rendering the sampled video stream to the target display area to perform video mixing processing.

[0124] According to one embodiment of this disclosure, determining the target display area based on the target video weight includes: determining the target display position and target display size of the video stream corresponding to the participant based on the target video weight; and determining the target display area based on the target display position and the target display size.

[0125] According to one embodiment of this disclosure, the step of mixing the audio and video streams corresponding to the participants based on the target audio and video mixing scheme to generate target audio and video data includes: performing audio mixing processing on the audio and video streams corresponding to the participants based on the target audio mixing scheme to generate target audio data, and performing video mixing processing on the target video mixing scheme based on the target video mixing scheme to generate target video data; and performing synchronization processing on the target audio data and the target video data to generate the target audio and video data.

[0126] According to one embodiment of this disclosure, the method further includes: receiving a viewing instruction from a participating terminal, the viewing instruction specifying the audio and video of a target participant; and adjusting the target audio weight and target video weight displayed on the participating terminal for the target participant to the upper limit values ​​of the target audio weight and the target video weight, respectively.

[0127] According to one embodiment of this disclosure, determining the target audio-video mixing scheme for any participant based on the target audio weight and the target video weight includes: determining the target audio mixing scheme based on the target audio weight and determining the target lightweight video mixing scheme based on the target audio weight when the participant's terminal coefficient is less than a processing threshold; and using the target audio mixing scheme and the target lightweight video mixing scheme as the participant's target audio-video mixing scheme.

[0128] By introducing a meeting topic recognition and participant role awareness mechanism, the audio and video mixing strategy has been made dynamic, personalized, and contextualized. This not only significantly improves the transmission efficiency of key information and user experience in multi-party video conferences, but also effectively reduces network bandwidth and system resource consumption. It solves the problems of rigid mixing strategies, resource waste, and easy submersion of key content in existing technologies, and provides key technical support for building the next generation of intelligent video network conferencing systems.

[0129] To implement the above embodiments, this disclosure also proposes an electronic device 700. Figure 7 This is a schematic diagram of an electronic device according to one embodiment of the present disclosure, such as... Figure 7 As shown, the electronic device 700 includes: a processor 701 and a memory 702 communicatively connected to the processor. The memory 702 stores instructions executable by at least one processor. The instructions are executed by at least one processor 701 to implement the functions described in this disclosure. Figures 1-5 The method for mixing conference audio and video in the embodiment.

[0130] To implement the above embodiments, this disclosure also proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to implement the present disclosure. Figures 1-5 The method for mixing conference audio and video in the embodiment.

[0131] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program, which, when executed by a processor, implements the features of this disclosure. Figures 1-5 The method for mixing conference audio and video in the embodiment.

[0132] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0133] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this disclosure is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.

[0134] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0135] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0136] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0137] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that contains, stores, communicates, propagates, or transmits programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0138] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0139] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0140] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0141] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for mixing audio and video streams in a conference, characterized in that, include: Upon receiving a meeting request, it acquires the audio and video streams and meeting information of each participant in the target meeting in real time. Based on the audio and video streams and the meeting information, identify the meeting topic information of the target meeting and the role information corresponding to each participant; Based on the topic information and the role information, the initial audio weight and initial video weight of the participants are adjusted to generate target audio weight and target video weight; For any participant, a target audio and video mixing scheme is determined based on the target audio weight and the target video weight. The audio and video streams corresponding to the participant are then mixed based on the target audio and video mixing scheme to generate target audio and video data.

2. The method according to claim 1, characterized in that, The process of determining the meeting topic information and the corresponding role information of each participant based on the audio and video streams and the meeting information includes: Obtain the session metadata of the audio and video streams; The meeting topic information and the role information are determined based on the meeting information and the session metadata.

3. The method according to claim 1, characterized in that, The step of adjusting the initial audio weight and initial video weight of the participants based on the topic information and the role information to generate target audio weight and target video weight includes: The weight and priority of the participants are determined based on the topic information and the role information. The initial audio weight and initial video weight of the participants are adjusted based on the weight priority to generate target audio weight and target video weight.

4. The method according to any one of claims 1-3, characterized in that, The method further includes: Determine the key semantics of the participants based on the audio and video streams; Based on the key semantics, the target audio weight and target video weight of the participants are adjusted to generate adjusted target audio weight and target video weight.

5. The method according to claim 4, characterized in that, The target audio weight and the target video weight satisfy the weight constraint conditions.

6. The method according to claim 1, characterized in that, The step of determining the target audio and video mixing scheme for any participant based on the target audio weight and the target video weight includes: The target audio mixing scheme is determined based on the target audio weights, and the target video mixing scheme is determined based on the target video weights; The target audio mixing scheme and the target video mixing scheme are used as the target audio and video mixing scheme for the participants.

7. The method according to claim 6, characterized in that, The determination of the target audio mixing scheme based on the target audio weights includes: The target audio weight is compared with the target audio weight judgment threshold; In response to the target audio weight being greater than or equal to the target audio weight judgment threshold, the target audio mixing scheme is determined to perform high-priority mixing processing on the audio streams of the participants; or, In response to the target audio weight being less than the target audio weight judgment threshold, the target audio mixing scheme is determined to perform low-priority mixing processing on the audio stream of the participant. The low-priority mixing processing includes at least one of downsampling, bandwidth compression, or silence suppression on the audio stream.

8. The method according to claim 6, characterized in that, The determination of the target video mixing scheme based on the target video weights includes: The target display area is determined based on the target video weight, and the target video weight is compared with the target video weight judgment threshold. In response to the target video weight being greater than or equal to the target video weight judgment threshold, the target video mixing scheme is determined to be: rendering the participant's video stream to the target display area and performing full frame rate video mixing processing; or, In response to the target video weight being less than the target video weight judgment threshold, the target video mixing scheme is determined as follows: intelligent sampling of the participants' video streams based on the target video weight to obtain the sampled video stream, and rendering the sampled video stream to the target display area to perform video mixing processing.

9. The method according to claim 8, characterized in that, Determining the target display area based on the target video weight includes: The target display position and target display size of the video stream corresponding to the participant are determined based on the target video weight; The target display area is determined based on the target display position and the target display size.

10. The method according to any one of claims 6-9, characterized in that, The process of mixing the audio and video streams corresponding to the participants based on the target audio and video mixing scheme to generate target audio and video data includes: Based on the target audio mixing scheme, the audio and video streams corresponding to the participants are subjected to audio mixing processing to generate target audio data; and based on the target video mixing scheme, the target video mixing scheme is subjected to video mixing processing to generate target video data. The target audio data and the target video data are processed synchronously to generate the target audio and video data.

11. The method according to claim 1, characterized in that, The method further includes: Receive viewing instructions from participating terminals, wherein the viewing instructions specify the audio and video of the target participant to be viewed; The target audio weight and target video weight displayed on the participant's terminal are adjusted to the upper limit of the target audio weight and the upper limit of the target video weight, respectively.

12. The method according to claim 1, characterized in that, The step of determining the target audio and video mixing scheme for any participant based on the target audio weight and the target video weight includes: When the terminal coefficient of the participant is less than the processing threshold, a target audio mixing scheme is determined based on the target audio weight, and a target lightweight video mixing scheme is determined based on the terminal coefficient. The target audio mixing scheme and the target lightweight video mixing scheme are used as the target audio and video mixing schemes for the participants.

13. A conference audio and video mixing device, characterized in that, include: The acquisition module is used to respond to a received meeting request by acquiring the audio and video streams and meeting information of each participant in the target meeting in real time. The determination module is used to identify the meeting topic information of the target meeting and the role information of each participant based on the audio and video stream and the meeting information; The allocation module is used to adjust the initial audio weight and initial video weight of the participants based on the topic information and the role information, so as to generate target audio weight and target video weight; The processing module is used to determine the target audio and video mixing scheme for any participant based on the target audio weight and the target video weight, and to perform mixing processing on the audio and video streams corresponding to the participant based on the target audio and video mixing scheme to generate target audio and video data.

14. An electronic device, characterized in that, Including memory and processor; The processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the method as described in any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-12.