Audio processing method and device, electronic equipment, storage medium and program product

By monitoring the number of interactive objects in real time and intelligently switching audio interaction modes, sub-virtual spaces are dynamically created to share the load, solving the performance bottleneck and resource limitation problems of the virtual space chorus solution when a large number of users flood in, and achieving an efficient and stable audio interaction experience.

CN121940707APending Publication Date: 2026-04-28HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
Filing Date
2025-12-15
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing virtual space chorus solutions struggle to provide a high-quality interactive experience while maintaining real-time performance and stability when faced with a large influx of users. This is especially true during popular songs or large online events, where a single virtual space is prone to performance bottlenecks and resource limitations.

Method used

By monitoring the number of interactive objects in real time, the system intelligently decides whether to enter the multi-space mode based on preset thresholds, dynamically creates sub-virtual spaces to share the load, and merges and displays the audio data of the main virtual space and sub-virtual spaces to achieve unified display of audio data.

Benefits of technology

It achieves high efficiency and scalability in large-scale audio interaction scenarios, avoids performance bottlenecks and resource limitations, ensures the continuity and consistency of user experience, and improves the system's processing power and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940707A_ABST
    Figure CN121940707A_ABST
Patent Text Reader

Abstract

The invention discloses an audio processing method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of multimedia, and the method comprises the steps: monitoring the number of first interaction objects in a main virtual space in a first audio interaction round; determining an audio interaction mode corresponding to the number of the first interaction objects based on a comparison result of the number of the first interaction objects and a preset threshold value; if the audio interaction mode is a multi-space mode, creating at least one sub-virtual space associated with the main virtual space; and merging the first audio data corresponding to the main virtual space and the second audio data corresponding to each sub virtual space, and displaying the merged target audio data in a display interface. By implementing the technical scheme of the invention, a plurality of virtual spaces can be dynamically created according to the number of real-time interaction people to share the audio processing pressure, and the audio data of each space is merged and then displayed in a unified manner, thereby supporting larger-scale real-time audio interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimedia technology, specifically to audio processing methods, apparatus, electronic devices, storage media, and program products. Background Technology

[0002] With the booming development of virtual reality, augmented reality, and online entertainment technologies, online choral singing based on virtual space has become an emerging form of interactive entertainment, gaining popularity among users. In this application, a main virtual space is typically used as a performance stage, allowing a large number of users to join as chorus members, singing the same song together, thereby gaining an immersive sense of collective participation and social experience.

[0003] Currently, to achieve this choral effect, the common practice is to gather all participants in a unified virtual space (or real-time communication room). Within this space, the audio data of all participants is collected, processed, and mixed to simulate a realistic choral atmosphere. However, as the user base of virtual space applications continues to expand, the number of users wishing to participate in a simultaneous chorus during popular songs or large online events may increase dramatically, potentially reaching tens of thousands or even more. Current technological architectures struggle to smoothly support such a massive influx of users simultaneously engaging in high-quality choral interaction while maintaining real-time performance and stability when dealing with this sudden surge in demand. Summary of the Invention

[0004] In view of this, this application provides an audio processing method, apparatus, electronic device, storage medium, and program product to solve the problem that current virtual space chorus solutions cannot simultaneously accommodate large-scale users and provide a high-quality interactive experience.

[0005] In a first aspect, this application provides an audio processing method, comprising: monitoring the number of first interactive objects in a main virtual space during a first audio interaction round; determining the audio interaction mode corresponding to the number of first interactive objects based on a comparison result between the number of first interactive objects and a preset threshold; if the audio interaction mode is a multi-space mode, creating at least one sub-virtual space associated with the main virtual space; merging the first audio data corresponding to the main virtual space and the second audio data corresponding to each sub-virtual space, and displaying the merged target audio data in a display interface.

[0006] In one optional implementation, creating at least one sub-virtual space associated with the main virtual space includes: obtaining a second audio interaction round corresponding to a first audio interaction round; and creating at least one sub-virtual space during the switching interval between the first audio interaction round and the second audio interaction round.

[0007] In one alternative implementation, creating at least one sub-virtual space includes: obtaining a baseline number of virtual space objects and space creation parameters; determining a target number of sub-virtual spaces to be created based on the ratio of the number of first interactive objects to the baseline number of virtual space objects; and creating the target number of sub-virtual spaces based on the space creation parameters.

[0008] In one alternative implementation, multiple second interactive objects that are added to the second audio interaction round are obtained; the multiple second interactive objects are assigned to the main virtual space and each sub-virtual space.

[0009] In one optional implementation, the third audio interaction round corresponding to the second audio interaction round is obtained; if the number of objects of the second interactive object is lower than a preset threshold, each sub-virtual space is closed during the switching interval between the second audio interaction round and the third audio interaction round; at least one third interactive object added to the third audio interaction round is assigned to the main virtual space.

[0010] In one optional implementation, the audio interaction mode corresponding to the number of first interactive objects is determined based on the comparison result between the number of first interactive objects and a preset threshold, including: in response to the number of first interactive objects exceeding the preset threshold, determining the audio interaction mode corresponding to the number of first interactive objects as a multi-space mode; in response to the number of first interactive objects not exceeding the preset threshold, determining the audio interaction mode corresponding to the number of first interactive objects as a single-space mode.

[0011] In one optional implementation, the method for determining the preset threshold includes: obtaining the number of virtual space reference objects; and determining a preset multiple of the number of virtual space reference objects as the preset threshold.

[0012] In one optional implementation, the first audio data includes the number of first target interactive objects and first audio interaction information corresponding to the main virtual space; the second audio data includes the number of second target interactive objects and second audio interaction information corresponding to each sub-virtual space; merging the first audio data corresponding to the main virtual space and the second audio data corresponding to each sub-virtual space, and displaying the merged target audio data on the display interface, includes: merging the number of first target interactive objects with the number of each second target interactive object to obtain an interactive object merging result; merging the first audio interaction information with each second audio interaction information to obtain an audio interaction information merging result; and displaying the interactive object merging result and the audio interaction information merging result on the display interface.

[0013] Secondly, this application provides an audio processing device, comprising: a monitoring module for monitoring the number of first interactive objects in a main virtual space during a first audio interaction round; a determining module for determining an audio interaction mode corresponding to the number of first interactive objects based on a comparison result between the number of first interactive objects and a preset threshold; a creating module for creating at least one sub-virtual space associated with the main virtual space if the audio interaction mode is a multi-space mode; and a merging module for merging the first audio data corresponding to the main virtual space and the second audio data corresponding to each sub-virtual space, and displaying the merged target audio data on a display interface.

[0014] Thirdly, this application provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the audio processing method of the first aspect or any corresponding embodiment described above.

[0015] Fourthly, this application provides a computer-readable storage medium storing computer instructions for causing a computer to perform the audio processing method of the first aspect or any corresponding embodiment described above.

[0016] Fifthly, this application provides a computer program product, including computer instructions for causing a computer to execute the audio processing method described in the first aspect or any corresponding embodiment thereof.

[0017] The audio processing method provided in this application achieves high efficiency and scalability in handling large-scale audio interaction scenarios by dynamically monitoring the number of interactive objects and intelligently switching audio interaction modes. This application monitors the number of interactive objects in the main virtual space in real time and compares it with a preset threshold to automatically decide whether to enter a multi-space mode. When the number of interactive objects exceeds the threshold, it dynamically creates sub-virtual spaces to distribute the load, avoiding performance bottlenecks and resource limitations that may occur with a single virtual space. Simultaneously, by merging and uniformly displaying the audio data from the main and sub-virtual spaces, it ensures that the audio experience perceived by the front-end user is always consistent and seamless, without needing to concern themselves with the technical details of the backend. This application not only improves the processing power and stability of the audio system but also optimizes resource utilization efficiency, enabling it to flexibly adapt to audio interaction needs of different scales, balancing the simplicity of small-scale scenarios with the scalability of large-scale scenarios. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this application, the drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this application; Figure 2 This is a schematic flowchart of a first audio processing method according to an embodiment of this application; Figure 3 This is a schematic diagram of a second type of audio processing method according to an embodiment of this application; Figure 4 This is a structural block diagram of an audio processing apparatus according to an embodiment of this application; Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0022] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0023] As one optional application scenario in this application embodiment, Figure 1 A schematic diagram illustrating an application scenario of an audio processing system is shown. For example... Figure 1 As shown, the system may include at least one terminal device and at least one server. Figure 1The system is illustrated in the example, which includes a computer 101, a mobile terminal 102, and a server 103, and the terminal devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 through a network 110.

[0024] Specifically, the terminal device can be a smartphone, tablet, laptop, PDA, desktop computer, game console, smart TV, smart wearable device, in-vehicle terminal, VR (Virtual Reality) device, AR (Augmented Reality) device, etc. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranet, local area network, wide area network, mobile communication network, and combinations thereof.

[0025] Taking online karaoke or live duet streaming as an example, the terminal device has a corresponding application installed. By running this application, users can enter a virtual online room (i.e., the main virtual space). In this room, some users can perform as lead singers or guests, while other users can apply to "go on stage" and become interactive participants, singing along with the lead singer to create audio interaction. The audio data of each interactive participant is collected and uploaded.

[0026] Popular interactive karaoke or live streaming apps typically aim to support as many users as possible singing together simultaneously to create a lively, collaborative atmosphere. However, current solutions usually place all participants in a single audio communication room. When the number of simultaneous participants surges, a single audio room faces challenges such as connection limits, server performance bottlenecks, and high operating costs. This makes it impossible to support ultra-large-scale scenarios with tens of thousands or even more people singing together at the same time, limiting business expansion and user experience upgrades.

[0027] The audio processing method provided in this application monitors the number of interactive objects in the main virtual space during an audio interaction cycle (such as the performance of a song) in real time, and intelligently decides whether to enter a multi-space mode based on a preset threshold. When the number of interactive objects exceeds the threshold, multiple sub-virtual spaces associated with the main virtual space are automatically created, and subsequently added interactive objects are reasonably distributed to these spaces, thereby dispersing the load pressure on individual spaces. Finally, the audio data generated by all spaces are merged and processed, and displayed as a unified audio stream to all users. This method achieves dynamic scaling and efficient utilization of system resources, seamlessly supporting various choral scenarios ranging from dozens to tens of thousands of people. While breaking through technical bottlenecks, it provides users with a large-scale and unified collective interactive experience.

[0028] According to an embodiment of this application, an audio processing method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0029] This embodiment provides an audio processing method that can be used in electronic devices, such as server 103. Figure 2 This is a flowchart of an audio processing method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps: Step S201: Monitor the number of first interactive objects in the main virtual space during the first audio interaction round.

[0030] The first audio interaction round refers to the complete performance of a song. In a live chorus room, each playback, performance, and ending of a song is considered an independent audio interaction round. The server monitors the number of interactive objects (i.e., users singing on the microphone) within this round and performs data statistics and decisions at the end of the round. The main virtual space refers to the main RTC room (main room), which is the core control center of the live stream. The main virtual space is created by the lead singer (host) to synchronize song control commands (such as play, pause, progress) and serves as the initial space for early chorus members to join. Specifically, by embedding a listening mechanism at the technical level of the main virtual space, key events that characterize changes in user interaction status are captured in real time. When a user joins (goes on the microphone) or leaves (goes off the microphone) the main virtual space, the server immediately receives the corresponding event notification. By listening to and counting these events, a real-time counter reflecting the current number of interactive objects on the microphone can be dynamically maintained. Throughout the entire audio interaction round, the value changes of this counter are continuously tracked, and the peak value that occurs within that round is recorded, i.e., the maximum number of interactive objects online simultaneously. This peak value is defined as the number of first interactive objects, which serves as a key indicator for assessing system load and deciding on the next course of action.

[0031] Step S202: Based on the comparison result between the number of first interactive objects and the preset threshold, determine the audio interaction mode corresponding to the number of first interactive objects.

[0032] The first interactive object count refers to the peak number of users singing along in the main virtual space during the first round of audio interaction. Audio interaction mode refers to the server's operating mode for handling audio interactions, which can include single-room mode, multi-room mode, etc. Specifically, one or more key performance and capacity thresholds are preset. These thresholds are predefined by the platform based on its technical architecture, resource costs, and user experience goals. The core of the decision-making logic is to compare the monitored first interactive object count (peak) with these preset thresholds. This comparison process follows a clear decision tree; if the peak exceeds a certain threshold range, the multi-room mode will be activated. The fundamental purpose of this mode switching is to ensure system stability, performance, and cost controllability through the elastic scaling of the architecture when the interaction scale expands.

[0033] Step S203: If the audio interaction mode is a multi-space mode, then create at least one sub-virtual space associated with the main virtual space.

[0034] Multi-space mode refers to the mode adopted when the server determines that rooms need to be split. Sub-virtual spaces refer to auxiliary RTC rooms associated with the main virtual space. Specifically, once the server decides to enter multi-space mode, it calculates the number of sub-virtual spaces to be created based on the peak number of interactive objects monitored above, using a predefined algorithm. Subsequently, the server dynamically instantiates the corresponding number of sub-virtual spaces in its resource pool. These sub-virtual spaces do not exist in isolation; they establish logical subordinate or associated relationships with the main virtual space, forming a space group. To ensure the consistency of the entire interactive experience, the server establishes a broadcast channel from the main virtual space to all sub-virtual spaces for core control signaling (such as play, pause, etc.). In this way, the main virtual space acts as the "brain" initiating the command, while each sub-virtual space acts as the "limbs" executing the command, working collaboratively.

[0035] Specifically, the synchronization path is implemented through a centralized signaling relay server. The lead client in the main virtual space sends control signals (such as play, pause, and progress commands) to the signaling relay server; each sub-virtual space deploys a signaling synchronization module (such as a headless robot client) that maintains a persistent connection with the signaling relay server. Upon receiving signals from the main virtual space, the signaling relay server uniformly and concurrently distributes them to the signaling synchronization modules in all sub-virtual spaces, ensuring synchronized commands and low latency across all rooms.

[0036] Step S204: Merge the first audio data corresponding to the main virtual space and the second audio data corresponding to each sub-virtual space, and display the merged target audio data on the display interface.

[0037] The first audio data characterizes the user interaction status occurring within the main virtual space. The second audio data characterizes the user interaction status occurring within each sub-virtual space. The target audio data refers to the overall audio data resulting from merging the first audio data and all second audio data; this is the final output of the entire data processing flow. The display interface refers to the live streaming interface on the user's end (such as a mobile app, webpage, etc.). Specifically, in the multi-space mode, user interaction status information is generated in a distributed manner. The main virtual space generates a status summary (i.e., the first audio data) about its internal interactive objects. This data package contains key information such as the number of online users, activity level, and collective score within the space. Simultaneously, each sub-virtual space independently generates a status summary (i.e., the second audio data) describing the status of its own internal user group. Its content is similar to the first audio data, but it only reflects the local situation of that sub-space. To present all participants with a unified global view of the entire interactive activity, the server establishes a centralized data aggregation service. This service is responsible for receiving these status data packages from the main virtual space and all sub-virtual spaces. Next, using data aggregation and fusion technology, these scattered, localized state data are merged into a single, comprehensive target audio data set. This target audio data can be understood as a composite data object that encapsulates the overall interactive state. Finally, this merged target audio data set, reflecting the overall situation, is passed to the display interface. The interface layer parses this data object and presents it to all users in an intuitive format (such as numbers, charts, and animations). In this way, although users are technically dispersed across multiple independent virtual spaces, what they perceive on the display interface is still a unified whole with numerous participants and a vibrant interactive atmosphere.

[0038] The audio processing method provided in this application achieves high efficiency and scalability in handling large-scale audio interaction scenarios by dynamically monitoring the number of interactive objects and intelligently switching audio interaction modes. This application monitors the number of interactive objects in the main virtual space in real time and compares it with a preset threshold to automatically decide whether to enter a multi-space mode. When the number of interactive objects exceeds the threshold, it dynamically creates sub-virtual spaces to distribute the load, avoiding performance bottlenecks and resource limitations that may occur with a single virtual space. Simultaneously, by merging and uniformly displaying the audio data from the main and sub-virtual spaces, it ensures that the audio experience perceived by the front-end user is always consistent and seamless, without needing to concern themselves with the technical details of the backend. This application not only improves the processing power and stability of the audio system but also optimizes resource utilization efficiency, enabling it to flexibly adapt to audio interaction needs of different scales, balancing the simplicity of small-scale scenarios with the scalability of large-scale scenarios.

[0039] This embodiment provides an audio processing method that can be used in electronic devices, such as server 103. Figure 3 This is a flowchart of an audio processing method according to an embodiment of this application, such as... Figure 3 As shown, the process includes the following steps: Step S301: Monitor the number of first interactive objects within the main virtual space during the first audio interaction round. For details, please refer to [link to details]. Figure 2 Step S201 of the illustrated embodiment will not be described again here.

[0040] Step S302: Based on the comparison result between the number of first interactive objects and the preset threshold, determine the audio interaction mode corresponding to the number of first interactive objects.

[0041] Specifically, step S302 includes: in response to the number of first interactive objects exceeding a preset threshold, determining that the audio interaction mode corresponding to the number of first interactive objects is a multi-space mode; in response to the number of first interactive objects not exceeding the preset threshold, determining that the audio interaction mode corresponding to the number of first interactive objects is a single-space mode.

[0042] The number of the first interactive objects detected is compared with a predefined threshold. This threshold represents the approximate capacity limit that a single virtual space can handle while maintaining a good experience and technical performance. When the server determines through logical comparison that the number of the first interactive objects has exceeded this threshold, it automatically triggers a state switch. Based on this, the server determines that the current scale of interaction is no longer suitable for continuing within a single space, and a distributed architecture must be introduced to share the load, thus officially establishing the audio interaction mode as a multi-space mode. The essence of this decision is the horizontal scaling strategy adopted by the server to maintain stability and experience after sensing capacity pressure.

[0043] Single-space mode refers to an operational mode where all interactive objects (chorus members) are concentrated in the same main virtual space for real-time audio interaction. Specifically, the number of the first interactive objects is compared with a preset threshold. When the logic determines that the number is lower than or equal to the threshold, it indicates that the current scale of interaction is well within the design capacity and performance range of the single virtual space. Therefore, the server decides not to start a complex multi-space architecture and continues to keep all interactive objects concentrated in the same main virtual space. Determining the audio interaction mode as single-space mode means that the server believes the current centralized architecture is sufficient to handle the current load efficiently and economically, thus choosing a simpler and more direct technical solution.

[0044] The audio processing method provided in this application establishes a clear, definite, and fully automated decision-making mechanism by using a preset threshold as a clear and unique judgment criterion for switching between single-space and multi-space modes. When the number of interactive objects exceeds the threshold, the server system automatically enables multi-space mode to cope with high-concurrency scenarios without manual intervention, ensuring the server system's powerful scalability. When the number does not exceed the threshold, single-space mode is automatically adopted, avoiding unnecessary architectural complexity and achieving optimal resource utilization. This binary judgment logic based on a single threshold makes the selection of server system architecture mode extremely efficient and reliable. It ensures both the simplicity and resource economy of the system when the load is low, and ensures that the system can promptly launch an expanded architecture to maintain performance stability when the load increases. This allows the entire system to switch accurately and promptly between two optimal states according to the actual load.

[0045] In some optional implementations, the method for determining the preset threshold includes: obtaining the number of virtual space reference objects; and determining a preset multiple of the number of virtual space reference objects as the preset threshold.

[0046] The baseline number of virtual space objects refers to the system's preset baseline expected number of users (N). It represents the ideal number of interactive objects that a virtual space (RTC room) can accommodate while ensuring the best interactive experience and technical performance. This value is the benchmark unit for capacity planning and splitting decisions. The preset multiplier is a multiplier coefficient set to calculate the room splitting trigger threshold. It is a configurable system parameter, such as 1.5, 2.0, or 2.2, allowing the platform to dynamically adjust the splitting threshold based on real-time performance monitoring, user experience feedback, or operational needs. This configurability avoids the flexibility limitations of hard-coding, enabling the server to adapt to load changes in different scenarios. Specifically, the baseline number of virtual space objects is a core parameter predefined and configured by the system administrator or platform based on technical architecture capabilities, user experience goals, and cost-benefit analysis. It is typically stored in a configuration file or a centralized parameter server. The server obtains this baseline value by accessing this storage location. Subsequently, a multiplication operation logic is invoked to multiply the baseline number of objects by another equally configurable preset multiplier. This preset multiplier is an empirical amplification factor used to provide the server with a certain buffer space, avoiding frequent mode switching near capacity limits. The result of multiplication is determined as the preset threshold used for mode switching decisions.

[0047] In the above implementation, a highly flexible and configurable threshold decision-making mechanism is constructed by directly associating the preset threshold with a basic virtual space baseline number of objects and its preset multiple. This method makes the trigger point for server system expansion no longer a fixed value, but can be flexibly scaled based on a verified baseline capacity. By adjusting the single parameter of the preset multiple, system administrators or algorithms can easily calibrate the server system's sensitivity and timing in responding to load increases without redefining the entire evaluation system. This setting method simplifies the configuration process and gives the server system strong policy adaptability, enabling it to accurately match expansion needs under different business scenarios, performance requirements, or resource constraints. Thus, while ensuring system responsiveness, it achieves simplicity and high customizability in policy management.

[0048] Step S303: If the audio interaction mode is a multi-space mode, then create at least one sub-virtual space associated with the main virtual space.

[0049] Specifically, creating at least one sub-virtual space associated with the main virtual space includes: Step a1: Obtain the second audio interaction round corresponding to the first audio interaction round.

[0050] The second audio interaction round refers to the next audio interaction cycle that immediately follows the first audio interaction round, such as the complete performance of the second song after the first song. The server typically prepares for resource allocation for the next round at the end of the previous round. Specifically, in business logic, audio interaction rounds (such as songs) are usually arranged in a predetermined order. The server internally maintains a state machine or program list representing the current progress. Once the first audio interaction round (such as the song currently playing or just finished) is identified, the immediately following second audio interaction round (i.e., the next song to be performed) can be logically obtained by querying its internal state machine or program list.

[0051] Step a2: During the switching interval between the first audio interaction round and the second audio interaction round, at least one sub-virtual space is created.

[0052] The server precisely utilizes the natural intervals between two audio interaction rounds, such as the preparation time between the end of one song and the start of the next. Within this time window, the server's room management module is triggered. Based on the needs determined in the decision-making phase (such as the number of sub-virtual spaces to be created), this module issues instructions to the underlying infrastructure services (such as the API of the RTC service provider) to request the instantiation of one or more new sub-virtual spaces logically associated with the main virtual space. The choice of this timing is crucial, ensuring that all technical adjustments and resource preparations are completed within a wait period imperceptible to the user, achieving a seamless switchover.

[0053] The audio processing method provided in this application cleverly utilizes the inherent, natural intervals in the business logic for system reconstruction by precisely setting the creation timing of the sub-virtual space within the switching interval between the first and second audio interaction rounds. This timing arrangement ensures that all spatial structure adjustment operations are completed during gaps in user-free interaction, thus achieving zero disruption to the user's audio interaction process. Users are unaware of the resource allocation and space creation processes performed in the background, effectively avoiding experience interruptions, audio stuttering, or synchronization issues that may occur when adjusting the system architecture during interaction. By deeply integrating key system management tasks with the business rhythm, this application maximizes the smoothness and consistency of the user experience while ensuring system scalability.

[0054] In some alternative implementations, at least one sub-virtual space is created, including: Step b1: Obtain the number of virtual space baseline objects and space creation parameters.

[0055] Space creation parameters refer to a set of configuration information required when creating a sub-virtual space. These parameters define the technical specifications and behavior of the sub-virtual space, and are typically used to quickly and consistently create rooms by combining pre-configured templates with dynamic parameters. Specifically, this operation is accomplished by calling its configuration management service. The virtual space baseline object count is read as a basic capacity configuration. Simultaneously, the space creation parameters are also retrieved. These parameters may include a series of metadata defining the technical characteristics of the virtual space, such as the room template ID, audio and video encoding formats, permission settings, and mixing strategies. This information is usually centrally stored in a database or configuration center for unified management and rapid retrieval, ensuring that the created virtual spaces meet the platform's standardization requirements.

[0056] Step b2: Determine the target number of sub-virtual spaces to be created based on the ratio of the number of first interactive objects to the number of virtual space baseline objects.

[0057] The target number refers to the specific number (M) of sub-virtual spaces that the server needs to create based on the current load. The formula is typically: M = ceil(Number of first interactive objects / Base number of virtual space objects), which rounds up based on the peak number of users and the base capacity from the previous round to ensure that all users can be accommodated. Specifically, the number of first interactive objects is used as the numerator, and the base number of virtual space objects (N) is used as the denominator, and a division operation is performed. The quotient theoretically indicates how many fully loaded base spaces are needed to accommodate all these users. However, since users cannot be divided, the server applies a rounding function to this quotient. This means that even if the quotient is not an integer, it will be rounded up to the next largest integer. This final integer value is the target number (M), which ensures that the number of sub-virtual spaces created is sufficient.

[0058] Step b3: Based on the space creation parameters, create the target number of sub-virtual spaces.

[0059] Once the target number of sub-virtual spaces to be created is determined, the server's room management module enters a loop or batch processing flow. For each sub-virtual space to be created, the module uses the acquired space creation parameters as a unified configuration template. It repeatedly calls the underlying infrastructure's creation interface a target number of times, passing these space creation parameters in with each call. This efficiently and consistently creates a set of sub-virtual spaces with identical technical specifications, preparing them for subsequent user allocation. This process emphasizes automation and standardization to achieve rapid, elastic resource scaling.

[0060] In the above implementation, by introducing two key factors—the number of virtual space baseline objects and space creation parameters—the creation process of sub-virtual spaces achieves a high degree of quantification and automation. The target number of required sub-virtual spaces is scientifically determined based on the precise ratio of the number of first interactive objects to the number of baseline objects. This calculation logic ensures that resource allocation closely matches the actual user scale requirements, avoiding performance bottlenecks caused by insufficient space and preventing resource waste caused by over-creation. Simultaneously, batch creation based on space creation parameters ensures the consistency and standardization of the basic configuration of all sub-virtual spaces, improving the manageability and maintainability of the system. This application transforms resource allocation from experience-based decision-making to an automated process based on explicit rules and parameters, thereby significantly improving the accuracy, efficiency, and overall resource utilization of system expansion.

[0061] In some alternative implementations, creating at least one sub-virtual space associated with the main virtual space, or creating at least one sub-virtual space, further includes: Step c1: Obtain multiple second interactive objects that will be added to the second audio interaction round.

[0062] The second interactive object refers to the interactive object that newly joins the chorus during the second audio interaction round. Specifically, as the second audio interaction round begins (i.e., the next song), user join requests are continuously generated. The server senses and captures these new join requests in real time through its user access service or gateway. These requests contain metadata such as user identifiers and session information. By listening to these access events and aggregating the identifiers of all users who initiated requests and passed verification during the second audio interaction round, the server logically obtains multiple second interactive objects. This is essentially a dynamic registration and list maintenance process for all legitimate new users by the server.

[0063] Step c2: Assign multiple second interactive objects to the main virtual space and each sub-virtual space.

[0064] The server first attempts to redirect newly joined secondary interactive objects to the main virtual space. It monitors the number of online users in the main virtual space in real time, and as long as the number of users has not reached or is close to the preset single-room capacity limit (i.e., the baseline expected number of users N), new users will continue to be assigned to the main virtual space. Only when the server detects that the number of users in the main virtual space is about to reach (or has reached) the capacity threshold N will subsequent newly joined users be switched to the next available sub-virtual space, and the filling of that sub-space will begin. This "fill-trigger-switch" process continues, filling each sub-virtual space in turn. In this way, the server ensures that, in most cases, each created room operates close to its ideal capacity, thereby supporting a large number of users while minimizing the number of active rooms, optimizing resource utilization and service costs.

[0065] Alternatively, an average distribution strategy can be adopted, distributing newly added interactive objects as evenly as possible across the main virtual space and various sub-virtual spaces. This includes treating the main virtual space itself as a sub-room for distribution, ensuring load balancing across rooms, avoiding overload in any single room, and optimizing resource utilization.

[0066] In the above implementation, by actively allocating multiple second interactive objects to the main virtual space and the already created sub-virtual spaces at the start of a new round of audio interaction, the user load is actively distributed among different virtual spaces. This allocation mechanism ensures that newly added interactive objects do not flood into a single virtual space, but are effectively distributed across multiple processing units of the system, thereby avoiding problems such as overload, performance degradation, or resource contention in a single space from the source. This application maintains the balance of the overall system load through system-level intelligent scheduling, providing a fundamental guarantee for a stable and smooth audio interaction experience under large-scale concurrent user participation, and enhancing the system's concurrent processing capabilities and overall robustness.

[0067] In some alternative implementations, creating at least one sub-virtual space associated with the main virtual space, or creating at least one sub-virtual space, further includes: Step d1: Obtain the third audio interaction round corresponding to the second audio interaction round.

[0068] The third audio interaction round refers to the next audio interaction cycle immediately following the second audio interaction round, such as the performance of the third song. After each round, the server assesses the overall scale and may adjust the architecture again. Specifically, this information is obtained from an activity sequence state machine or program list maintained internally by the server. When the second audio interaction round is identified as currently active or recently ended, the server logically retrieves the following third audio interaction round by querying the predefined next item in the state machine or program list. This ensures that the server can proactively prepare for the next interaction cycle.

[0069] Step d2: If the number of objects in the second interactive object is lower than a preset threshold, then each sub-virtual space is closed during the switching interval between the second and third audio interactive rounds.

[0070] At the end of the second audio interaction round, the overall interaction scale of that round is evaluated (i.e., the number of objects in the second interaction, referring to the peak number of people). This value is compared with a preset shrinkage threshold (such as 2N mentioned above). If it is determined that the current scale has fallen below the threshold, it indicates that a multi-space architecture is no longer needed to support it. Therefore, during the interval between the second and third rounds, a resource cleanup process will be triggered. The room management module will issue instructions to the underlying infrastructure to close or destroy all sub-virtual spaces created for traffic splitting, releasing related computing and network resources to prepare for a return to single-room mode.

[0071] The preset threshold for the shrinking operation can be the same as the splitting threshold (e.g., 2N), or it can be a separately configured shrinking threshold. When the peak number of users drops below this threshold, the server automatically triggers the shrinking process during song switching intervals, closing all sub-virtual spaces and returning to single-room mode to achieve elastic resource release.

[0072] Step d3 involves assigning at least one third interactive object, which is added to the third audio interaction round, to the main virtual space.

[0073] The third interactive object refers to the new interactive object that joins the chorus during the third audio interaction round. Specifically, after the server performs a shrinking operation, the architecture has reverted to a single-room mode. At this point, the server's allocation strategy for the third interactive object added in the third audio interaction round becomes extremely simple and direct. Since the sub-virtual spaces have been closed, the main virtual space becomes the only available target. Therefore, the server will direct all new user access requests to the main virtual space without exception. This is typically achieved by returning a uniform access configuration pointing to the main virtual space to all third interactive objects.

[0074] In the above implementation, by continuously monitoring the number of interactive objects and automatically closing all sub-virtual spaces and converging subsequent interactive objects into the main virtual space during the switching interval before the start of a new round of interaction when the number falls below a preset threshold, a complete, bidirectional, elastic system scaling mechanism is achieved. This mechanism ensures that system resources can be dynamically reclaimed and centralized according to real-time load. When the scale of interaction shrinks, by simplifying the system architecture and closing redundant sub-virtual spaces, the idleness and waste of computing, storage, and network resources are effectively avoided. At the same time, this convergence process is also completed during business downtime, ensuring that resource release operations do not interfere with the normal interactive experience of users, achieving a smooth and seamless transition between the two states of system expansion and contraction, and ultimately achieving a balance between efficient resource utilization and a stable user experience.

[0075] Step S304: Merge the first audio data corresponding to the main virtual space and the second audio data corresponding to each sub-virtual space, and display the merged target audio data on the display interface.

[0076] Specifically, the first audio data includes the number of first target interactive objects corresponding to the main virtual space and the first audio interaction information; the second audio data includes the number of second target interactive objects corresponding to each sub-virtual space and the second audio interaction information; the above step S304 includes: Step S3041: Combine the number of first target interactive objects with the number of each second target interactive object to obtain the interactive object merging result.

[0077] The first target number of interactive objects refers to the number of interactive objects collected and counted from the main virtual space for display. It represents the determined number of chorus members within the main virtual space. The second target number of interactive objects refers to the number of interactive objects collected and counted from each sub-virtual space for display. Each sub-space reports its own number of members. The final result of the interactive object aggregation is the total number of interactive objects obtained by summing the first target number of interactive objects in the main virtual space with the second target number of interactive objects from all sub-virtual spaces. Specifically, the server's data aggregation module collects the first target number of interactive objects reported from the main virtual space (i.e., the number of people in the main room) and the second target number of interactive objects reported from each sub-virtual space (i.e., the number of people in each sub-room). Then, the module performs a simple arithmetic addition, adding the number of people in the main room to the number of people in each sub-room. The final sum is the final result of the interactive object aggregation, representing the scale of all interactive objects, such as the total number of chorus members.

[0078] Step S3042: Merge the first audio interaction information with each of the second audio interaction information to obtain the audio interaction information merging result.

[0079] The first audio interaction information refers to the interaction data collected from the main virtual space, excluding the number of users, such as the cumulative score, average activity level, and chorus accuracy of all users in that space. The second audio interaction information refers to the interaction data collected from each sub-virtual space, excluding the number of users. The audio interaction information merging result refers to the global interaction data obtained by fusing the first audio interaction information from the main virtual space with the second audio interaction information from all sub-virtual spaces. Specifically, in addition to the number of users, the server also collects the first audio interaction information from the main virtual space (such as the total score of the main room) and the second audio interaction information from each sub-virtual space (such as the total score of each sub-room). The data aggregation module merges these information based on their properties: for additive indicators (such as total score, total number of flowers, etc.), it sums them up; for indicators that require averaging (such as average accuracy), it performs a weighted average based on the number of users in each room. Through this targeted data fusion rule, a unified audio interaction information merging result is ultimately generated, reflecting the overall performance of the entire interactive activity.

[0080] Step S3043: Display the result of merging interactive objects and the result of merging audio interactive information on the display interface.

[0081] The front-end display interface maintains communication with the back-end data service. When the interface receives the combined results of interactive objects (such as the total number of people) and audio interaction information (such as the total score) pushed from the back-end, it calls its rendering engine to bind these data results to preset UI components. For example, the total number of people is displayed in a prominent counter, and the total score is displayed on a progress bar or leaderboard. Through this unified visual presentation, all users (regardless of which technical room they are in) see unified data about the entire large group on the display interface, thereby eliminating the sense of isolation caused by technical fragmentation and creating a strong sense of unity.

[0082] In some alternative implementations, data aggregation employs a hybrid mechanism of sub-room reporting and periodic fetching from a central server. Sub-virtual spaces periodically (e.g., every 5 seconds) report key data (such as online users, scores, etc.) to the central server, and immediately report any changes in key data (e.g., when a room is full) to ensure timeliness. Simultaneously, the central server proactively fetches data from all sub-virtual spaces periodically (e.g., every 30 seconds) for backup and consistency verification. This mechanism balances real-time requirements with system load.

[0083] The audio processing method provided in this application constructs a clear and systematic data aggregation process by explicitly distinguishing and specifically merging the number of first target interactive objects and first audio interaction information from the main virtual space with the number of second target interactive objects and second audio interaction information from each sub-virtual space. This application accumulates all the interactive object numbers to present a unified, overall participation scale on the display interface, effectively creating a sense of collective atmosphere. Simultaneously, it integrates all independent audio interaction information to generate a comprehensive result representing the entire interactive activity. This structured merging and display mechanism ensures that although users are technically assigned to different virtual spaces for interaction, from the user's perspective, the key data received is always complete and consistent. This successfully distributes the load in the technical implementation while maintaining an inseparable overall activity perception in the user experience, achieving effective decoupling and perfect integration between the technical architecture and user perception.

[0084] This embodiment also provides an audio processing apparatus for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0085] This embodiment provides an audio processing device, such as... Figure 4 As shown, it includes: Monitoring module 401 is used to monitor the number of first interactive objects in the main virtual space during the first audio interaction round; The determining module 402 is used to determine the audio interaction mode corresponding to the number of first interactive objects based on the comparison result between the number of first interactive objects and a preset threshold. Create module 403 to create at least one sub-virtual space associated with the main virtual space if the audio interaction mode is a multi-space mode; The merging module 404 is used to merge the first audio data corresponding to the main virtual space and the second audio data corresponding to each sub-virtual space, and display the merged target audio data on the display interface.

[0086] In some alternative implementations, creation module 403 includes: The first acquisition submodule is used to acquire the second audio interaction round corresponding to the first audio interaction round; Create a submodule to create at least one sub-virtual space during the switching interval between the first and second audio interaction rounds.

[0087] In some alternative implementations, creating a submodule includes: The first acquisition unit is used to acquire the number of virtual space baseline objects and space creation parameters; The determining unit is used to determine the target number of sub-virtual spaces to be created based on the ratio of the number of first interactive objects to the number of virtual space baseline objects; Create a unit to create a target number of sub-virtual spaces based on spatial creation parameters.

[0088] In some alternative implementations, creating module 403 or creating submodules includes: The second acquisition unit is used to acquire multiple second interactive objects that are added to the second audio interaction round; The first allocation unit is used to allocate multiple second interactive objects to the main virtual space and various sub-virtual spaces.

[0089] In some alternative implementations, creating module 403 or creating submodules includes: The third acquisition unit is used to acquire the third audio interaction round corresponding to the second audio interaction round; The closing unit is used to close each sub-virtual space during the switching interval between the second audio interaction round and the third audio interaction round if the number of objects of the second interactive object is lower than a preset threshold. The second allocation unit is used to allocate at least one third interactive object that has been added to the third audio interactive round to the main virtual space.

[0090] In some alternative implementations, the determining module 402 includes: The first determining submodule is used to determine the audio interaction mode corresponding to the number of first interactive objects as a multi-space mode in response to the number of first interactive objects exceeding a preset threshold. The second determining submodule is used to determine the audio interaction mode corresponding to the number of first interactive objects as single-space mode in response to the fact that the number of first interactive objects does not exceed a preset threshold.

[0091] In some alternative implementations, the determining module 402 further includes: The second acquisition submodule is used to acquire the number of virtual space baseline objects; The third determining submodule is used to determine the preset threshold as a preset multiple of the number of virtual space reference objects.

[0092] In some optional implementations, the first audio data includes the number of first target interactive objects corresponding to the main virtual space and the first audio interaction information; the second audio data includes the number of second target interactive objects corresponding to each sub-virtual space and the second audio interaction information; the merging module 404 includes: The first merging submodule is used to merge the number of first target interactive objects with the number of each second target interactive object to obtain the interactive object merging result; The second merging submodule is used to merge the first audio interaction information with each of the second audio interaction information to obtain the audio interaction information merging result; The display submodule is used to show the merged results of interactive objects and the merged results of audio interactive information on the display interface.

[0093] The audio processing apparatus provided in this application can execute the audio processing method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0094] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0095] The following is a detailed reference. Figure 5The diagram illustrates a structural schematic suitable for implementing the electronic device described in the embodiments of this application. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from memory 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0096] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0097] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a memory 508, or installed from a ROM 502. When the computer program is executed by the processor 501, it performs the functions defined in the audio processing method of the embodiments of this application.

[0098] Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0099] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the audio processing method shown in the above embodiments is implemented.

[0100] A portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0101] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.

Claims

1. An audio processing method, characterized in that, The method includes: Monitor the number of first interactive objects within the main virtual space during the first round of audio interaction; Based on the comparison between the number of the first interactive objects and the preset threshold, the audio interaction mode corresponding to the number of the first interactive objects is determined. If the audio interaction mode is a multi-space mode, then at least one sub-virtual space associated with the main virtual space is created; The first audio data corresponding to the main virtual space and the second audio data corresponding to each of the sub-virtual spaces are merged, and the merged target audio data is displayed on the display interface.

2. The method according to claim 1, characterized in that, The creation of at least one sub-virtual space associated with the main virtual space includes: Obtain the second audio interaction round corresponding to the first audio interaction round; During the switching interval between the first audio interaction round and the second audio interaction round, at least one sub-virtual space is created.

3. The method according to claim 2, characterized in that, Creating the at least one sub-virtual space includes: Obtain the number of virtual space baseline objects and space creation parameters; Based on the ratio of the number of the first interactive objects to the number of the virtual space baseline objects, the target number of the sub-virtual spaces to be created is determined; Based on the space creation parameters, create the target number of sub-virtual spaces.

4. The method according to claim 2 or 3, characterized in that, Also includes: Obtain multiple second interactive objects that will be added to the second audio interaction round; The second interactive objects are assigned to the main virtual space and each of the sub-virtual spaces.

5. The method according to claim 4, characterized in that, The method further includes: Obtain the third audio interaction round corresponding to the second audio interaction round; If the number of objects in the second interactive object is lower than the preset threshold, then each of the sub-virtual spaces is closed during the switching interval between the second audio interaction round and the third audio interaction round. At least one third interactive object will be added to the third audio interactive round and assigned to the main virtual space.

6. The method according to claim 1, characterized in that, The step of determining the audio interaction mode corresponding to the number of the first interactive objects based on the comparison result between the number of the first interactive objects and a preset threshold includes: In response to the number of the first interactive objects exceeding the preset threshold, the audio interaction mode corresponding to the number of the first interactive objects is determined to be the multi-space mode; In response to the fact that the number of the first interactive objects does not exceed the preset threshold, the audio interaction mode corresponding to the number of the first interactive objects is determined to be a single-space mode.

7. An audio processing device, characterized in that, The device includes: The monitoring module is used to monitor the number of first interactive objects in the main virtual space during the first audio interaction round; The determining module is used to determine the audio interaction mode corresponding to the number of the first interactive objects based on the comparison result between the number of the first interactive objects and the preset threshold. A creation module is used to create at least one sub-virtual space associated with the main virtual space if the audio interaction mode is a multi-space mode. The merging module is used to merge the first audio data corresponding to the main virtual space and the second audio data corresponding to each of the sub-virtual spaces, and display the merged target audio data on the display interface.

8. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the audio processing method of any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the audio processing method according to any one of claims 1 to 6.

10. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the audio processing method according to any one of claims 1 to 6.