Audio stream data processing method, device, cloud server and readable storage medium

By creating virtual audio devices in the cloud server and establishing a connection relationship with the target application, the problem of lack of flexibility in audio stream data processing in the prior art is solved, and support for complex application scenarios is achieved.

CN115801740BActive Publication Date: 2025-05-02BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211371481.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2025-05-02
Estimated Expiration
2042-11-03

AI Technical Summary

Technical Problem

The existing real-time cloud rendering technology lacks flexibility in audio streaming data processing and cannot meet the needs of complex application scenarios.

Method used

Create a virtual audio device in a cloud server, establish a connection relationship between the virtual audio device and the target application based on the audio transmission relationship of the target application, and realize flexible processing of audio stream data.

Benefits of technology

Through the use of virtual audio equipment, cloud servers can flexibly process audio streaming data, meet the needs of complex application scenarios, and improve the flexibility of audio streaming data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115801740B_ABST
    Figure CN115801740B_ABST
Patent Text Reader

Abstract

The present disclosure discloses an audio stream data processing method, device, cloud server and readable storage medium, wherein the method includes: receiving a session request initiated by a terminal device, and creating a user session matching the session request; creating a virtual audio device required by the target application in the user session according to the audio transmission relationship of one or more target applications in the user session, and establishing a connection relationship between the virtual audio device and the target application; running the target application so that the virtual audio device connected to the target application transmits audio stream data to the target application and / or captures the audio stream data output by the target application. The technical solution provided by one or more embodiments of the present disclosure can improve the flexibility of audio stream data processing in the cloud server, thereby meeting the needs of users for complex application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of Internet technology, and in particular to an audio stream data processing method, device, cloud server and readable storage medium. Background Art

[0002] With the continuous development of Internet technology, real-time cloud rendering technology, which combines real-time communication technology and high-quality graphics and image rendering algorithms, has emerged. At present, real-time cloud rendering technology has been applied in cloud games and cloud extended reality (eXtrendedReality, XR), etc.

[0003] When applying real-time cloud rendering technology, users usually only need to provide necessary input and output devices locally, without the need to provide high-performance graphics processing equipment. Figure 1 As shown, the user's local input device can capture the user's input signal and transmit it to the cloud server through real-time communication (RTC). The cloud server responds to the user's input signal, completes audio and video rendering, and transmits the rendered audio and video data back to the user terminal through RTC.

[0004] In this process, the processing of audio stream data is usually a one-way transmission from the cloud server to the user's local computer. The reason is that the application that performs audio and video rendering in the cloud server cannot directly control the user's local audio input device (such as a microphone) and audio output device (such as a speaker). This results in applications running through real-time cloud rendering technology only supporting relatively simple application scenarios and failing to meet the actual needs of users. Summary of the invention

[0005] In view of this, one or more embodiments of the present disclosure provide an audio stream data processing method, device, cloud server and readable storage medium, which can improve the flexibility of audio stream data processing in the cloud server and thus meet user needs for complex application scenarios.

[0006] On one hand, the present disclosure provides an audio stream data processing method, the method comprising: receiving a session request initiated by a terminal device, and creating a user session matching the session request; creating a virtual audio device required by one or more target applications in the user session according to an audio transmission relationship between the target application or applications in the user session, and establishing a connection relationship between the virtual audio device and the target application; and running the target application so that the virtual audio device connected to the target application transmits audio stream data to the target application and / or captures audio stream data output by the target application.

[0007] On the other hand, the present disclosure further provides an audio stream data processing device, the device comprising: a session creation unit, used to receive a session request initiated by a terminal device, and create a user session matching the session request; a virtual audio device creation unit, used to create a virtual audio device required by the target application in the user session according to the audio transmission relationship of one or more target applications in the user session, and establish a connection relationship between the virtual audio device and the target application; an application running unit, used to run the target application so that the virtual audio device connected to the target application transmits audio stream data to the target application and / or captures the audio stream data output by the target application.

[0008] On the other hand, the present disclosure further provides a cloud server, which includes a memory and a processor, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the above-mentioned audio stream data processing method is implemented.

[0009] On the other hand, the present disclosure further provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the above-mentioned audio stream data processing method is implemented.

[0010] According to the technical solution provided by one or more embodiments of the present disclosure, after receiving a session request from a user's terminal device, a cloud server can create a corresponding user session. In the user session, one or more target applications required by the user can be run. In order to enable the target application to flexibly process audio stream data, before running the target application, a virtual audio device required by the target application can be created for the audio transmission relationship of the target application. The virtual audio device can capture the audio stream data output by the target application, and can also transmit the audio stream data sent by the terminal device to the target application for processing.

[0011] After creating a virtual audio device, a connection relationship between the virtual audio device and the target application can be established to obtain a transmission topology of audio stream data. Subsequently, after running the target application, the virtual audio device can transmit audio stream data to the target application and / or capture audio stream data output by the target application, thereby realizing a flexible audio stream data processing process.

[0012] It can be seen that the technical solution provided by one or more embodiments of the present disclosure can flexibly implement the audio stream processing process through a virtual audio device in a user session according to the actual audio transmission relationship of the target application, thereby meeting the user's needs for complex application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The features and advantages of the present disclosure will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present disclosure in any way. In the accompanying drawings:

[0014] Figure 1 A schematic diagram of the process of real-time cloud rendering in an application scenario is shown;

[0015] Figure 2 A schematic diagram showing the steps of an audio stream data processing method in one embodiment of the present disclosure is shown;

[0016] Figure 3 An operation schematic diagram of a cloud rendering platform in one embodiment of the present disclosure is shown;

[0017] Figure 4 A schematic diagram of transmission of audio stream data in one embodiment of the present disclosure is shown;

[0018] Figure 5 A schematic diagram of processing within a user session in one embodiment of the present disclosure is shown;

[0019] Figure 6 A schematic diagram of processing within a user session in another embodiment of the present disclosure is shown;

[0020] Figure 7 A schematic diagram of processing within a user session in another embodiment of the present disclosure is shown;

[0021] Figure 8 A schematic diagram of processing within a user session in another embodiment of the present disclosure is shown;

[0022] Fig. 9 A schematic diagram of functional modules of an audio stream data processing device in one embodiment of the present disclosure is shown;

[0023] Fig.10 A schematic diagram of the structure of a cloud server in one embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.

[0025] The audio stream data processing method provided in one embodiment of the present disclosure can be applied to a cloud server. The cloud server can be a server with a high-performance graphics processing device. The cloud server can receive a session request from a user's terminal device and run a target application required by the user in response to the session request. After rendering and processing the audio and video data in the target application, the cloud server can feed back the audio and video data to the user's terminal device.

[0026] In actual applications, a cloud server can be an independent server or a server cluster consisting of multiple servers. Generally speaking, a cloud rendering platform is usually used to provide cloud rendering services to user terminal devices. The cloud server can be deployed in the cloud rendering platform, and multiple virtual machines can be run in the cloud server to isolate different user data and provide users with stable services.

[0027] See also Figure 2 , an audio stream data processing method provided by an embodiment of the present disclosure may include the following steps.

[0028] S1: receiving a session request initiated by a terminal device, and creating a user session matching the session request.

[0029] In this embodiment, the terminal device used by the user may be an electronic device such as a smart phone, a personal computer, or virtual reality glasses. In the user's terminal device, an application (application, App) that supports cloud rendering technology may be installed. When the data in the application needs to be processed by the cloud rendering platform, the terminal device may initiate a session request to the cloud rendering platform. Depending on the application scenario, the session request may be initiated using protocols such as HTTP and HTTPS, and of course, may also be initiated using protocols such as TCP or UDP. Generally speaking, in the session request initiated by the terminal device, parameters such as a user identifier, an application identifier, and preference data may be carried. Among them, the user identifier may be used to uniquely indicate the user's identity, and the user identifier may be, for example, a user account pre-registered by the user in the application. The application identifier may uniquely indicate the application identity, and the application identifier may be, for example, an application name, or a background code corresponding to the application name. The preference data may be parameters such as resolution, frame rate, and bit rate configured when the application is running.

[0030] See also Figure 3In one embodiment, after receiving a session request initiated by a terminal device, the cloud rendering platform can select a suitable virtual machine in the cloud rendering platform according to the virtual resources required for the session request through the platform session scheduling service. Specifically, the platform session scheduling service can identify the virtual resources required for the session request by parsing the parameters carried in the session request. The virtual resources can be, for example, CPU resources, memory resources, storage resources, network resources, etc. obtained through virtualization technology. After identifying the virtual resources required for the session request, the platform session scheduling service can filter out a target virtual machine with the virtual resources from a large number of virtual machines, and schedule the session request to the target virtual machine. In the target virtual machine, a user session management service can be run. After the session request is scheduled to the target virtual machine, a user session matching the session request can be created in the target virtual machine through the user session management service. Among them, the matching degree between the user session and the session request can be reflected in: the application running environment set in the user session is set by preference data in the session request, etc.

[0031] S3: According to the audio transmission relationship of one or more target applications in the user session, a virtual audio device required by the target application is created in the user session, and a connection relationship between the virtual audio device and the target application is established.

[0032] In this embodiment, by parsing the application execution parameters in the user session, the audio transmission relationship of one or more target applications to be run can be obtained. The audio transmission relationship can indicate the transmission process of audio stream data between target applications. For example, Figure 4 The target application to be run in the user session is application 1. After the user inputs the audio stream data through the audio input device of the terminal device, the audio stream data can be transmitted to the user session of the cloud server through the real-time communication module (RTC module). The audio stream data can be processed by application 1 and then returned to the user's terminal device through the RTC module. In this case, the audio transmission relationship parsed in the user session can represent the transmission process of the audio stream data from the RTC module to application 1 and then from application 1 to the RTC module.

[0033] In this embodiment, after clarifying the audio transmission relationship of one or more target applications in the user session, a virtual audio device required by the target application can be created in the user session, and a connection relationship between the virtual audio device and the target application can be established. Wherein, the virtual audio device can be implemented based on a virtual audio driver. Depending on the operating system, the development method of the virtual audio driver will also be different. For example, under the windows operating system, the virtual audio driver can be developed using the DevCon tool in the Windows Driver Kit (WDK). Specifically, when developing a virtual audio driver, a kernel audio stream (Kernel Streaming, KS) interface can be created based on a preset driver. The preset driver can be, for example, a WDM (Windows Driver Model, Windows Driver Model) development specification in the Windows platform. The created KS interface has one or more resource endpoints (EndPoints), which can connect to the created virtual audio device.

[0034] In actual applications, the virtual audio driver can be built into the virtual machine image together with the operating system. In this way, when the virtual machine is running, the virtual audio driver will be available in the operating system, which facilitates the subsequent creation of virtual audio devices. In addition, the virtual audio driver can be provided to the upper-level audio interface (such as WASAPI, MME, DirectSound, etc.) through the system audio engine (Audio Engine) for access. The endpoint information of each resource endpoint of the virtual audio driver can be obtained by the system audio service (Audio Service), so that the connected virtual audio device can be called.

[0035] In one embodiment, the virtual audio device may include a virtual audio output device and a virtual audio pipe device. Among them, the virtual audio output device can be used to capture the audio stream data output by the target application, and feed the captured audio stream data back to the terminal device through the RTC module. The virtual audio pipe device can be used to redirect the audio stream data of the terminal device or the audio stream data of the first target application to the second target application. Depending on the type of virtual audio device, the creation process will also be different.

[0036] Specifically, taking the Windows platform as an example, for the Windows WDM / KS driver, the virtual audio output device can be implemented as a virtual device of the KSNODETYPE_SPEAKER type, which can transfer audio stream data to the output endpoint and allow audio data capture through the audio interface provided by the system (such as WASAPI).

[0037] The function of a virtual audio pipe device is equivalent to an audio cable. Taking the Windows platform as an example, for Windows WDM / KS drivers, a virtual audio pipe device can be implemented as two virtual devices of the KSNODETYPE_LINE_CONNECTOR type, where one endpoint of the virtual audio pipe device serves as the pipe inlet and the other endpoint serves as the pipe outlet, thereby transmitting audio stream data from one endpoint to another.

[0038] It should be noted that, although the above embodiments of the present disclosure are all described using the Windows platform as an example to illustrate the principle of the solution, those skilled in the art should be aware that the technical solution of the present disclosure is not limited to the Windows platform. Those skilled in the art, on the premise of understanding the essence of the technical solution of the present disclosure, can implement virtual audio devices with the same or similar functions on other operating system platforms based on the same or similar concepts, and these technical solutions should also be included in the protection scope of the present disclosure.

[0039] In this embodiment, based on the audio transmission relationship, a virtual audio device required by the target application can be created at the corresponding position of the user session, and a connection relationship is established between the virtual audio device and the target application according to the transmission process of the audio stream data. Figure 4 In the example, the audio transmission relationship represents the transmission process of audio stream data from the RTC module to application 1 and then from application 1 to the RTC module. Then, a virtual audio pipeline device can be created between the RTC module and application 1, and a virtual audio output device can be created between application 1 and the RTC module, thereby obtaining Figure 5 The audio stream data transmission diagram shown in the figure. Figure 5 In the figure, the RTC module on the left can receive the audio stream data uploaded by the terminal device, which can be captured by the virtual audio pipeline device A and transmitted to the application 1. After the application 1 processes the audio stream data, the output audio stream device can be captured by the virtual audio output device B and output by the virtual audio output device B to the RTC module on the right. The RTC module on the right can feed back the received audio stream data to the terminal device.

[0040] It should be noted that although Figure 5 Two RTC modules are shown, but this is just an exemplary expression used for the convenience of explaining the technical solution of the present disclosure. In practical applications, the same RTC module can be used to interact with the terminal device for audio stream data. Of course, different RTC modules can also be used to send and receive audio stream data, and the present disclosure does not limit this.

[0041] S5: Run the target application so that the virtual audio device connected to the target application transmits audio stream data to the target application and / or captures audio stream data output by the target application.

[0042] In this embodiment, after a virtual audio device is created in a user session and a connection relationship is established between the virtual audio device and the target application, the target application can be run in the user session, thereby providing cloud rendering services for the terminal device. In the user session, the running target applications can be managed by the application execution module, thereby controlling the running and closing of each target application. In addition, in the user session, each virtual audio device can also be managed by the device management module, thereby controlling the creation and deregistration of each virtual audio device.

[0043] In this embodiment, after the target application in the user session is running, the audio stream data can be processed in sequence according to the audio stream transmission process represented by the audio transmission relationship, so as to realize the interaction of audio stream data between the terminal device and the cloud rendering platform. Specifically, the virtual audio device can transmit audio stream data to the target application, and can also capture the audio stream data output by the target application. For example, Figure 5 In the example, virtual audio pipeline device A can capture the audio stream data output by the left RTC module and transmit the audio stream data to application 1. Virtual audio output device B can capture the audio stream data output by application 1 and output the audio stream data to the right RTC module.

[0044] According to the technical solution provided by one or more embodiments of the present disclosure, after receiving a session request from a user's terminal device, a cloud server can create a corresponding user session. In the user session, one or more target applications required by the user can be run. In order to enable the target application to flexibly process audio stream data, before running the target application, a virtual audio device required by the target application can be created for the audio transmission relationship of the target application. The virtual audio device can capture the audio stream data output by the target application, and can also transmit the audio stream data sent by the terminal device to the target application for processing.

[0045] After creating a virtual audio device, a connection relationship between the virtual audio device and the target application can be established to obtain a transmission topology of audio stream data. Subsequently, after running the target application, the virtual audio device can transmit audio stream data to the target application and / or capture audio stream data output by the target application, thereby realizing a flexible audio stream data processing process.

[0046] It can be seen that the technical solution provided by one or more embodiments of the present disclosure can flexibly implement the audio stream processing process through a virtual audio device in a user session according to the actual audio transmission relationship of the target application, thereby meeting the user's needs for complex application scenarios.

[0047] In one embodiment, after establishing the connection relationship between the virtual audio device and the target application, the virtual audio device and the target application with the connection relationship can be bound. Specifically, in a user session, the binding relationship between the virtual audio device and the target application can be maintained, and the binding relationship can be a one-to-many relationship. For example, the same virtual audio device can be bound to one or more target applications, indicating that the one or more target applications all have a connection relationship with the virtual audio device. Similarly, the same target application can also be bound to one or more virtual audio devices, indicating that the target application all has a connection relationship with the one or more virtual audio devices.

[0048] Of course, in some application scenarios, in order to ensure that the audio stream data between different target applications are isolated from each other, different virtual audio output devices can be created for different target applications. In this way, when feeding back audio stream data to the terminal device, the interference of other target applications can be avoided. Specifically, in the prior art, usually the audio output device is captured by the Loopback mode to the audio stream data, thereby easily obtaining the audio output stream. If however a single audio output device is shared by a plurality of applications, the audio output streams of different applications will be mixed together. For this problem, the technical scheme disclosed in the present invention can bind different virtual audio devices for different target applications, thereby realizing the mutual isolation between the audio data streams. Subsequently, the RTC module can adopt the Loopback mode to capture the audio stream data for specific virtual audio output device, thereby avoiding the interference of other audio stream data.

[0049] In one implementation, when binding the virtual audio device to the target application, taking Windows as an example, the application execution module may bind the target application to the virtual audio device through the SetPersistedDefaultAudioEndpoint method of the IAudioPolicyConfigFactory interface.

[0050] In one embodiment, the virtual audio device in the user session can also be managed dynamically. Specifically, when the target application exits operation, the device management module can determine whether the target virtual audio device bound to the target application is connected to other target applications. If the target virtual audio device is not connected to other target applications, the target virtual audio device can be deregistered to save resources of the virtual machine.

[0051] In one embodiment, by creating a virtual audio device, scenarios of complex user sessions can be supported. Specifically, if the audio transmission relationship represents the audio transmission from the terminal device to the target application, a virtual audio pipe device can be created between the RTC module and the target application. If the audio transmission relationship represents the audio transmission from the target application to the terminal device, a virtual audio output device can be created between the target application and the RTC module. If the audio transmission relationship represents the audio transmission from the first target application to the second target application, a virtual audio pipe device can be created between the first target application and the second target application. Of course, the above-mentioned cases of creating virtual audio devices may only exist in one or more cases in the same user session.

[0052] See also Figure 6 In this application scenario example, there are application 1 and application 6, where application 1 can be a cloud game application and application 6 can be a voice special effects processing application. Among them, the virtual audio pipeline device A can capture the user voice of the current user uploaded by the terminal device from the RTC module, and the voice is transmitted to application 6. After voice special effects processing, the obtained special effects voice can be captured by the virtual audio pipeline device F and transmitted to application 1. In this way, the original game scene sound and the received special effects voice in application 1 can be captured by the virtual audio output device B and pushed to the terminal devices of other game users through the RTC module. In this way, other users can hear the special effects voice and game scene sound of the current user.

[0053] In one embodiment, if the audio transmission relationship represents that the audio stream data output by multiple first target applications is synthesized and then input into a second target application, a virtual audio pipe device can be created between the multiple first target applications and the second target application, wherein one end of the virtual audio pipe device is connected to the multiple first target applications, and the other end of the virtual audio pipe device is connected to the second target application.

[0054] See also Figure 7 In this application scenario example, there are application 2, application 3 and application 4, where application 2 and application 3 can be the multiple first target applications mentioned above, and application 4 can be the second target application mentioned above. The audio stream data output by application 2 and application 3 can be captured and synthesized by the virtual audio pipeline device D, and then transmitted to application 4 for processing. The processed audio stream data can be captured by the virtual audio output device E and fed back to the terminal device through RTC.

[0055] In one embodiment, the dynamic loading process of the application may also be supported. Specifically, if a new application is run in a user session, the audio transmission relationship between the new application and the target application already running in the user session may be determined, and based on the determined audio transmission relationship, a new virtual audio device may be created for the new application or the new application may be connected to the created virtual audio device.

[0056] See also Figure 8 , when the terminal device is running the cloud game application (application 1), it can also run the cloud music player application (application 5), and the cloud music player application can be a new application running in the user session. According to the audio transmission relationship of the cloud music player application, it can be determined that the audio stream data output by the cloud music player application should be transmitted to the cloud game application. In view of this, the cloud music player application can be connected to the created virtual audio pipe device A. In this way, the audio stream data output by the cloud music player application can be transmitted to the cloud game application through the virtual audio pipe device A. The audio stream data output by the cloud game application includes not only the game scene sound, but also the music that the user wants to listen to. The audio stream data composed of the game scene sound and music will be captured by the virtual audio output device B, and finally fed back to the user's terminal device. In this way, the user can also hear music while playing the game. Figure 6 It can be used as a complex application scenario that combines the two user needs of voice special effects processing and cloud music playback. Figure 6 In the figure, application 1 is a cloud game application, application 5 is a cloud music player application, and application 6 is a voice special effect processing application.

[0057] For different application scenarios, when creating a new virtual audio device for a new application or connecting a new application to an already created virtual audio device, at least one of the following situations may be included:

[0058] (a) If the determined audio transmission relationship indicates that the new application processes the audio stream data input to the already running target application, a virtual audio pipe device is created between the new application and the already running target application.

[0059] (b) If the determined audio transmission relationship indicates that the new application processes the audio stream data output by the running target application, a virtual audio pipe device is created between the running target application and the new application.

[0060] (c) If the determined audio transmission relationship represents that the audio stream data output by the new application is synthesized with other audio stream data, the audio stream data is input into the running target application, a virtual audio pipe device for inputting the other audio stream data into the running target application is identified, and the new application is connected to the identified virtual audio pipe device.

[0061] (d) If the determined audio transmission relationship represents synthesis of first audio stream data output by the new application and second audio stream data output by the running target application, identify a virtual audio device for capturing the second audio stream data, and connect the new application to the identified virtual audio device.

[0062] Of course, as more application scenarios are developed, more situations can be included, which are not listed here one by one.

[0063] As can be seen from the above, the technical solution provided by the present disclosure can provide flexible audio stream data processing capabilities and provide users with rich audio processing scenarios, such as cloud music, cloud audio effects and other capability expansions. The isolation of audio stream data between multiple applications is achieved, which facilitates the cloud rendering platform to process the input and output of audio stream data, thereby improving the comprehensive processing capabilities of the cloud rendering platform and meeting the needs of users for complex application scenarios. In addition, the technical solution provided by the present disclosure also allows users to dynamically create or cancel virtual audio devices and flexibly adjust the connection relationship between applications and virtual audio devices according to changes in the user's application call combination during the user session. During the execution of each application, there is no perception of upstream and downstream applications, and only the focus needs to be on the interaction between itself and the virtual audio device.

[0064] See also Fig. 9 The present disclosure also provides an audio stream data processing device, the device comprising:

[0065] The session creation unit 100 is used to receive a session request initiated by a terminal device and create a user session matching the session request;

[0066] A virtual audio device creation unit 200 is used to create a virtual audio device required by the target application in the user session according to the audio transmission relationship of one or more target applications in the user session, and establish a connection relationship between the virtual audio device and the target application;

[0067] The application running unit 300 is used to run the target application so that the virtual audio device connected to the target application transmits audio stream data to the target application and / or captures audio stream data output by the target application.

[0068] In one embodiment, the session creation unit 100 is further configured to identify virtual resources required by the session request and select a target virtual machine having the virtual resources to create a user session matching the session request in the target virtual machine.

[0069] In one embodiment, the device further comprises:

[0070] The interface creation unit is used to create a kernel audio stream interface based on a preset driver program, wherein the kernel audio stream interface has one or more resource endpoints, and the resource endpoints are used to connect to the created virtual audio device.

[0071] In one embodiment, the virtual audio device includes a virtual audio output device and a virtual audio pipe device; wherein the virtual audio output device is used to capture the audio stream data output by the target application and feed the captured audio stream data back to the terminal device; the virtual audio pipe device is used to redirect the audio stream data of the terminal device or the audio stream data of the first target application to the second target application.

[0072] In one embodiment, the device further comprises:

[0073] A binding unit, used to bind the virtual audio device and the target application that have a connection relationship;

[0074] The deregistration unit is used to determine whether the target virtual audio device bound to the target application is connected to other target applications when the target application exits operation, and to deregister the target virtual audio device if the target virtual audio device is not connected to other target applications.

[0075] In one embodiment, the virtual audio device creation unit 200 is also used to create a virtual audio pipe device between the instant communication module and the target application if the audio transmission relationship represents audio transmission from the terminal device to the target application; create a virtual audio output device between the target application and the instant communication module if the audio transmission relationship represents audio transmission from the target application to the terminal device; and create a virtual audio pipe device between the first target application and the second target application if the audio transmission relationship represents audio transmission from the first target application to the second target application.

[0076] In one embodiment, the virtual audio device creation unit 200 is also used to create a virtual audio pipe device between the multiple first target applications and the second target application if the audio transmission relationship represents that the audio stream data output by multiple first target applications is synthesized and input into a second target application; wherein one end of the virtual audio pipe device is connected to the multiple first target applications, and the other end of the virtual audio pipe device is connected to the second target application.

[0077] In one embodiment, the device further comprises:

[0078] A dynamic processing unit is used to determine, if a new application is run in the user session, an audio transmission relationship between the new application and a target application already running in the user session, and based on the determined audio transmission relationship, create a new virtual audio device for the new application or connect the new application to an already created virtual audio device.

[0079] In one embodiment, the dynamic processing unit is also used to, if the determined audio transmission relationship represents that the new application processes the audio stream data input to the running target application, create a virtual audio pipe device between the new application and the running target application; if the determined audio transmission relationship represents that the new application processes the audio stream data output by the running target application, create a virtual audio pipe device between the running target application and the new application; if the determined audio transmission relationship represents that the audio stream data output by the new application is synthesized with other audio stream data and then input into the running target application, identify the virtual audio pipe device used to input the other audio stream data into the running target application, and connect the new application to the identified virtual audio pipe device; if the determined audio transmission relationship represents that the first audio stream data output by the new application is synthesized with the second audio stream data output by the running target application, identify the virtual audio device used to capture the second audio stream data, and connect the new application to the identified virtual audio device.

[0080] Each unit described in the above embodiments may be implemented by a computer chip or a product having a certain function. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0081] For the convenience of description, the above devices are described in terms of functions and are described separately in various units. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0082] See also Fig.10 The present disclosure also provides a cloud server, which includes a memory and a processor, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the above-mentioned audio stream data processing method is implemented.

[0083] The present disclosure also provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the above-mentioned audio stream data processing method is implemented.

[0084] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.

[0085] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory, that is, implementing the method in the above method embodiment.

[0086] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0087] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memory.

[0088] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, cloud server, and storage medium, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0089] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

[0090] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for processing audio stream data, characterized in that: The method comprises: Receiving a session request initiated by a terminal device, and creating a user session matching the session request; According to the audio transmission relationship of one or more target applications in the user session, create a virtual audio device required by the target application in the user session, and establish a connection relationship between the virtual audio device and the target application; Running the target application so that the virtual audio device connected to the target application transmits audio stream data to the target application and / or captures audio stream data output by the target application; Wherein, creating a user session matching the session request includes: Identifying virtual resources required for the session request, and selecting a target virtual machine having the virtual resources, so as to create a user session matching the session request in the target virtual machine; Creating a virtual audio device required by the target application in the user session includes at least one of the following: If the audio transmission relationship represents audio transmission from the terminal device to the target application, a virtual audio pipeline device is created between the instant messaging module and the target application; If the audio transmission relationship represents audio transmission from the target application to the terminal device, creating a virtual audio output device between the target application and the instant messaging module; If the audio transmission relationship represents audio transmission from a first target application to a second target application, a virtual audio pipe device is created between the first target application and the second target application.

2. The method according to claim 1, characterized in that Before creating the virtual audio device required by the target application in the user session, the method further includes: A kernel audio stream interface is created based on a preset driver, wherein the kernel audio stream interface has one or more resource endpoints, and the resource endpoints are used to connect to the created virtual audio device.

3. The method according to claim 1 or 2, characterized in that: The virtual audio device includes a virtual audio output device and a virtual audio pipe device; wherein the virtual audio output device is used to capture the audio stream data output by the target application and feed the captured audio stream data back to the terminal device; the virtual audio pipe device is used to redirect the audio stream data of the terminal device or the audio stream data of the first target application to the second target application.

4. The method according to claim 1, characterized in that: After establishing a connection relationship between the virtual audio device and the target application, the method further includes: Bind a virtual audio device and a target application that are in a connection relationship; wherein, when the target application exits, determine whether the target virtual audio device bound to the target application is connected to other target applications, and if the target virtual audio device is not connected to other target applications, unregister the target virtual audio device.

5. The method according to claim 1, characterized in that Creating a virtual audio device required by the target application in the user session and establishing a connection relationship between the virtual audio device and the target application includes: If the audio transmission relationship represents that the audio stream data output by multiple first target applications are synthesized and then input into a second target application, a virtual audio pipe device is created between the multiple first target applications and the second target application; wherein one end of the virtual audio pipe device is connected to the multiple first target applications, and the other end of the virtual audio pipe device is connected to the second target application.

6. The method according to claim 1, characterized in that After running the target application, the method further includes: If a new application is run in the user session, an audio transmission relationship between the new application and a target application already running in the user session is determined, and based on the determined audio transmission relationship, a new virtual audio device is created for the new application or the new application is connected to an already created virtual audio device.

7. The method according to claim 6, characterized in that Creating a new virtual audio device for the new application or connecting the new application to the created virtual audio device includes at least one of the following: If the determined audio transmission relationship indicates that the new application processes the audio stream data input to the already running target application, creating a virtual audio pipe device between the new application and the already running target application; If the determined audio transmission relationship indicates that the new application processes the audio stream data output by the running target application, creating a virtual audio pipe device between the running target application and the new application; If the determined audio transmission relationship indicates that the audio stream data output by the new application is synthesized with other audio stream data, the audio stream data is input into the running target application, a virtual audio pipe device for inputting the other audio stream data into the running target application is identified, and the new application is connected to the identified virtual audio pipe device; If the determined audio transmission relationship represents the synthesis of the first audio stream data output by the new application and the second audio stream data output by the running target application, identify the virtual audio device for capturing the second audio stream data, and connect the new application to the identified virtual audio device.

8. An audio stream data processing device, characterized in that: The device comprises: A session creation unit, configured to receive a session request initiated by a terminal device and create a user session matching the session request; a virtual audio device creation unit, configured to create a virtual audio device required by the target application in the user session according to the audio transmission relationship of one or more target applications in the user session, and establish a connection relationship between the virtual audio device and the target application; An application running unit, configured to run the target application so that the virtual audio device connected to the target application transmits audio stream data to the target application and / or captures audio stream data output by the target application; The session creation unit is further configured to identify virtual resources required by the session request and select a target virtual machine having the virtual resources to create a user session matching the session request in the target virtual machine; The virtual audio device creation unit is also used to create a virtual audio pipe device between the instant communication module and the target application if the audio transmission relationship represents audio transmission from the terminal device to the target application; create a virtual audio output device between the target application and the instant communication module if the audio transmission relationship represents audio transmission from the target application to the terminal device; and create a virtual audio pipe device between the first target application and the second target application if the audio transmission relationship represents audio transmission from the first target application to the second target application.

9. A cloud server, characterized in that: The cloud server includes a memory and a processor, the memory is used to store a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Audio data streaming method and device of cloud host, equipment and storage medium

    CN112631736A