Media data stream processing method and device, cluster, medium and program product

By adding traceless logos and performing traceless processing in audio and video conferencing, the problem of uncontrollable user media content recording permissions is solved, ensuring that sensitive content is not recorded, and the privacy protection of audio and video conferencing is achieved.

CN120302087APending Publication Date: 2025-07-11HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410342096.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-09
Filing Date
2024-03-22
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In audio and video conferencing, the user's media content recording permissions are not controllable, resulting in the possibility of privacy or sensitive content being leaked and unable to meet the user's privacy protection needs.

Method used

By adding traceless identifiers to the media data stream, identifying and processing them with computing devices, blocking sensitive content, ensuring that the recording server only records media content without traceless identifiers.

Benefits of technology

It realizes controllability of media data stream recording permissions in audio and video conferencing, avoids the leakage of private content, and expands the privacy protection function of audio and video conferencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302087A_ABST
    Figure CN120302087A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a media data stream processing method and device, a cluster, a medium and a program product, relates to the technical field of data processing, avoids leakage of media contents belonging to privacy contents in an audio and video conference, and further expands the privacy protection function of the audio and video conference. The method comprises the following steps: acquiring a media data stream, the media data stream being data transmitted by a first terminal to other participating terminals in an audio and video conference, and the other participating terminals being used for displaying media content corresponding to the media data stream; if the media data stream comprises the traceless identifier, traceless processing is carried out on the media data stream according to the traceless identifier to obtain a processed media data stream, and the traceless processing is used for shielding media content corresponding to part or all of media data in the media data stream; and sending the processed media data stream to a recording server, so that the recording server records the media content indicated by the processed media data stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a method, apparatus, cluster, medium, and program product for processing media data streams. Background Art

[0002] With the continuous development of Internet technology, users in different locations can access audio and video conferences through the network via online audio and video conferences, achieving the purpose of holding meetings online without being restricted by the location of the users. Online audio and video conferences can also provide various post-meeting review functions. For example, functions such as recording audio and video during the meeting, generating intelligent meeting minutes for the meeting content, and recording chat records during the meeting. However, there may be some sensitive content in the media content of each participating user in the audio and video conference, and the participating users do not want the sensitive part of the content to be recorded or saved after the meeting and then spread after the meeting.

[0003] Currently, the meeting host in an audio and video conference has the permission to pause recording or edit and delete the corresponding media content after the meeting. That is to say, when there is some sensitive content in the media content of a participating user who is not the host in the audio and video conference, the participating user needs to notify the host of the media content that needs to be stopped recording or deleted after the meeting.

[0004] In the related art, by notifying the host to stop recording at a certain node or notifying the host to delete the corresponding media content after the meeting, it cannot be guaranteed that the host fully understands the processing requirements of the participating users for the media content, and at the same time, it cannot be guaranteed that the host can correctly complete the corresponding operations. Since the participating users in the audio and video conference do not have controllability over the recording permission of the media data stream sent through the terminal, there is a risk of leakage of media content that belongs to privacy or sensitive content in the audio and video conference, and thus the post-meeting review function of the audio and video conference cannot meet the needs of users. Summary of the Invention

[0005] Embodiments of this application provide a method, apparatus, cluster, medium, and program product for processing media data streams, which ensure the controllability of the recording permission of the media data stream sent by the terminal of the participating user in the audio and video conference, avoid the leakage of media content that belongs to private content in the audio and video conference, and thus expand the privacy protection function of the audio and video conference.

[0006] In a first aspect, the present application provides a method for processing media data streams, which is applied to a computing device. The method includes: obtaining a media data stream, where the media data stream is data transmitted by a first terminal to other participating terminals during an audio-video conference, and the other participating terminals are used to display the media content corresponding to the media data stream; if the media data stream includes a markless identifier, performing markless processing on the media data stream according to the markless identifier to obtain a processed media data stream, and the markless processing is used to mask part or all of the media content corresponding to the media data in the media data stream; sending the processed media data stream to a recording server so that the recording server records the media content indicated by the processed media data stream.

[0007] It can be understood that a user terminal that needs to enable the markless function adds a markless identifier to the media data stream, and sends the media data stream with the markless identifier added to a computing device in an audio-video conference system. The computing device determines the media data stream in the audio-video conference that needs markless processing by identifying the markless identifier, performs markless processing on the media data stream, and then sends it to the recording server, so that the recording server can record the media content indicated by the media data stream with the markless identifier masked. Thus, it ensures the controllability of the recording permission of the media data stream sent by the participating user terminals in the audio-video conference, avoids the leakage of media content that belongs to private content in the audio-video conference, and further expands the privacy protection function of the audio-video conference.

[0008] In a possible implementation manner, the media data stream includes multiple data packets. Performing markless processing on the media data stream according to the markless identifier to obtain a processed media data stream includes: determining a target data packet from the multiple data packets, where the target data packet is a data packet that includes the markless identifier in the packet header field; performing markless processing on the target data packet to obtain a processed media data stream, and the processed media data stream includes the target data packet that has undergone markless processing and the other data packets in the multiple data packets except the target data packet.

[0009] It can be understood that by determining whether the packet header field of the data packet includes a markless identifier, the target data packet that needs markless processing in the media data stream can be determined, so as to perform markless processing on the media content corresponding to the target data packet, and the media content that needs markless processing and the media content that does not need markless processing in the media data stream can be finely divided, improving the processing effect of the markless processing.

[0010] In a possible implementation manner, the method further includes: sending the processed media data stream to an artificial intelligence (AI) server, where the AI server is used to perform AI intelligent analysis on the media content indicated by the processed media data stream; receiving the result of the AI intelligent analysis on the processed media data stream returned by the AI server, and the result is used to be displayed in the audio-video conference.

[0011] It is understandable that by sending the processed media data stream to the AI server, it can be ensured that the media content that needs to be processed without a trace will not be subjected to AI intelligent analysis by the AI server, thus ensuring the privacy of the media content.

[0012] In a possible implementation, the media data stream includes one or more of the following: audio data stream, video data stream, shared media stream, message stream.

[0013] In a possible implementation, if the media data stream is a media stream based on the scalable video coding (SVC) mode, the computing device includes a media stream router (SFU); or, if the media data stream is a media stream based on the advanced video coding (AVC) mode, the computing device includes a media server.

[0014] In a possible implementation, the method further includes: if it is determined that the media data stream includes a trace-free identifier, generating status information, where the status information is used to indicate that the media data stream sent by the first terminal needs to be processed without a trace; sending the status information to other terminals, so that the other terminals display specified information to the user, and the specified information is used to prompt the user that the media data stream sent by the first terminal needs to be processed without a trace.

[0015] It is understandable that by sending the status information to other terminals, it can be ensured that the users on the other terminal sides can obtain the user terminals that have enabled trace-free in the audio and video conference, thus improving the user experience in the audio and video conference.

[0016] In a possible implementation, if the status information is further used to indicate the types of media data streams in the media data stream sent by the first terminal that need to be processed without a trace, the specified information is further used to prompt the user of the types of media data streams in the media data stream sent by the first terminal that need to be processed without a trace.

[0017] It is understandable that the status information received by other terminals may also include the types of media data streams that are processed without a trace, which can further prompt the user and further improve the user experience in the audio and video conference.

[0018] Second aspect, the present application provides a method for processing media data stream, which is applied to a first client. The first client is a client of an audio-video conference running on a first terminal. The method includes: displaying a markless application control in the conference interface of the audio-video conference. The markless application control is used to provide a user with an application to perform markless processing on the media data stream. The media data stream is data transmitted by the first terminal to other participating terminals in the audio-video conference. The other participating terminals are used to display the media content corresponding to the media data stream. Markless processing is used to mask the media content corresponding to some or all of the media data in the media data stream; in response to receiving a trigger operation on the markless application control, adding a markless identifier to the media data stream; sending the media data stream with the markless identifier added to a computing device. The computing device is used to perform markless processing on the media data stream according to the markless identifier to obtain a processed media data stream, and sending the processed media data stream to a recording server, so that the recording server records the media content indicated by the processed media data stream.

[0019] It can be understood that a user terminal that needs to enable the markless function adds a markless identifier to the media data stream, and sends the media data stream with the markless identifier added to a computing device in the audio-video conference system. The computing device determines the media data stream in the audio-video conference that needs markless processing by identifying the markless identifier, performs markless processing on the media data stream, and then sends it to the recording server, so that the recording server can record the media content that masks the media content indicated by the media data stream with the markless identifier added. Thus, the controllability of the recording permission of the media data stream sent by the participating user terminals in the audio-video conference is ensured, the leakage of media content that belongs to private content in the audio-video conference is avoided, and the privacy protection function of the audio-video conference is further extended.

[0020] In a possible implementation manner, selection controls corresponding to different types of media content are further displayed in the conference interface of the audio-video conference; in response to receiving a trigger operation on the markless application control, adding a markless identifier to the media data stream includes: in response to receiving a trigger operation on the markless application control and receiving a selection operation on the selection control corresponding to the first type of media content, adding a markless identifier to the header fields of each data packet of the media data stream corresponding to the first type of media content in the media data stream.

[0021] It can be understood that by displaying selection controls corresponding to different types of media content, the purpose of adding markless identifiers for different types of media content can be achieved, thereby refining the media content for which markless processing is performed and improving the user experience in the audio-video conference.

[0022] Third aspect, an embodiment of the present application provides a media data stream processing device, and the media data stream processing device is used to execute any one of the media data stream processing methods provided in the first aspect above.

[0023] In a possible implementation manner, embodiments of the present application may divide functional modules of the media data stream processing device according to the method provided in the first aspect above. For example, each functional module may be divided corresponding to each function, or two or more functions may be integrated into one processing module. Exemplarily, embodiments of the present application may divide the media data stream processing device into an acquisition module, a processing module, a sending module, etc. according to functions. Descriptions of possible technical solutions and beneficial effects executed by each of the above-divided functional modules may refer to the technical solutions provided in the first aspect above or its corresponding possible implementation manners, and will not be elaborated here.

[0024] In a fourth aspect, embodiments of the present application provide a media data stream processing device, which is used to execute any one of the media data stream processing methods provided in the second aspect above.

[0025] In a possible implementation manner, embodiments of the present application may divide functional modules of the media data stream processing device according to the method provided in the second aspect above. For example, each functional module may be divided corresponding to each function, or two or more functions may be integrated into one processing module. Exemplarily, embodiments of the present application may divide the media data stream processing device into a setting module, an adding module, a sending module, etc. according to functions. Descriptions of possible technical solutions and beneficial effects executed by each of the above-divided functional modules may refer to the technical solutions provided in the second aspect above or its corresponding possible implementation manners, and will not be elaborated here.

[0026] In a fifth aspect, embodiments of the present application provide a computing device, which includes a processor and a memory, and the processor is coupled to the memory; the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor so that the computing device implements the media data stream processing method as described in the above aspects.

[0027] In a sixth aspect, embodiments of the present application provide a computing device cluster, which includes at least one computing device, and each computing device includes: a processor and a memory, and the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the media data stream processing method provided in various optional implementation manners of the first aspect or the second aspect above.

[0028] In a seventh aspect, embodiments of the present application provide a computer-readable storage medium, in which at least one computer program instruction is stored, and the computer program instruction is loaded and executed by a processor to implement the media data stream processing method as described in the above aspects.

[0029] In an eighth aspect, an embodiment of the present application provides a computer program product, which includes computer instructions stored in a computer-readable storage medium. A processor of a computing device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computing device cluster executes the media data stream processing method provided in the various optional implementations of the first aspect or the second aspect above.

[0030] For the specific descriptions of the third aspect to the eighth aspect and their various implementations in the present application, reference may be made to the detailed descriptions in the first aspect and its various implementations or the second aspect and its various implementations; moreover, for the beneficial effects of the third aspect to the eighth aspect and their various implementations, reference may be made to the beneficial effect analysis in the first aspect and its various implementations or the second aspect and its various implementations, which will not be elaborated herein.

[0031] These aspects or other aspects of the present application will be more clearly understood in the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a schematic diagram of an audio and video conferencing system based on the SVC media mode shown according to an exemplary embodiment;

[0033] Figure 2 is a schematic diagram of an audio and video conferencing system based on the AVC media mode shown according to an exemplary embodiment;

[0034] Figure 3 is a schematic diagram of a scenario of media data stream processing shown according to an exemplary embodiment;

[0035] Figure 4 is a schematic diagram of a scenario of media data stream processing shown according to an exemplary embodiment;

[0036] Figure 5 is a schematic flowchart of a media data stream processing method shown according to an exemplary embodiment;

[0037] Figure 6 is Figure 5 a schematic diagram after seamless processing of a video data stream involved in the shown embodiment;

[0038] Figure 7 is Figure 5 a schematic diagram of the interface display of another participating terminal involved in the shown embodiment;

[0039] Figure 8 is a schematic flowchart of a media data stream processing method shown according to an exemplary embodiment;

[0040] Figure 9 is Figure 8 a schematic diagram showing a display of a conference interface involved in the illustrated embodiment;

[0041] Figure 10 a flowchart showing a method for processing media data streams according to an exemplary embodiment;

[0042] Figure 11 a schematic diagram showing the structure of a media data stream processing apparatus according to an exemplary embodiment;

[0043] Figure 12 a schematic diagram showing the structure of a media data stream processing apparatus according to an exemplary embodiment;

[0044] Figure 13 a schematic diagram of a computing device according to an exemplary embodiment;

[0045] Figure 14 a schematic diagram of a computing device cluster according to an exemplary embodiment;

[0046] Figure 15 a schematic diagram of a connection method between computing device clusters according to an exemplary embodiment. Detailed implementation manners

[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.

[0048] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship between associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0049] Moreover, in the description of this application, unless otherwise specified, "a plurality of" means two or more than two. "At least one (item)" or a similar expression thereof refers to any combination of these items, including any combination of a single item or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0050] In addition, to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner for easy understanding.

[0051] First, an exemplary introduction to the application scenarios of the embodiments of the present application is provided.

[0052] With the continuous development of Internet technology, through online audio and video conferencing, users in different locations can access the audio and video conference through the network, achieving the purpose of holding a conference online regardless of the location of the users. Through online audio and video conferencing, various post-meeting review functions can also be provided, such as the recording function of audio and video during the meeting, the intelligent minutes function for the meeting content, the function of recording the chat records during the meeting, etc.

[0053] The current audio and video conference is implemented through an audio and video conference system, which can be an audio and video conference system based on the scalable video coding (SVC) media mode, or an audio and video conference system based on the advanced video coding (AVC) media mode.

[0054] Among them, Figure 1 shows a schematic diagram of an audio and video conference system provided by an embodiment of the present application based on the SVC media mode. As Figure 1 shown, this audio and video conference system includes multiple participating user terminals 10, a media forwarding server 11, a media gateway 12, an intelligent recognition server 13, and a recording server 14.

[0055] Among them, the participating user terminal 10 is the user terminal used for logging in by each user account participating in the audio and video conference. The participating user terminal 10 can be a terminal device such as a computer device or a smart phone, which is not limited here.

[0056] The media forwarding server 11 can be used to receive and forward subscriptions of media data streams, and the media forwarding server 11 may not encode the media data streams. For example, the media forwarding server 11 can be a media stream routing unit (SFU).

[0057] The media gateway 12 can be used to encode and decode media data streams into other formats. For example, a video data stream encoded in H264 can be saved as a picture format and sent to an artificial intelligence (AI) service for image recognition; or, an audio data stream can be saved as a pulse code modulation (PCM) raw audio file and sent to an AI service for speech recognition, etc.

[0058] The intelligent recognition server 13 can be used to receive the data files encoded and transmitted by the media gateway 12, and perform artificial intelligence analysis on the received data files. For example, image recognition or speech recognition, etc.

[0059] The recording server 14 can be used to record the content of an audio-visual conference. The media forwarding server 11 sends the media data stream to the recording server 14, and the recording server 14 can record the media data stream including video data stream, audio data stream, shared media stream or message stream, etc. into video files, audio files, etc.

[0060] Exemplarily, when any participating user terminal 10 in the audio-visual conference system sends a media data stream to the media forwarding server 11, on the one hand, the media forwarding server 11 can forward the media data stream to each of the other participating user terminals 10, so that the other participating user terminals 10 can synchronously display the media content indicated by the media data stream. On the other hand, the media forwarding server 11 can also send the media data stream to the media gateway 12, and the media gateway 12 encodes and decodes the media data stream into an image file or an audio file, etc., and sends the processed image file or audio file to the intelligent recognition server 13, so that the intelligent recognition server 13 can perform image recognition on the image file and speech recognition on the audio file respectively. In addition, the media forwarding server 11 can also send the media data stream to the recording server 14, and the recording server 14 records the corresponding media content according to the media data stream to obtain video files, audio files or shared files, etc.

[0061] In addition, Figure 2 shows a schematic diagram of an audio-visual conference system provided by an embodiment of the present application based on the AVC media mode. As Figure 2As shown, the audio and video conferencing system includes multiple participating user terminals 10, a media server 15, a conference management service 16, an intelligent recognition server 13, and a recording server 14.

[0062] Among them, the participating user terminal 10 is the user terminal used for logging in by each user account participating in the audio and video conference. The participating user terminal 10 can be a terminal device such as a computer device or a smart phone, and there is no limitation here.

[0063] The media server 15 can be used to perform encoding and decoding processing on the received media data stream and convert the media data stream into a format required by other services and terminals for forwarding. For example, the media server 15 can be a multipoint control unit (MCU).

[0064] The conference management service 16 can be used to provide conference reservation and conference control functions for the audio and video conferencing system. It can be provided to the client for access in the form of an application programming interface (API), or a portal form can be provided to provide an access interface, that is, a conference interface is provided to the participating user terminals.

[0065] The intelligent recognition server 13 can be used to receive the data files encoded and transmitted by the media server 15 and perform artificial intelligence analysis on the received data files. For example, image recognition or speech recognition, etc.

[0066] The recording server 14 can be used to record the conference content in the audio and video conference. The media server 15 sends the media data stream to the recording server 14, and the recording server 14 can record the video data stream, audio data stream, shared media stream, or message stream, etc. in the media data stream into video files, audio files, etc.

[0067] Exemplarily, when any participating user terminal 10 in the audio-video conferencing system sends a media data stream to the media server 15, on the one hand, the media server 15 can forward the media data stream to each of the other participating user terminals 10 so that the other participating user terminals 10 can synchronously display the media content indicated by the media data stream. On the other hand, the media server 15 can also perform encoding and decoding processing on the media data stream into an image file, an audio file, etc., and send the obtained image file or audio file after processing to the intelligent recognition server 13, so that the intelligent recognition server 13 can perform image recognition on the image file and speech recognition on the audio file respectively. In addition, the media server 15 can also send the encoded and decoded video files, audio files, etc. to the recording server 14, and the recording server 14 records according to the media content corresponding to the video files, audio files, etc. to obtain videos, audios, or shared content, etc. While any participating user terminal 10 in the audio-video conferencing system sends a media data stream to the media server 15, the participating user terminal 10 can send the media data stream to the conference management service 16, so that the conference management service provides a conference interface to each of the participating user terminals 10.

[0068] Through the above audio-video conferencing system, each participating user terminal can achieve online audio-video synchronous transmission to complete the conference, and can also achieve the function of reviewing after the meeting, reviewing the complete conference video, audio, or shared files and other content. However, since there may be problems with some sensitive content in the media content of each participating user terminal in the audio-video conference, that is, the user does not want the media content sent by the participating user terminal in the conference to be retained, and the user does not want the sensitive part of the content to be recorded or saved after the meeting and then spread after the meeting.

[0069] Currently, the conference host in the audio-video conference has the permission to pause recording or edit and delete the corresponding media content after the meeting. That is to say, in the case where there is some sensitive content in the media content of a participating user who is not the host in the audio-video conference, the participating user needs to notify the host of the media content that needs to be stopped recording or deleted after the meeting. By notifying the host to perform a stop recording operation at a certain node or notifying the host to delete the corresponding media content after the meeting, it cannot be guaranteed that the host fully understands the processing requirements of the participating user for the media content, and at the same time, it cannot be guaranteed that the host can correctly complete the corresponding operation. Due to the uncontrollability of the recording permission of the media data stream sent by the participating user through the terminal in the audio-video conference, there is a risk of leakage of media content that belongs to privacy or sensitive content in the audio-video conference, and as a result, the function of reviewing after the audio-video conference cannot meet the needs of users.

[0070] In view of this, a user terminal that needs to enable the trace-free function adds a trace-free identifier to the media data stream, and sends the media data stream with the trace-free identifier added to a computing device in an audio-video conferencing system. The computing device determines the media data stream in the audio-video conference that needs to be processed in a trace-free manner by identifying the trace-free identifier, and sends the media data stream to a recording server after performing trace-free processing on the media data stream, so that the recording server can block the media content indicated by the media data stream with the trace-free identifier added, that is, the recording server will not record the media content indicated by the media data stream with the trace-free identifier, and the recording server can normally record the media content indicated by the media data stream without the trace-free identifier added. Thus, the controllability of the recording permission of the media data stream sent by the participating user terminals in the audio-video conference is ensured, the leakage of the media content that belongs to the private content in the audio-video conference is avoided, and the privacy protection function of the audio-video conference is further extended.

[0071] Among them, Figure 3 shows a schematic diagram of a scenario for processing media data streams provided by an embodiment of the present application. As Figure 3 shown, for an audio-video conferencing system based on the SVC media mode, in practical applications, a user terminal 10 that starts the trace-free function can send a media data stream with a trace-free identifier added to a media forwarding server 11, and the media forwarding server 11 can forward the media data stream with the trace-free identifier added to other participating user terminals 10. After receiving the media data stream with the trace-free identifier added, other participating user terminals 10 cannot perform local recording, screenshotting, sharing, etc. on the media data stream with the trace-free identifier added to retain the media content indicated by the trace-free identifier. On the other hand, the media forwarding server 11 can also send the media data stream with the trace-free identifier added to a recording server 14. Since the media data stream received by the recording server 14 contains the trace-free identifier, the recording server 14 cannot perform operations such as cloud recording, screenshotting, sharing, etc. on the media content corresponding to the media data stream with the trace-free identifier added to retain the media content indicated by the trace-free identifier.

[0072] Among them, the computing device can be the media forwarding server 11.

[0073] In addition, Figure 4 shows a schematic diagram of a scenario for processing media data streams provided by an embodiment of the present application. As Figure 4As shown in the figure, for an audio and video conferencing system based on the AVC media mode, in practical applications, a user terminal 10 that needs to activate the trace-free function can send a trace-free application to the conference management service 16. After receiving the trace-free application, the conference management service 16 adds a trace-free identifier to the media data stream corresponding to the user terminal that needs to activate the trace-free function. The conference management service 16 sends the media data stream with the trace-free identifier added to the media server 15. The media server 15 can encode and decode the media data stream with the trace-free identifier added to obtain media content that does not contain the media content with the trace-free identifier. If there is media content without the trace-free identifier, the media content will be sent to the recording server 14 and the intelligent recognition server 13. Otherwise, the media content will not be sent to the recording server 14 and the intelligent recognition server 13. In addition, the media server 15 sends the media data stream with the trace-free identifier added to each of the other participating user terminals 10, so that each of the other participating user terminals 10 can normally synchronously display the media content, and ensure that each user terminal 10 cannot perform operations such as local recording, screenshotting, and sharing on the media content indicated by the trace-free identifier to retain the media content.

[0074] Among them, the computing device can be the media server 15.

[0075] It should be noted that the application scenarios and system architectures described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0076] For ease of understanding, the media data stream processing method provided by the present application is introduced exemplarily below in conjunction with the accompanying drawings. This media data stream processing method is applicable to Figure 3 、 Figure 4 the audio and video conferencing system shown in the figure.

[0077] Figure 5 The flowchart of the media data stream processing method provided by an exemplary embodiment of the present application is shown. This media data stream processing method can be executed by a computing device, and the computing device can be Figure 3 the media forwarding server 11 shown in the figure or Figure 4 the media server 15 shown in the figure. This media data stream processing method includes the following steps:

[0078] S101, the computing device obtains the media data stream.

[0079] In an embodiment of the present application, a computing device may receive a media data stream sent by a first terminal. The media data stream may be data transmitted by the first terminal to other participating terminals during an audio-video conference, and the other participating terminals are used to display the media content corresponding to the media data stream.

[0080] Among them, the media data stream may include at least one of an audio data stream, a video data stream, a shared media stream, or a message stream. The shared media stream may be a data stream for sharing screen content or a data stream for sharing files by a user through the first terminal during an audio-video conference. The message stream may be a data stream corresponding to an instant message sent by a user through the first terminal during an audio-video conference.

[0081] In a possible implementation manner, the media content corresponding to the media data stream can be classified into audio content, video content, shared content, or message content according to the type. The media data stream is classified into an audio data stream, a video data stream, a shared media stream, or a message stream according to the type of the indicated media content.

[0082] Exemplarily, a participating user may send, through the first terminal, the video content and audio content of the participating user's speech during an audio-video conference, as well as the data of sharing the screen and files, to a computing device in an audio-video conference system in the form of a data stream. The computing device may be a media forwarding server or a media server.

[0083] S102. If the media data stream includes an incognito flag, the computing device performs incognito processing on the media data stream according to the incognito flag to obtain a processed media data stream.

[0084] In an embodiment of the present application, after receiving the media data stream, the computing device may detect whether each data packet in the data stream includes an incognito flag. If it is determined that the incognito flag is included, the media data stream including the incognito flag may be subjected to incognito processing to obtain a processed media data stream.

[0085] Among them, the incognito flag may be identification information of "incogito=1" added to the header field of the data packet. That is, when the computing device detects that "incogito=1" exists in the header field of the data packet, it is determined that the data packet is added with an incognito flag.

[0086] In a possible implementation manner, the computing device may determine a target data packet from multiple data packets. The target data packet may be a data packet including an incognito flag in the packet header field. Then, the computing device may perform incognito processing on the target data packet to obtain a processed media data stream. The processed media data stream may include the target data packet subjected to incognito processing and other data packets in the multiple data packets except the target data packet.

[0087] Among them, the trace-less processing can be used to mask part or all of the media content corresponding to the media data stream. The masking processing is a processing means that makes users unable to obtain information related to the media content from the processed media content by replacing or covering the media content. Specifically, the masking processing can be to replace the media content with a preset specific content, or to cover the media content with a preset specific content, and the preset specific content can be content unrelated to the media content. By masking the media content, the original media content cannot be clearly displayed to users, and the processed media content can achieve the purpose of making users unable to obtain content-related information from it.

[0088] For example, Figure 6 is a schematic diagram after trace-less processing of a video data stream involved in an embodiment of the present application. As Figure 6 shown, if it is necessary to perform trace-less processing on the video data stream, the video content can be replaced with a specific picture or specific words to achieve the purpose of masking the video data stream. Alternatively, the video content corresponding to the video data stream can also be processed by adding mosaics or watermarks, etc. The method of masking the video content is not limited here.

[0089] If it is necessary to perform trace-less processing on the audio data stream, the audio content can be replaced with a specific audio, or noise can be added to the audio content, or the audio content can be directly discarded. Similarly, the method of masking the audio content is not limited here.

[0090] S103, the computing device sends the processed media data stream to the recording server.

[0091] In the embodiment of the present application, the computing device sends the processed media data stream to the recording server, so that the recording server only records the media content corresponding to the processed media data stream, and does not record the masked media content, thereby ensuring the protection of sensitive content involved in the meeting.

[0092] In a possible implementation manner, the computing device can also send the processed media data stream to the AI server, and the AI server can be used to perform AI intelligent analysis on the media content indicated by the processed media data stream.

[0093] That is to say, since the media content that needs to be processed without a trace has been masked in the processed media data stream, the AI server will not use the masked media content for AI recognition and analysis, thereby ensuring the protection of sensitive content involved in the meeting.

[0094] In a possible implementation, if it is determined that the media data stream includes a trace - free identifier, the computing device may generate status information, which can be used to indicate that the media data stream sent by the first terminal needs to be processed without a trace. Then, the computing device may send the status information to other terminals so that the other terminals can display the specified information to the user, and the specified information can be used to prompt the user that the media data stream sent by the first terminal needs to be processed without a trace.

[0095] For example, Figure 7 is a schematic diagram of the interface display of another participating terminal involved in the embodiments of the present application. As Figure 7 shown, if the computing device determines that the media data stream sent by user A through the corresponding terminal includes a trace - free identifier, the computing device generates status information and sends the status information to other participating user terminals, that is, the terminals corresponding to user B and user C. In the conference interface displayed by the terminal corresponding to user B, the specified information 31 can be displayed at the position corresponding to user A in the user list of the participating conference, so as to achieve the purpose that user B can determine that user A has applied for trace - free processing through the displayed specified information 31.

[0096] In summary, the user terminal that needs to enable the trace - free function adds a trace - free identifier to the media data stream, and sends the media data stream with the added trace - free identifier to the computing device in the audio - video conferencing system. The computing device determines the media data stream in the audio - video conference that needs to be processed without a trace by identifying the trace - free identifier, and sends the media data stream to the recording server after processing it without a trace, so that the recording server can record and shield the media content indicated by the media data stream with the added trace - free identifier. Thus, the controllability of the recording permission of the media data stream sent by the participating user terminals in the audio - video conference is ensured, and the media content that belongs to the private content in the audio - video conference is prevented from being leaked, thereby expanding the privacy protection function of the audio - video conference.

[0097] Taking the application of the embodiments of the present application to an audio - video conferencing system based on the SVC media mode as an example, Figure 8 shows a schematic flowchart of a media data stream processing method provided by an exemplary embodiment of the present application. This media data stream processing method can be executed by an audio - video conferencing system, and the audio - video conferencing system can be Figure 3 or Figure 4 the audio - video conferencing system shown. This media data stream processing method includes the following steps:

[0098] S201, the first terminal sends a media data stream containing a trace - free identifier to the media forwarding server.

[0099] In a possible implementation, the client running the audio and video conference on the first terminal is the first client. The first client can set a stealth application control in the conference interface of the audio and video conference. After the user triggers the stealth application control, that is, after the first terminal receives the trigger operation of the stealth application control through the first client, the first terminal can add a stealth identifier to the corresponding media data stream through the first client.

[0100] Among them, the stealth application control can be used to provide the user with an application to perform stealth processing on the media data stream. The media data stream can be the data transmitted by the first terminal to other participating terminals in the audio and video conference. Other participating terminals can support synchronously displaying the corresponding media content according to the received media data stream. Stealth processing can be used to mask part or all of the media content corresponding to the media data in the media data stream.

[0101] In a possible implementation, selection controls corresponding to different types of media content are also displayed in the conference interface of the audio and video conference. In response to receiving the trigger operation of the stealth application control and receiving the selection operation of the selection control corresponding to the first type of media content, a stealth identifier is added to the header fields of each data packet of the media data stream corresponding to the first type of media content.

[0102] Exemplarily, the user can select to set the media content in the conference, such as audio, video, and sharing, to the stealth mode through the stealth application control displayed on the first terminal and the selection controls corresponding to different types of media content. When the terminal sends the corresponding media data stream, a stealth identifier can be carried in the packet header or in the header field of the data packet, such as incogito = 1.

[0103] For example, Figure 9 is a schematic diagram showing the display of a conference interface involved in an embodiment of the present application. As Figure 9 shown, a stealth mode setting control 40 can be displayed in the conference interface. By triggering the stealth mode setting control 40, a pop-up window or page for setting the stealth mode can be displayed in the conference interface. As Figure 9 shown, a pop-up window for setting the stealth mode is displayed in the conference interface. The pop-up window includes a stealth application control 41. If the user triggers the stealth application control 41, it can be determined that the user A's own media content in the conference includes media content with the stealth mode enabled. The pop-up window for setting the stealth mode in the conference interface can also include selection controls 42 corresponding to different types of media content, such as a selection control for audio, a selection control for video, and a selection control for sharing. If the user selects the selection control for sharing, the first terminal can add a stealth identifier to the data stream corresponding to the shared content.

[0104] S202, The media forwarding server sends the media data stream carrying the trace - free identifier to the recording server.

[0105] S203, The recording server performs trace - free processing on the media content carrying the trace - free identifier.

[0106] S204, The media forwarding server sends the media data stream carrying the trace - free identifier to the media gateway.

[0107] S205, The media gateway performs trace - free processing on the media content carrying the trace - free identifier.

[0108] S206, The media gateway sends the media content after trace - free processing to the AI server.

[0109] Among them, after receiving the media content after trace - free processing, the AI server performs AI intelligent analysis on the processed media content to obtain an analysis result. The AI server sends the analysis result to the media gateway. After receiving the result of AI intelligent analysis on the processed media data stream returned by the AI server, the media gateway sends it to the media forwarding server. The media forwarding server can send this result to other participating terminals, so that this result can be displayed in the audio - video conference.

[0110] S207, The media forwarding server sends the media data stream carrying the trace - free identifier to other participating terminals.

[0111] Among them, the execution order of the above S202, S204, and S207 is not limited.

[0112] Taking the application of the embodiment of the present application to an audio - video conference system based on the AVC media mode as an example, Figure 10 shows a schematic flowchart of a media data stream processing method provided by an exemplary embodiment of the present application. This media data stream processing method can be executed by an audio - video conference system, and this audio - video conference system can be Figure 3 or Figure 4 the audio - video conference system shown. This media data stream processing method includes the following steps:

[0113] S301, The first terminal applies to the conference management server for the first terminal to enter the trace - free state.

[0114] Among them, the first terminal can apply to the conference management server for media content such as audio, video, and sharing to enter the trace - free state.

[0115] S302, The conference management server sends the status information of the first terminal to the media server.

[0116] S303, the media server sends the status information of the first terminal to other participating terminals.

[0117] Among them, the first terminal can display the specified information of the first terminal according to the status information of the first terminal to prompt the user that the first terminal has entered the traceless state.

[0118] S304, the first terminal sends the media data stream to the media server.

[0119] S305, the media server adds a traceless identifier to the media data stream and performs traceless processing.

[0120] S306, the media server sends the media content after traceless processing to the recording server.

[0121] S307, the media server sends the media content after traceless processing to the AI server.

[0122] Among them, after receiving the media content after traceless processing, the AI server performs AI intelligent analysis on the processed media content to obtain the analysis result. The AI server sends the analysis result to the media server. After receiving the result of the AI intelligent analysis of the processed media data stream returned by the AI server, the media server can send the result to other participating terminals, so that the result is displayed in the audio and video conference.

[0123] S308, the media server sends the media data stream to other participating terminals.

[0124] Since other participating terminals have determined that the first terminal has entered the traceless state after S303, they cannot perform operations such as screen recording, screenshotting, and sharing on the corresponding media content after receiving the media data stream sent by the media server.

[0125] S309, the meeting host sends an application for the meeting to enter the traceless state to the meeting management server through the host terminal.

[0126] S310, the meeting management server can determine that the meeting has entered the traceless state and notify the media server.

[0127] S311, the media server adds a traceless identifier to each media data stream of the meeting and performs traceless processing.

[0128] S312, the media server notifies each participating terminal that the meeting is in traceless mode.

[0129] In summary, the user terminal that needs to enable the trace - free function adds a trace - free identifier to the media data stream, and sends the media data stream with the trace - free identifier added to the computing device in the audio - video conferencing system. The computing device determines the media data stream in the audio - video conference that needs to be processed for trace - free by identifying the trace - free identifier, and sends the media data stream after trace - free processing to the recording server, so that the recording server can record and shield the media content indicated by the media data stream with the trace - free identifier added. Thus, the controllability of the recording permission of the media data stream sent by the participating user terminals in the audio - video conference is ensured, the leakage of media content that belongs to private content in the audio - video conference is avoided, and the privacy protection function of the audio - video conference is further extended.

[0130] The above mainly introduces the solution of the embodiment of the present application from the perspective of the method. It can be understood that in order to implement the above functions, the media data stream processing device includes at least one of the corresponding hardware structures and software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0131] The embodiments of the present application can divide the media data stream processing device into functional units according to the above - mentioned method examples. For example, each functional unit can be divided corresponding to each function, or two or more functions can be integrated into one processing unit. The above - integrated unit can be implemented in the form of hardware or in the form of a software functional unit. It should be noted that the division of units in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0132] Exemplarily, Figure 11 FIG. shows a schematic structural diagram of a media data stream processing device 500 provided by an exemplary embodiment of the present application. The media data stream processing device 500 is applied to a computing device, or the media data stream processing device 500 can be a computing device. The media data stream processing device 500 includes:

[0133] An acquisition module 510, configured to acquire a media data stream, where the media data stream is data transmitted by a first terminal to other participating terminals in an audio - video conference, and the other participating terminals are used to display media content corresponding to the media data stream;

[0134] A processing module 520, configured to, if a trace - free identifier is included in the media data stream, perform trace - free processing on the media data stream according to the trace - free identifier to obtain a processed media data stream, where the trace - free processing is used to mask part or all of the media content corresponding to the media data in the media data stream;

[0135] A sending module 530, configured to send the processed media data stream to a recording server, so that the recording server records the media content indicated by the processed media data stream.

[0136] For example, in combination with Figure 3 , the obtaining module 510 may be configured to execute S101 as shown in Figure 5 , the processing module 520 may be configured to execute S102 as shown in Figure 5 , and the sending module 530 may be configured to execute S103 as shown in Figure 5 .

[0137] In a possible implementation manner, the media data stream includes multiple data packets, and the processing module 520 is further configured to determine a target data packet from the multiple data packets, where the target data packet is the data packet that includes the trace - free identifier in the packet header field; perform trace - free processing on the target data packet to obtain the processed media data stream, and the processed media data stream includes the processed target data packet and other data packets in the multiple data packets except the target data packet.

[0138] In a possible implementation manner, the sending module 530 is further configured to send the processed media data stream to an artificial intelligence (AI) server, where the AI server is configured to perform AI intelligent analysis on the media content indicated by the processed media data stream; receive the result of the AI intelligent analysis on the processed media data stream returned by the AI server, and the result is used for display in the audio - video conference.

[0139] In a possible implementation manner, the media data stream includes one or more of the following: at least one of an audio data stream, a video data stream, a shared media stream, or a message stream.

[0140] In a possible implementation manner, if the media data stream is a media stream based on the scalable video coding (SVC) mode, the computing device includes a media stream router (SFU); or, if the media data stream is a media stream based on the advanced video coding (AVC) mode, the computing device includes a media server.

[0141] In a possible implementation manner, the apparatus further includes:

[0142] A generation module, configured to generate status information if it is determined that the media data stream includes a markless identifier, where the status information is used to indicate that the media data stream sent by the first terminal needs to be processed marklessly;

[0143] The sending module 530 is further configured to send the status information to the other terminal, so that the other terminal displays specified information to the user, and the specified information is used to prompt the user that the media data stream sent by the first terminal needs to be processed marklessly.

[0144] In a possible implementation manner, if the status information is further used to indicate the types of media data streams in the media data stream sent by the first terminal that need to be processed marklessly, the specified information is further used to prompt the user of the types of media data streams in the media data stream sent by the first terminal that need to be processed marklessly.

[0145] Exemplarily, Figure 12 FIG. shows a schematic structural diagram of a media data stream processing apparatus 600 provided by an exemplary embodiment of the present application. The media data stream processing apparatus 600 is applied to a first client, or the media data stream processing apparatus 600 may be the first client. The first client is a client of an audio and video conference running on a first terminal. The media data stream processing apparatus 600 includes:

[0146] A setting module 610, configured to set a markless application control in the conference interface of the audio and video conference, where the markless application control is used to provide an application for the user to perform markless processing on the media data stream. The media data stream is data transmitted by the first terminal to other participating terminals in the audio and video conference, and the other participating terminals are used to display the media content corresponding to the media data stream. The markless processing is used to mask the media content corresponding to some or all of the media data in the media data stream;

[0147] An adding module 620, configured to add a markless identifier to the media data stream in response to receiving a trigger operation on the markless application control;

[0148] A sending module 630, configured to send the media data stream with the markless identifier added to a computing device, where the computing device is configured to perform markless processing on the media data stream according to the markless identifier to obtain a processed media data stream, and send the processed media data stream to a recording server, so that the recording server records the media content indicated by the processed media data stream.

[0149] In a possible implementation manner, selection controls corresponding to different types of media content are further displayed in the conference interface of the audio and video conference;

[0150] The adding module 620 is further configured to, in response to receiving a triggering operation on the traceless application control and a selection operation on the selection control corresponding to the media content of the first type, add the traceless identifier to the header fields of each data packet of the media data stream corresponding to the media content of the first type in the media data stream.

[0151] For the specific description of the above optional manner, reference may be made to the foregoing method embodiments, which will not be elaborated herein. In addition, the explanations and descriptions of the beneficial effects of any of the above-provided media data stream processing devices may refer to the corresponding method embodiments above and will not be elaborated.

[0152] Among them, the obtaining module 510, the processing module 520, and the sending module 530 can all be implemented by software or by hardware. Exemplarily, next, taking the obtaining module 510 as an example, the implementation manner of the obtaining module 510 will be introduced. Similarly, the implementation manners of the processing module 520 and the sending module 530 can refer to the implementation manner of the obtaining module 510.

[0153] As an example of a software functional unit, the obtaining module 510 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the obtaining module 510 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running this code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically close data centers. Among them, generally, one region may include multiple AZs.

[0154] Similarly, the multiple hosts / virtual machines / containers for running this code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, generally, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.

[0155] As an example of a hardware functional unit, the obtaining module 510 may include at least one computing device, such as a server or the like. Alternatively, the obtaining module 510 may also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0156] The multiple computing devices included in the obtaining module 510 may be distributed in the same region or in different regions. The multiple computing devices included in the obtaining module 510 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the obtaining module 510 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0157] It should be noted that in other embodiments, the obtaining module 510 may be used to execute any step in the media data stream processing method, the processing module 520 may be used to execute any step in the media data stream processing method, and the sending module 530 may be used to execute any step in the media data stream processing method. The steps to be implemented by the obtaining module 510, the processing module 520, and the sending module 530 can be specified as needed. The entire function of the media data stream processing device is realized by respectively implementing different steps in the media data stream processing method through the obtaining module 510, the processing module 520, and the sending module 530. This application also provides a computing device 100. As Figure 13 shown, the computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.

[0158] The bus 102 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 13 only one line is used in Figure 13 , but this does not mean that there is only one bus or one type of bus. The bus 104 can include paths for transmitting information between various components of the computing device 100 (for example, the memory 106, the processor 104, and the communication interface 108).

[0159] The processor 104 can include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0160] The memory 106 can include volatile memory, such as random access memory (RAM). The processor 104 can also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0161] The memory 106 stores executable program code, and the processor 104 executes the executable program code to implement the functions of the aforementioned acquisition module, processing module 520, and sending module 530 respectively, thereby implementing the media data stream processing method. That is, the memory 106 stores instructions for executing the media data stream processing method.

[0162] Alternatively, the memory 106 stores executable code, and the processor 104 executes the executable code to implement the functions of the aforementioned media data stream processing device respectively, thereby implementing the media data stream processing method. That is, the memory 106 stores instructions for executing the media data stream processing method.

[0163] The communication interface 103 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or communication networks.

[0164] The embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0165] As Figure 14 shown, the computing device cluster includes at least one computing device 100. Instructions for executing the media data stream processing method may be stored in the same manner in the memories 106 of one or more of the computing devices 100 in the computing device cluster.

[0166] In some possible implementation manners, partial instructions for executing the media data stream processing method may also be stored separately in the memories 106 of one or more of the computing devices 100 in the computing device cluster. In other words, a combination of one or more computing devices 100 may jointly execute the instructions for executing the media data stream processing method.

[0167] It should be noted that the memories 106 in different computing devices 100 in the computing device cluster may store different instructions, respectively for executing partial functions of the media data stream processing apparatus. That is, the instructions stored in the memories 106 of different computing devices 100 may implement the functions of one or more of the obtaining module 510, the processing module 520, and the sending module 530.

[0168] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 15 Shows a possible implementation manner. As Figure 15 shown, two computing devices 100A and 100B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manner, instructions for executing the function of the obtaining module 510 are stored in the memory 106 of the computing device 100A. At the same time, instructions for executing the functions of the processing module 520 and the sending module 530 are stored in the memory 106 of the computing device 100B.

[0169] Figure 15 The connection manner between the computing device clusters shown may be considered that since the media data stream processing method provided in the present application requires a large amount of data storage and data calculation, the functions implemented by the processing module 520 and the sending module 530 are considered to be executed by the computing device 100B.

[0170] It should be understood that Figure 15The functions of the computing device 100A shown can also be completed by multiple computing devices 100. Similarly, the functions of the computing device 100B can also be completed by multiple computing devices 100.

[0171] The embodiments of the present application also provide another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similarly referred to Figure 14 and Figure 15 the connection method of the computing device cluster. The difference is that the same instructions for executing the media data stream processing method may be stored in the memory 106 of one or more of the computing devices 100 in the computing device cluster.

[0172] In some possible implementation manners, the memory 106 of one or more of the computing devices 100 in the computing device cluster may also separately store partial instructions for executing the media data stream processing method. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the media data stream processing method.

[0173] It should be noted that the memories 106 in different computing devices 100 in the computing device cluster may store different instructions for executing partial functions of the data processing system. That is, the instructions stored in the memories 106 of different computing devices 100 can implement the functions of one or more of the devices in the media data stream processing apparatus.

[0174] The embodiments of the present application also provide a computer program product including instructions. The computer program product may be software or a program product including instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, at least one computing device is caused to execute the media data stream processing method.

[0175] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center including one or more available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that instruct the computing device to execute the media data stream processing method, or instruct the computing device to execute the media data stream processing method.

[0176] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing media data streams, characterized in that, Applied to a computing device, the method includes: Obtain a media data stream, which is data transmitted by a first terminal to other participating terminals during an audio-video conference, and the other participating terminals are used to display media content corresponding to the media data stream; If the media data stream includes a trace-free identifier, perform trace-free processing on the media data stream according to the trace-free identifier to obtain a processed media data stream, and the trace-free processing is used to mask part or all of the media content corresponding to the media data in the media data stream; Send the processed media data stream to a recording server so that the recording server records the media content indicated by the processed media data stream.

2. The method according to claim 1, characterized in that, The media data stream includes multiple data packets, and performing trace-free processing on the media data stream according to the trace-free identifier to obtain a processed media data stream includes: Determine a target data packet from the multiple data packets, and the target data packet is the data packet that includes the trace-free identifier in the packet header field; Perform trace-free processing on the target data packet to obtain the processed media data stream, and the processed media data stream includes the target data packet that has undergone trace-free processing and other data packets in the multiple data packets except the target data packet.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Send the processed media data stream to an artificial intelligence (AI) server, and the AI server is used to perform AI intelligent analysis on the media content indicated by the processed media data stream; Receive the result of the AI intelligent analysis of the processed media data stream returned by the AI server, and the result is used to be displayed during the audio-video conference.

4. The method according to any one of claims 1 to 3, characterized in that, The media data stream includes one or more of the following: audio data stream, video data stream, shared media stream, message stream.

5. The method according to any one of claims 1 to 4, characterized in that If the media data stream is a media stream based on the scalable video coding (SVC) mode, the computing device includes a media stream router (SFU); or, if the media data stream is a media stream based on the advanced video coding (AVC) mode, the computing device includes a media server.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: If it is determined that the media data stream includes a trace-free identifier, generate status information, and the status information is used to indicate that the media data stream sent by the first terminal needs to be processed without a trace; Send the status information to the other terminals so that the other terminals display specified information to the user, and the specified information is used to prompt the user that the media data stream sent by the first terminal needs to be processed without a trace.

7. The method according to claim 6, wherein If the status information is further used to indicate the type of the media data stream that needs to be processed without a trace in the media data stream sent by the first terminal, the specified information is further used to prompt the user of the type of the media data stream that needs to be processed without a trace in the media data stream sent by the first terminal.

8. A method for processing media data stream, characterized in that, Applied to a first client, the first client is a client of an audio-video conference running on a first terminal, and the method includes: Set a traceless application control in the conference interface of the audio and video conference. The traceless application control is used to provide a user with an application to perform traceless processing on the media data stream. The media data stream is data transmitted by the first terminal to other participating terminals in the audio and video conference, and the other participating terminals are used to display the media content corresponding to the media data stream. The traceless processing is used to mask the media content corresponding to some or all of the media data in the media data stream; In response to receiving a trigger operation on the traceless application control, add a traceless identifier to the media data stream; Send the media data stream with the traceless identifier added to a computing device. The computing device is used to perform traceless processing on the media data stream according to the traceless identifier to obtain a processed media data stream, and send the processed media data stream to a recording server, so that the recording server records the media content indicated by the processed media data stream.

9. The method according to claim 8, wherein Selection controls corresponding to different types of media content are also displayed in the conference interface of the audio and video conference; The adding a traceless identifier to the media data stream in response to receiving a trigger operation on the traceless application control includes: In response to receiving a trigger operation on the traceless application control and a selection operation on the selection control corresponding to the first type of media content, add the traceless identifier to the header fields of each data packet of the media data stream corresponding to the first type of media content in the media data stream.

10. A media data stream processing device, characterized in that, Applied to a computing device, the apparatus includes: An acquisition module, configured to acquire a media data stream. The media data stream is data transmitted by a first terminal to other participating terminals in an audio and video conference, and the other participating terminals are used to display the media content corresponding to the media data stream; A processing module, configured to, if the media data stream includes a traceless identifier, perform traceless processing on the media data stream according to the traceless identifier to obtain a processed media data stream. The traceless processing is used to mask the media content corresponding to some or all of the media data in the media data stream; A sending module, configured to send the processed media data stream to a recording server, so that the recording server records the media content indicated by the processed media data stream.

11. The device according to claim 10, wherein, The media data stream includes a plurality of data packets. The processing module is further configured to determine a target data packet from the plurality of data packets. The target data packet is the data packet that includes the traceless identifier in the packet header field; Perform traceless processing on the target data packet to obtain the processed media data stream. The processed media data stream includes the target data packet that has undergone traceless processing and the other data packets in the plurality of data packets except the target data packet.

12. The device according to claim 10 or 11, characterized in that, The sending module is further configured to send the processed media data stream to an artificial intelligence (AI) server, where the AI server is configured to perform AI intelligent analysis on the media content indicated by the processed media data stream; and receive the result of the AI intelligent analysis on the processed media data stream returned by the AI server, where the result is used for display in the audio-video conference.

13. The device according to any one of claims 10 to 12, characterized in that, The media data stream includes one or more of the following: audio data stream, video data stream, shared media stream, and message stream.

14. The device according to any one of claims 10 to 13, characterized in that If the media data stream is a media stream based on the scalable video coding (SVC) mode, the computing device includes a media stream router (SFU); or, if the media data stream is a media stream based on the advanced video coding (AVC) mode, the computing device includes a media server.

15. The device according to any one of claims 10 to 14, characterized in that, The device further includes: A generating module, configured to generate status information if it is determined that the media data stream includes a trace-free identifier, where the status information is used to indicate that the media data stream sent by the first terminal needs to be processed without a trace; The sending module is further configured to send the status information to the other terminals, so that the other terminals display specified information to the user, where the specified information is used to prompt the user that the media data stream sent by the first terminal needs to be processed without a trace.

16. The device according to claim 15, characterized in that, If the status information is further used to indicate the type of the media data stream that needs to be processed without a trace in the media data stream sent by the first terminal, the specified information is further used to prompt the user of the type of the media data stream that needs to be processed without a trace in the media data stream sent by the first terminal.

17. A media data stream processing device, characterized in that, Applied to a first client, where the first client is a client of an audio-video conference running on a first terminal, the device includes: A setting module, configured to set a trace-free application control in the conference interface of the audio-video conference, where the trace-free application control is used to provide the user with an application to process the media data stream without a trace, and the media data stream is data transmitted by the first terminal to other participating terminals in the audio-video conference, and the other participating terminals are used to display the media content corresponding to the media data stream, and the trace-free processing is used to mask part or all of the media content corresponding to the media data in the media data stream; An adding module, configured to add a trace-free identifier to the media data stream in response to receiving a trigger operation on the trace-free application control; A sending module, configured to send the media data stream with the added trace-free identifier to a computing device, where the computing device is configured to perform trace-free processing on the media data stream according to the trace-free identifier to obtain a processed media data stream, and send the processed media data stream to a recording server, so that the recording server records the media content indicated by the processed media data stream.

18. The device according to claim 17, characterized in that, Selection controls corresponding to different types of media content are further displayed in the conference interface of the audio-video conference; The adding module is further configured to, in response to receiving a trigger operation on the traceless application control and a selection operation on the selection control corresponding to the media content of the first type, add the traceless identifier to the header fields of the data packets of the media data stream corresponding to the media content of the first type in the media data stream.

19. A cluster of computing devices, characterized in that, It includes at least one computing device, and each computing device includes a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the media data stream processing method according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that, It includes computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the media data stream processing method according to any one of claims 1 to 9.

21. A computer program product, characterized in that, The computer program product includes instructions. When the instructions are run by a computing device cluster, the computing device cluster is caused to execute the media data stream processing method according to any one of claims 1 to 9.