Media data stream processing method and apparatus, cluster, medium, and program product
By adding traceless logos and performing traceless processing in audio and video conferencing, the problem of uncontrollable user media data stream recording permissions is solved, and the privacy protection of audio and video conferencing is realized to prevent the leakage of sensitive content.
Patent Information
- Application Number
- PCT/CN2024/141193
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-22
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-17
AI Technical Summary
In audio and video conferencing, users' recording permissions for media data streams are not controllable, resulting in privacy or sensitive content that may be leaked, and the prior art cannot effectively protect user privacy.
By adding traceless identifiers to the media data stream, identifying and processing them with computing devices, blocking sensitive content, ensuring that the recording server only records unidentified media content and avoiding private content being leaked.
It realizes controllability of recording permissions for user media data streams in audio and video conferencing, prevents private content from being leaked, and expands the privacy protection function of audio and video conferencing.
Smart Images

Figure CN2024141193_17072025_PF_FP_ABST
Abstract
Description
Media data stream processing method, device, cluster, medium and program product
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 9, 2024, with application number 202410032624.5 and application name “A method and device for seamless communication in meetings”, the entire contents of which are incorporated by reference into this application.
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on March 22, 2024, with application number 202410342096.3 and application name “Media data stream processing method, device, cluster, medium and program product”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of data processing technology, and in particular to a method, device, cluster, medium, and program product for processing media data streams. Background Art
[0004] With the continuous development of internet technology, online audio and video conferencing allows users in different locations to access conferences via the internet, enabling online meetings regardless of location. Online audio and video conferencing also offers a variety of post-meeting review features, such as recording in-meeting audio and video, taking intelligent minutes of meeting content, and recording chat logs. However, some of the media content shared by participants in audio and video conferences may contain sensitive content, and participants may not want sensitive content to be recorded or stored after the meeting, potentially allowing it to be disseminated afterward.
[0005] Currently, the conference host has the ability to pause recording or edit and delete relevant media content after the meeting. This means that if a non-host participant's media content contains sensitive content, they will need to notify the host to stop recording or delete the content after the meeting.
[0006] In related technologies, notifying the host to stop recording at a certain point or deleting the corresponding media content after the meeting does not guarantee that the host is fully aware of the media content processing requirements of the participating users, nor does it guarantee that the host can correctly complete the corresponding operations. Because the recording permissions of the media data streams sent by the terminal in the audio and video conference are not controllable by the participating users, private or sensitive media content in the audio and video conference may be leaked, which in turn makes the post-conference review function of the audio and video conference unable to meet the needs of users. Summary of the Invention
[0007] The embodiments of the present application provide a media data stream processing method, device, cluster, medium and program product, which ensure the controllability of the recording permissions of the media data streams sent by the participating user terminals in the audio and video conference, avoid the leakage of media content that is private content in the audio and video conference, and thus expand the privacy protection function of the audio and video conference.
[0008] In a first aspect, the present application provides a media data stream processing method, which is applied to a computing device, and the method includes: obtaining a media data stream, which is data transmitted by a first terminal to other participating terminals in an audio and video conference, and the other participating terminals are used to display the media content corresponding to the media data stream; if the media data stream includes a traceless identifier, the media data stream is tracelessly processed according to the traceless identifier to obtain a processed media data stream, and the traceless processing is used to shield the media content corresponding to part or all of the media data in the media data stream; and sending the processed media data stream to a recording server so that the recording server records the media content indicated by the processed media data stream.
[0009] It is understandable that a user terminal that needs to enable the traceless function adds a traceless identifier to the media data stream and sends the media data stream with the traceless identifier to the computing device in the audio and video conferencing system. The computing device, by identifying the traceless identifier, determines the media data stream in the audio and video conference that requires traceless processing, and sends the media data stream after traceless processing to the recording server, so that the recording server can record the media content indicated by the media data stream with the traceless identifier. This ensures the controllability of the recording permissions of the media data streams sent by the participating user terminals in the audio and video conference, prevents the leakage of private media content in the audio and video conference, and further expands the privacy protection function of the audio and video conference.
[0010] In one possible implementation, a media data stream includes multiple data packets, and the media data stream is processed seamlessly according to a seamless identifier to obtain a processed media data stream, including: determining a target data packet from multiple data packets, where the target data packet is a data packet containing a seamless identifier in a packet header field; processing the target data packet seamlessly to obtain a processed media data stream, where the processed media data stream includes the target data packet that has been processed seamlessly and other data packets among the multiple data packets except the target data packet.
[0011] It can be understood that by determining whether the message header field of the data message contains a traceless identifier, the target data message in the media data stream that needs to be processed tracelessly can be determined, thereby achieving traceless processing of the media content corresponding to the target data message. The media data stream can be finely divided into media content that needs to be processed tracelessly and media content that does not need to be processed tracelessly, thereby improving the processing effect of traceless processing.
[0012] In one possible implementation, the method further includes: sending the processed media data stream to an artificial intelligence (AI) server, where the AI server is used to perform AI intelligent analysis on the media content indicated by the processed media data stream; and receiving the results of the AI intelligent analysis of the processed media data stream returned by the AI server, where the results are used to be displayed in the audio and video conference.
[0013] It can be understood that by sending the processed media data stream to the AI server, it can be ensured that the media content that needs to be processed seamlessly will not be subjected to AI intelligent analysis by the AI server, thereby ensuring the privacy of the media content.
[0014] In a possible implementation, the media data stream includes one or more of the following: an audio data stream, a video data stream, a shared media stream, and a message stream.
[0015] In one possible implementation, if the media data stream is a media stream based on the scalable video coding SVC mode, the computing device includes a media stream router SFU; or, if the media data stream is a media stream based on the advanced video coding AVC mode, the computing device includes a media server.
[0016] In one possible implementation, the method also includes: if it is determined that the media data stream includes a traceless identifier, generating status information, the status information is used to indicate that the media data stream sent by the first terminal needs to be processed tracelessly; sending status information to other terminals so that the other terminals display specified information to the user, the specified information is used to prompt the user that the media data stream sent by the first terminal needs to be processed tracelessly.
[0017] It is understandable that by sending status information to other terminals, it can be ensured that users on the other terminal sides can obtain the user terminal that has enabled traceless in the audio and video conference, thereby improving the user experience in the audio and video conference.
[0018] In one possible implementation, if the status information is also used to indicate the type of media data stream in the media data stream sent by the first terminal that needs to be seamlessly processed, the designated information is also used to prompt the user of the type of media data stream in the media data stream sent by the first terminal that needs to be seamlessly processed.
[0019] It is understandable that the status information received by other terminals may also include the type of media data stream to be processed seamlessly, which can further prompt the user and further improve the user experience in the audio and video conference.
[0020] In the second aspect, the present application provides a media data stream processing method, which is applied to a first client, and the first client is a client of an audio and video conference running on a first terminal. The method includes: displaying a traceless application control in the conference interface of the audio and video conference, and the traceless application control is used to provide the user with an application for traceless processing of the media data stream. The media data stream is the data transmitted by the first terminal to other participating terminals in the audio and video conference, and the other participating terminals are used to display the media content corresponding to the media data stream. The traceless processing is used to shield the media content corresponding to part or all of the media data in the media data stream; in response to receiving a trigger operation for the traceless application control, adding a traceless identifier to the media data stream; sending the media data stream with the traceless identifier added to the computing device, the computing device is used to tracelessly process the media data stream according to the traceless identifier to obtain the processed media data stream, and send the processed media data stream to the recording server, so that the recording server records the media content indicated by the processed media data stream.
[0021] It is understandable that a user terminal that needs to enable the traceless function adds a traceless identifier to the media data stream and sends the media data stream with the traceless identifier to the computing device in the audio and video conferencing system. The computing device, by identifying the traceless identifier, determines the media data stream in the audio and video conference that requires traceless processing, and sends the media data stream after traceless processing to the recording server, so that the recording server can record the media content indicated by the media data stream with the traceless identifier. This ensures the controllability of the recording permissions of the media data streams sent by the participating user terminals in the audio and video conference, prevents the leakage of private media content in the audio and video conference, and further expands the privacy protection function of the audio and video conference.
[0022] In one possible implementation, the conference interface of the audio and video conference also displays selection controls corresponding to different types of media content; in response to receiving a trigger operation for the traceless application control, a traceless identifier is added to the media data stream, including: in response to receiving a trigger operation for the traceless application control and receiving a selection operation for the selection control corresponding to the first type of media content, a traceless identifier is added to the header field of each data message of the media data stream corresponding to the first type of media content in the media data stream.
[0023] It can be understood that by displaying selection controls corresponding to different types of media content, the purpose of adding invisible labels for different types of media content can be achieved, thereby refining the media content to be processed invisible and improving the user experience in audio and video conferences.
[0024] In a third aspect, an embodiment of the present application provides a media data stream processing device, which is used to execute any one of the media data stream processing methods provided in the first aspect.
[0025] In one possible implementation, the embodiment of the present application can divide the media data stream processing device into functional modules according to the method provided in the first aspect above. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. Exemplarily, the embodiment of the present application can divide the media data stream processing device into an acquisition module, a processing module, and a sending module, etc. according to the function. The description of the possible technical solutions and beneficial effects executed by each of the functional modules divided above can refer to the technical solutions provided in the first aspect above or its corresponding possible implementation, and will not be repeated here.
[0026] In a fourth aspect, an embodiment of the present application provides a media data stream processing device, which is used to execute any one of the media data stream processing methods provided in the second aspect above.
[0027] In one possible implementation, the embodiment of the present application can divide the media data stream processing device into functional modules according to the method provided in the second aspect above. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. Exemplarily, the embodiment of the present application can divide the media data stream processing device into a setting module, an adding module, and a sending module, etc. according to the function. The description of the possible technical solutions and beneficial effects executed by each of the functional modules divided above can refer to the technical solutions provided in the second aspect above or its corresponding possible implementation, and will not be repeated here.
[0028] In a fifth aspect, an embodiment of the present application provides a computing device, which includes a processor and a memory, and the processor is coupled to the memory; the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to enable the computing device to implement the media data stream processing method described in the above aspects.
[0029] In the sixth aspect, an embodiment of the present application provides a computing device cluster, which includes at least one computing device, each computing device including: a processor and a memory, the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the media data stream processing method provided in the various optional implementation methods of the first aspect or the second aspect above.
[0030] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, which stores at least one computer program instruction, and the computer program instruction is loaded and executed by a processor to implement the media data stream processing method as described in the above aspects.
[0031] In an eighth aspect, embodiments of the present application provide a computer program product, comprising computer instructions stored in a computer-readable storage medium. A processor of a computing device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computing device cluster to perform the media data stream processing method provided in various optional implementations of the first or second aspects.
[0032] For the specific description of the third to eighth aspects and their various implementations in this application, reference may be made to the detailed description in the first aspect and its various implementations or the second aspect and its various implementations; and for the beneficial effects of the third to eighth aspects and their various implementations, reference may be made to the analysis of the beneficial effects in the first aspect and its various implementations or the second aspect and its various implementations, which will not be repeated here.
[0033] These and other aspects of the present application will become more readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] FIG1 is a schematic diagram of an audio and video conferencing system based on an SVC media mode according to an exemplary embodiment;
[0035] FIG2 is a schematic diagram of an audio and video conferencing system based on the AVC media mode according to an exemplary embodiment;
[0036] FIG3 is a schematic diagram showing a scenario of media data stream processing according to an exemplary embodiment;
[0037] FIG4 is a schematic diagram showing a scenario of media data stream processing according to an exemplary embodiment;
[0038] FIG5 is a flow chart showing a method for processing a media data stream according to an exemplary embodiment;
[0039] FIG6 is a schematic diagram of a video data stream involved in the embodiment shown in FIG5 after seamless processing;
[0040] FIG7 is a schematic diagram of an interface display of another participating terminal involved in the embodiment shown in FIG5;
[0041] FIG8 is a schematic flow chart showing a method for processing a media data stream according to an exemplary embodiment;
[0042] FIG9 is a schematic diagram showing a conference interface according to the embodiment shown in FIG8 ;
[0043] FIG10 is a schematic flow chart showing a method for processing a media data stream according to an exemplary embodiment;
[0044] FIG11 is a schematic structural diagram of a media data stream processing device according to an exemplary embodiment;
[0045] FIG12 is a schematic structural diagram of a media data stream processing device according to an exemplary embodiment;
[0046] FIG13 is a schematic diagram of a computing device according to an exemplary embodiment;
[0047] FIG14 is a schematic diagram showing a computing device cluster according to an exemplary embodiment;
[0048] FIG15 is a schematic diagram showing a connection method between computing device clusters according to an exemplary embodiment. DETAILED DESCRIPTION
[0049] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0050] In this document, "plurality" refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.
[0051] Furthermore, in the description of this application, unless otherwise specified, "plurality" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0052] In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit differences. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.
[0053] First, the application scenarios of the embodiments of the present application are exemplarily introduced.
[0054] With the continuous development of internet technology, online audio and video conferencing allows users in different locations to access the conference via the internet, enabling online meetings regardless of location. Online audio and video conferencing also offers a variety of post-meeting review features, such as recording audio and video, taking intelligent minutes of meeting content, and recording chat logs.
[0055] Current audio and video conferencing is implemented through an audio and video conferencing system, which can be an audio and video conferencing system based on the scalable video coding (SVC) media mode or an audio and video conferencing system based on the advanced video coding (AVC) media mode.
[0056] Figure 1 shows a schematic diagram of an audio and video conferencing system based on the SVC media mode provided by an embodiment of the present application. As shown in Figure 1, the audio and video conferencing system includes multiple participating user terminals 10, a media forwarding server 11, a media gateway 12, an intelligent recognition server 13, and a recording server 14.
[0057] The participating user terminal 10 is a user terminal used by each user account participating in the audio and video conference to log in. The participating user terminal 10 can be a computer device, a smart phone or other terminal device, which is not limited here.
[0058] The media forwarding server 11 may be used to receive subscription forwarding of media data streams, and the media forwarding server 11 may not encode the media data streams. For example, the media forwarding server 11 may be a media stream routing unit (SFU).
[0059] The media gateway 12 can be used to encode and decode media data streams into other formats. For example, an H264-encoded video data stream can be saved as an image format and sent to an artificial intelligence (AI) service for image recognition; or an audio data stream can be saved as a pulse code modulation (PCM) raw audio file and sent to an AI service for speech recognition.
[0060] The intelligent recognition server 13 can be used to receive the data files encoded and transmitted by the media gateway 12, and perform artificial intelligence analysis on the received data files, such as image recognition or speech recognition.
[0061] Recording server 14 can be used to record the content of audio and video conferences. Media forwarding server 11 sends media data streams to recording server 14, which can record the media data streams, including video data streams, audio data streams, shared media streams, or message streams, into video files, audio files, etc.
[0062] For example, when any participating user terminal 10 in the audio and video conferencing system sends a media data stream to the media forwarding server 11, on the one hand, the media forwarding server 11 can forward the media data stream to other participating user terminals 10 so that other participating user terminals 10 can synchronously display the media content indicated by the media data stream. On the other hand, the media forwarding server 11 can also send the media data stream to the media gateway 12, which encodes and decodes the media data stream into an image file or an audio file, etc., and sends the processed image file or audio file to the intelligent recognition server 13, so that the intelligent recognition server 13 can perform image recognition on the image file and voice recognition on the audio file. In addition, the media forwarding server 11 can also send the media data stream to the recording server 14, and the recording server 14 records the corresponding media content according to the media data stream to obtain a video file, an audio file or a shared file, etc.
[0063] In addition, Figure 2 shows a schematic diagram of an audio and video conferencing system based on the AVC media mode provided by an embodiment of the present application. As shown in Figure 2, the audio and video conferencing system includes multiple participating user terminals 10, a media server 15, a conference management service 16, an intelligent recognition server 13, and a recording server 14.
[0064] The participating user terminal 10 is a user terminal used by each user account participating in the audio and video conference to log in. The participating user terminal 10 can be a computer device, a smart phone or other terminal device, which is not limited here.
[0065] The media server 15 can be used to perform encoding and decoding processing on the received media data stream, convert the media data stream into a format required by other services and terminals, and forward it. For example, the media server 15 can be a multipoint control unit (MCU).
[0066] The conference management service 16 can be used to provide conference reservation and conference control functions for the audio and video conferencing system. It can be provided to the client for access in the form of an application programming interface (API), or provide an access interface in the form of a portal, that is, provide a conference interface to the participating user terminals.
[0067] The intelligent recognition server 13 can be used to receive the data files encoded and transmitted by the media server 15, and perform artificial intelligence analysis on the received data files, such as image recognition or voice recognition.
[0068] Recording server 14 can be used to record the content of the audio and video conference. Media server 15 sends the media data stream to recording server 14, which can record the media data stream, including video data stream, audio data stream, shared media stream or message stream, into video files, audio files, etc.
[0069] For example, when any participating user terminal 10 in the audio and video conferencing system sends a media data stream to the media server 15, the media server 15 can forward the media data stream to the other participating user terminals 10 so that the other participating user terminals 10 can synchronously display the media content indicated by the media data stream. Furthermore, the media server 15 can also encode and decode the media data stream into an image file or an audio file, and send the processed image file or audio file to the intelligent recognition server 13, so that the intelligent recognition server 13 can perform image recognition on the image file and speech recognition on the audio file. Furthermore, the media server 15 can also send the encoded and decoded video file or audio file to the recording server 14, which records the corresponding media content of the video file or audio file to obtain video, audio, or shared content. When any participating user terminal 10 in the audio and video conferencing system sends a media data stream to the media server 15, the participating user terminal 10 can also send the media data stream to the conference management service 16, so that the conference management service provides a conference interface to each participating user terminal 10.
[0070] Through the above-mentioned audio and video conferencing system, each participating user terminal can complete the meeting through online audio and video synchronous transmission, and can also realize the post-meeting review function to review the complete conference video, audio, or shared files. However, due to the problem that some of the media content of each participating user terminal in the audio and video conference may contain sensitive content, that is, the media content sent by the participating user terminal during the meeting may not be retained by the user, and the participating user does not want the sensitive content to be recorded or saved after the meeting and disseminated after the meeting.
[0071] Currently, the conference host in an audio and video conference has the authority to pause recording or edit and delete the corresponding media content after the meeting. In other words, if some of the media content in the audio and video conference of a non-host participant is sensitive, the participant needs to notify the host to stop recording or delete the media content after the meeting. By notifying the host to stop recording at a certain point or notifying the host to delete the corresponding media content after the meeting, it cannot be guaranteed that the host is fully aware of the participant's need to process the media content, nor can it be guaranteed that the host can correctly complete the corresponding operation. Since the participant in the audio and video conference has no control over the recording authority of the media data stream sent through the terminal, the media content that is private or sensitive in the audio and video conference may be leaked, which in turn causes the post-meeting review function of the audio and video conference to fail to meet the needs of users.
[0072] In view of this, user terminals that need to enable the traceless function add a traceless identifier to the media data stream and send the media data stream with the traceless identifier to the computing device in the audio and video conferencing system. The computing device identifies the traceless identifier, determines the media data stream in the audio and video conference that requires traceless processing, and sends the media data stream after traceless processing to the recording server. This allows the recording server to block the media content indicated by the media data stream with the traceless identifier. In other words, the recording server will not record the media content indicated by the media data stream with the traceless identifier, and the recording server can normally record the media content indicated by the media data stream without the traceless identifier. This ensures the controllability of the recording permissions of the media data streams sent by the participating user terminals in the audio and video conference, prevents the leakage of private media content in the audio and video conference, and further expands the privacy protection function of the audio and video conference.
[0073] Among them, Figure 3 shows a scenario diagram of a media data stream processing provided by an embodiment of the present application. As shown in Figure 3, for an audio and video conferencing system based on the SVC media mode, in actual application, the user terminal 10 that starts the traceless function can send the media data stream with the traceless identifier to the media forwarding server 11, and the media forwarding server 11 can forward the media data stream with the traceless identifier to other participating user terminals 10. After receiving the media data stream with the traceless identifier, the other participating user terminals 10 cannot perform local recording or screenshot, sharing, etc. on the media data stream with the traceless identifier to retain the media content indicated by the traceless identifier. On the other hand, the media forwarding server 11 can also send the media data stream with the traceless identifier to the recording server 14. Since the media data stream received by the recording server 14 contains the traceless identifier, the recording server 14 cannot perform cloud recording or screenshot, sharing, etc. on the media content corresponding to the media data stream with the traceless identifier to retain the media content indicated by the traceless identifier.
[0074] The computing device may be a media forwarding server 11 .
[0075] In addition, Figure 4 shows a scenario diagram of a media data stream processing provided by an embodiment of the present application. As shown in Figure 4, for an audio and video conferencing system based on the AVC media mode, in actual application, a user terminal 10 that needs to start the traceless function can send a traceless application to the conference management service 16. After receiving the traceless application, the conference management service 16 adds a traceless identifier to the media data stream corresponding to the user terminal that needs to start the traceless function. The conference management service 16 sends the media data stream with the traceless identifier to the media server 15. The media server 15 can encode and decode the media data stream with the traceless identifier to obtain media content that does not contain the traceless identifier. If there is media content without the traceless identifier, the media content will be sent to the recording server 14 and the intelligent recognition server 13. Otherwise, the media content will not be sent to the recording server 14 and the intelligent recognition server 13. In addition, the media server 15 will send the media data stream with the invisible identifier to the other user terminals 10 participating in the meeting, so that the other user terminals 10 participating in the meeting can display the media content normally and synchronously, and ensure that the user terminals 10 cannot locally record or take screenshots, share, or perform other operations on the media content with the invisible identifier to retain the media content indicated by the invisible identifier.
[0076] The computing device may be a media server 15 .
[0077] It should be noted that the application scenarios and system architectures described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0078] For ease of understanding, the media data stream processing method provided in the present application is exemplarily introduced below with reference to the accompanying drawings. The media data stream processing method is applicable to the audio and video conferencing system shown in FIG3 and FIG4.
[0079] FIG5 is a flow chart showing a method for processing a media data stream according to an exemplary embodiment of the present application. The method can be executed by a computing device, which can be the media forwarding server 11 shown in FIG3 or the media server 15 shown in FIG4 . The method includes the following steps:
[0080] S101: A computing device obtains a media data stream.
[0081] In an embodiment of the present application, a computing device may receive a media data stream sent by a first terminal. The media data stream may be data transmitted by the first terminal to other participating terminals in an audio and video conference, and the other participating terminals are used to display media content corresponding to the media data stream.
[0082] The media data stream may include at least one of an audio data stream, a video data stream, a shared media stream, or a message stream. A shared media stream may be a data stream for sharing screen content or files via the first terminal in an audio or video conference. A message stream may be a data stream corresponding to instant messages sent by the user via the first terminal in an audio or video conference.
[0083] In a possible implementation, the media content corresponding to the media data stream can be divided into audio content, video content, shared content, or message content according to the type. The media data stream is divided into audio data stream, video data stream, shared media stream, or message stream according to the type of the indicated media content.
[0084] For example, participating users can use the first terminal to send the video content, audio content, and shared screen and file data of their speeches in the audio and video conference in the form of a data stream to a computing device in the audio and video conferencing system, which can be a media forwarding server or a media server.
[0085] S102: If the media data stream includes a traceless identifier, the computing device performs traceless processing on the media data stream according to the traceless identifier to obtain a processed media data stream.
[0086] In an embodiment of the present application, after receiving a media data stream, the computing device can detect whether each data packet in the data stream includes a traceless identifier. If it is determined that the traceless identifier is included, the media data stream including the traceless identifier can be processed tracelessly to obtain a processed media data stream.
[0087] The traceless identifier may be the identifier "incogito=1" added to the header of the data message. That is, when the computing device detects the presence of "incogito=1" in the header of the data message, it determines that the data message has been added with the traceless identifier.
[0088] In one possible implementation, a computing device may determine a target data packet from a plurality of data packets, where the target data packet may be a data packet containing a traceless identifier in a packet header field, and then the computing device may perform traceless processing on the target data packet to obtain a processed media data stream, which may include the target data packet that has been tracelessly processed and other data packets from the plurality of data packets except the target data packet.
[0089] Among them, traceless processing can be used to shield the media content corresponding to part or all of the media data in the media data stream. Shielding is a processing method that replaces or overwrites the media content so that the user cannot obtain information related to the media content from the processed media content. Specifically, shielding can be replacing the media content with preset specific content, or overwriting the media content with preset specific content, and the preset specific content can be content unrelated to the media content. By shielding the media content, the original media content cannot be clearly displayed to the user, and the shielded media content can achieve the purpose of preventing the user from obtaining content-related information.
[0090] For example, Figure 6 is a schematic diagram of a video data stream according to an embodiment of the present application after traceless processing. As shown in Figure 6, if a video data stream needs to be traceless processed, the video content can be replaced with a specific image or specific text to achieve the purpose of shielding the video data stream. Alternatively, the video content corresponding to the video data stream can be processed by adding mosaics or adding watermarks, etc. The method of shielding the video content is not limited here.
[0091] If the audio data stream needs to be processed seamlessly, the audio content can be replaced with a specific audio, or noise can be added to the audio content, or the audio content can be directly discarded. The same method of shielding the audio content is not limited here.
[0092] S103: The computing device sends the processed media data stream to the recording server.
[0093] In an embodiment of the present application, the computing device sends the processed media data stream to the recording server, so that the recording server only records the media content corresponding to the processed media data stream and does not record the blocked media content, thereby ensuring the protection of sensitive content involved in the meeting.
[0094] In one possible implementation, the computing device may further send the processed media data stream to the AI server, which may be used to perform AI intelligent analysis on the media content indicated by the processed media data stream.
[0095] In other words, since the media content that needs to be processed seamlessly has been blocked in the processed media data stream, the AI server will not use the blocked media content for AI recognition and analysis, thereby ensuring the protection of sensitive content involved in the meeting.
[0096] In one possible implementation, if it is determined that the media data stream includes a traceless identifier, the computing device can generate status information, which can be used to indicate that the media data stream sent by the first terminal needs to be processed tracelessly. Then, the computing device can send the status information to other terminals so that the other terminals display specified information to the user. The specified information can be used to prompt the user that the media data stream sent by the first terminal needs to be processed tracelessly.
[0097] For example, Figure 7 is a schematic diagram of an interface display of another participating terminal involved in an embodiment of the present application. As shown in Figure 7, if the computing device determines that the media data stream sent by user A through the corresponding terminal includes a traceless identifier, the computing device generates status information and sends the status information to the other participating user terminals, that is, the terminals corresponding to users B and C. The conference interface displayed by the terminal corresponding to user B can display the specified information 31 at the position corresponding to user A in the user list participating in the conference, so that user B can determine that user A has applied for traceless processing through the displayed specified information 31.
[0098] In summary, a user terminal that requires the traceless function to be enabled adds a traceless identifier to the media data stream and sends the media data stream with the traceless identifier to a computing device in the audio and video conferencing system. The computing device then identifies the traceless identifier, determines the media data streams in the audio and video conference that require traceless processing, and sends the media data streams to the recording server after performing traceless processing. This allows the recording server to record the media content indicated by the media data stream with the traceless identifier blocked. This ensures controllable recording permissions for the media data streams sent by participating user terminals in the audio and video conference, prevents the leakage of private media content in the audio and video conference, and further expands the privacy protection function of the audio and video conference.
[0099] Taking the embodiment of the present application as an example of an audio and video conferencing system based on the SVC media mode, FIG8 shows a flow chart of a media data stream processing method provided by an exemplary embodiment of the present application. The media data stream processing method can be executed by the audio and video conferencing system, which can be the audio and video conferencing system shown in FIG3 or FIG4. The media data stream processing method includes the following steps:
[0100] S201: A first terminal sends a media data stream including a traceless identifier to a media forwarding server.
[0101] In one possible implementation, the client running the audio and video conference on the first terminal is the first client. The first client can set a traceless application control in the conference interface of the audio and video conference. The user triggers the traceless application control, that is, the first terminal receives the triggering operation of the traceless application control through the first client. The first terminal can add a traceless identifier in the corresponding media data stream through the first client.
[0102] Among them, the traceless application control can be used to provide users with an application for traceless processing of media data streams. The media data stream can be the data transmitted by the first terminal to other participating terminals in an audio and video conference. The other participating terminals can support the synchronous display of corresponding media content according to the received media data stream. Traceless processing can be used to shield the media content corresponding to part or all of the media data in the media data stream.
[0103] In one possible implementation, the conference interface of the audio and video conference also displays selection controls corresponding to different types of media content. In response to receiving a trigger operation for the traceless application control and receiving a selection operation for the selection control corresponding to the first type of media content, a traceless identifier is added to the header field of each data message of the media data stream corresponding to the first type of media content in the media data stream.
[0104] For example, the user can use the incognito application control displayed on the first terminal and the selection controls corresponding to different types of media content to select their own audio, video, shared and other media content in the meeting to be incognito mode. When the terminal sends the corresponding media data stream, it can carry an incognito flag in the packet header or in the header field of the data message, such as incogito=1.
[0105] For example, Figure 9 is a display diagram of a conference interface involved in an embodiment of the present application. As shown in Figure 9, an incognito mode setting control 40 can be displayed in the conference interface. By triggering the incognito mode setting control 40, the user can display a pop-up window or page for setting the incognito mode in the conference interface. As shown in Figure 9, a pop-up window for setting the incognito mode is displayed in the conference interface, and the pop-up window includes an incognito application control 41. If the user triggers the incognito application control 41, it can be determined that user A's own media content in the conference includes media content that enables incognito mode. The pop-up window for setting the incognito mode in the conference interface can also include selection controls 42 corresponding to different types of media content, such as selection controls corresponding to audio, selection controls corresponding to video, and selection controls corresponding to sharing. If the user selects the selection control corresponding to sharing, the first terminal can add an incognito identifier to the data stream corresponding to the shared content.
[0106] S202: The media forwarding server sends a media data stream carrying a traceless identifier to the recording server.
[0107] S203: The recording server performs traceless processing on the media content carrying the traceless identifier.
[0108] S204: The media forwarding server sends a media data stream carrying a traceless identifier to the media gateway.
[0109] S205: The media gateway performs traceless processing on the media content carrying the traceless identifier.
[0110] S206: The media gateway sends the seamlessly processed media content to the AI server.
[0111] Among them, after receiving the media content after seamless processing, the AI server performs AI intelligent analysis on the processed media content to obtain the analysis results. The AI server sends the analysis results to the media gateway. After the media gateway receives the results of the AI intelligent analysis of the processed media data stream returned by the AI server, it sends them to the media forwarding server. The media forwarding server can send the results to other participating terminals so that the results can be displayed in the audio and video conference.
[0112] S207: The media forwarding server sends the media data stream carrying the traceless identifier to other participating terminals.
[0113] The order in which the above S202, S204 and S207 are executed is not limited.
[0114] Taking the embodiment of the present application as an example of an audio and video conferencing system based on the AVC media mode, FIG10 shows a flow chart of a media data stream processing method provided by an exemplary embodiment of the present application. The media data stream processing method can be executed by the audio and video conferencing system, which can be the audio and video conferencing system shown in FIG3 or FIG4. The media data stream processing method includes the following steps:
[0115] S301: A first terminal applies to a conference management server for entering a stealth state.
[0116] The first terminal may request the conference management server to enable audio, video, shared and other media contents to enter a traceless state.
[0117] S302: The conference management server sends the status information of the first terminal to the media server.
[0118] S303: The media server sends the status information of the first terminal to other participating terminals.
[0119] The first terminal may display designated information of the first terminal according to the state information of the first terminal, so as to prompt the user that the first terminal has entered the invisible state.
[0120] S304: The first terminal sends a media data stream to the media server.
[0121] S305: The media server adds a traceless identifier to the media data stream and performs traceless processing.
[0122] S306: The media server sends the media content after seamless processing to the recording server.
[0123] S307: The media server sends the seamlessly processed media content to the AI server.
[0124] Among them, after receiving the media content after seamless processing, the AI server performs AI intelligent analysis on the processed media content to obtain the analysis results. The AI server sends the analysis results to the media server. After the media server receives the results of the AI intelligent analysis of the processed media data stream returned by the AI server, it can send the results to other participating terminals so that the results can be displayed in the audio and video conference.
[0125] S308: The media server sends the media data stream to other participating terminals.
[0126] Since the other participating terminals have determined that the first terminal has entered the invisible state after S303, they cannot perform operations such as recording, screenshot, and sharing on the corresponding media content after receiving the media data stream sent by the media server.
[0127] S309: The conference host sends a request for entering a stealth state to the conference management server via the host terminal.
[0128] S310: The conference management server may determine that the conference enters the invisible state and notify the media server.
[0129] S311: The media server adds a traceless identifier to each media data stream of the conference and performs traceless processing.
[0130] S312: The media server notifies each participating terminal that the conference is in incognito mode.
[0131] In summary, a user terminal that requires the traceless function to be enabled adds a traceless identifier to the media data stream and sends the media data stream with the traceless identifier to a computing device in the audio and video conferencing system. The computing device then identifies the traceless identifier, determines the media data streams in the audio and video conference that require traceless processing, and sends the media data streams to the recording server after performing traceless processing. This allows the recording server to record the media content indicated by the media data stream with the traceless identifier blocked. This ensures controllable recording permissions for the media data streams sent by participating user terminals in the audio and video conference, prevents the leakage of private media content in the audio and video conference, and further expands the privacy protection function of the audio and video conference.
[0132] The above mainly introduces the scheme of the embodiment of the present application from the perspective of method. It can be understood that in order to realize the above functions, the media data stream processing device includes at least one of the hardware structure and software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0133] The embodiment of the present application can divide the media data stream processing device into functional units according to the above method example. For example, each functional unit can be divided according to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units. It should be noted that the division of units in the embodiment of the present application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.
[0134] For example, FIG11 shows a schematic diagram of the structure of a media data stream processing apparatus 500 provided by an exemplary embodiment of the present application. The media data stream processing apparatus 500 is applied to a computing device, or the media data stream processing apparatus 500 can be a computing device. The media data stream processing apparatus 500 includes:
[0135] An acquisition module 510 is configured to acquire a media data stream, where the media data stream is data transmitted by the first terminal to other participating terminals in an audio or video conference, and the other participating terminals are configured to display media content corresponding to the media data stream;
[0136] A processing module 520 is configured to, if the media data stream includes a traceless identifier, perform traceless processing on the media data stream according to the traceless identifier to obtain a processed media data stream, wherein the traceless processing is configured to mask media content corresponding to part or all of the media data in the media data stream;
[0137] The sending module 530 is configured to send the processed media data stream to a recording server, so that the recording server records the media content indicated by the processed media data stream.
[0138] For example, in conjunction with FIG3 , the acquisition module 510 may be used to execute S101 as shown in FIG5 , the processing module 520 may be used to execute S102 as shown in FIG5 , and the sending module 530 may be used to execute S103 as shown in FIG5 .
[0139] In one possible implementation, the media data stream includes multiple data packets, and the processing module 520 is further used to determine a target data packet from the multiple data packets, where the target data packet is the data packet containing the traceless identifier in the packet header field; the target data packet is processed tracelessly to obtain the processed media data stream, where the processed media data stream includes the target data packet that has been processed tracelessly and other data packets among the multiple data packets except the target data packet.
[0140] In one possible implementation, the sending module 530 is further used to send the processed media data stream to an artificial intelligence (AI) server, and the AI server is used to perform AI intelligent analysis on the media content indicated by the processed media data stream; and receive the result of the AI intelligent analysis of the processed media data stream returned by the AI server, and the result is used to be displayed in the audio and video conference.
[0141] In a possible implementation, the media data stream includes one or more of the following: at least one of an audio data stream, a video data stream, a shared media stream, or a message stream.
[0142] In one possible implementation, if the media data stream is a media stream based on a scalable video coding SVC mode, the computing device includes a media stream router SFU; or, if the media data stream is a media stream based on an advanced video coding AVC mode, the computing device includes a media server.
[0143] In a possible implementation, the apparatus further includes:
[0144] a generating module configured to generate status information if it is determined that the media data stream includes a traceless identifier, wherein the status information is used to indicate that the media data stream sent by the first terminal needs to be processed tracelessly;
[0145] The sending module 530 is further configured to send the status information to the other terminal, so that the other terminal displays specified information to the user, where the specified information is used to prompt the user that the media data stream sent by the first terminal needs to be seamlessly processed.
[0146] In one possible implementation, if the status information is also used to indicate the type of media data stream in the media data stream sent by the first terminal that needs to be seamlessly processed, the specified information is also used to prompt the user of the type of media data stream in the media data stream sent by the first terminal that needs to be seamlessly processed.
[0147] For example, FIG12 shows a schematic diagram of the structure of a media data stream processing device 600 provided in an exemplary embodiment of the present application. The media data stream processing device 600 is applied to a first client, or the media data stream processing device 600 can be the first client. The first client is a client for an audio or video conference running on a first terminal. The media data stream processing device 600 includes:
[0148] A setting module 610 is configured to set a traceless application control in the conference interface of the audio and video conference, wherein the traceless application control is configured to provide a user with an application for traceless processing of a media data stream, wherein the media data stream is data transmitted by the first terminal to other participating terminals in the audio and video conference, and the other participating terminals are configured to display media content corresponding to the media data stream, and the traceless processing is configured to shield the media content corresponding to part or all of the media data in the media data stream;
[0149] An adding module 620 is configured to add a traceless identifier to the media data stream in response to receiving a triggering operation on the traceless application control;
[0150] The sending module 630 is used to send the media data stream with the traceless identifier added to the computing device, and the computing device is used to perform traceless processing on the media data stream according to the traceless identifier to obtain the processed media data stream, and send the processed media data stream to the recording server so that the recording server records the media content indicated by the processed media data stream.
[0151] In a possible implementation, the conference interface of the audio and video conference further displays selection controls corresponding to different types of media content;
[0152] The adding module 620 is also used to add the seamless identifier in the header field of each data message of the media data stream corresponding to the first type of media content in the media data stream in response to receiving a trigger operation for the seamless application control and receiving a selection operation for the selection control corresponding to the first type of media content.
[0153] For the detailed description of the above optional methods, please refer to the above method embodiments, which will not be repeated here. In addition, the explanation and beneficial effects of any of the above media data stream processing devices can be referred to the above corresponding method embodiments, which will not be repeated here.
[0154] The acquisition module 510, processing module 520, and sending module 530 can all be implemented in software or hardware. For example, the implementation of acquisition module 510 will be described below using acquisition module 510 as an example. Similarly, the implementation of processing module 520 and sending module 530 can refer to the implementation of acquisition module 510.
[0155] As an example of a software functional unit, the acquisition module 510 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the acquisition module 510 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.
[0156] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0157] As an example of a hardware functional unit, the acquisition module 510 may include at least one computing device, such as a server. Alternatively, the acquisition module 510 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0158] The multiple computing devices included in acquisition module 510 can be distributed in the same region or in different regions. The multiple computing devices included in acquisition module 510 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in acquisition module 510 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0159] It should be noted that, in other embodiments, the acquisition module 510 can be used to execute any step in the media data stream processing method, the processing module 520 can be used to execute any step in the media data stream processing method, and the sending module 530 can be used to execute any step in the media data stream processing method. The steps that the acquisition module 510, the processing module 520, and the sending module 530 are responsible for implementing can be specified as needed. By having the acquisition module 510, the processing module 520, and the sending module 530 respectively implement different steps in the media data stream processing method, the full functionality of the media data stream processing device is achieved. This application also provides a computing device 100. As shown in Figure 13, the computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other via the bus 102. The computing device 100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.
[0160] Bus 102 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG13 shows only one line, but this does not imply a single bus or type of bus. Bus 104 may include a path for transmitting information between various components of computing device 100 (e.g., memory 106, processor 104, and communication interface 108).
[0161] The processor 104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0162] The memory 106 may include volatile memory, such as random access memory (RAM). The processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0163] The memory 106 stores executable program code, and the processor 104 executes the executable program code to respectively implement the functions of the aforementioned acquisition module, processing module 520, and sending module 530, thereby implementing the media data stream processing method. In other words, the memory 106 stores instructions for executing the media data stream processing method.
[0164] Alternatively, the memory 106 stores executable codes, and the processor 104 executes the executable codes to respectively implement the functions of the aforementioned media data stream processing apparatus, thereby implementing the media data stream processing method. That is, the memory 106 stores instructions for executing the media data stream processing method.
[0165] The communication interface 103 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or a communication network.
[0166] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0167] As shown in Figure 14, the computing device cluster includes at least one computing device 100. The memory 106 in one or more computing devices 100 in the computing device cluster may store the same instructions for executing the media data stream processing method.
[0168] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the media data stream processing method. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for executing the media data stream processing method.
[0169] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions, each for executing a portion of the functions of the media data stream processing apparatus. In other words, the instructions stored in the memory 106 in different computing devices 100 can implement the functions of one or more of the acquisition module 510, processing module 520, and sending module 530.
[0170] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network or a local area network, etc. FIG15 shows a possible implementation. As shown in FIG15 , two computing devices 100A and 100B are connected via a network. Specifically, the connection to the network is made via a communication interface in each computing device. In this type of possible implementation, the memory 106 in the computing device 100A stores instructions for executing the functions of the acquisition module 510. At the same time, the memory 106 in the computing device 100B stores instructions for executing the functions of the processing module 520 and the sending module 530.
[0171] The connection method between the computing device clusters shown in Figure 15 can be considered to be that the media data stream processing method provided in this application requires a large amount of storage data and calculation data, so it is considered to entrust the functions implemented by the processing module 520 and the sending module 530 to the computing device 100B for execution.
[0172] It should be understood that the functions of the computing device 100A shown in FIG15 may also be completed by multiple computing devices 100. Similarly, the functions of the computing device 100B may also be completed by multiple computing devices 100.
[0173] The present application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection method of the computing device cluster described in Figures 14 and 15. However, the memory 106 in one or more computing devices 100 in this computing device cluster can store the same instructions for executing the media data stream processing method.
[0174] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the media data stream processing method. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for executing the media data stream processing method.
[0175] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions for executing partial functions of the data processing system. In other words, the instructions stored in the memory 106 in different computing devices 100 can implement the functions of one or more devices in the media data stream processing device.
[0176] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the media data stream processing method.
[0177] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to execute a media data stream processing method, or instructs a computing device to execute a media data stream processing method.
[0178] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for processing media data streams, characterized in that, Applied to a computing device, the method includes: Obtain a media data stream, which is data transmitted by a first terminal to other participating terminals during an audio-video conference, and the other participating terminals are used to display media content corresponding to the media data stream; If the media data stream includes a markless identifier, perform markless processing on the media data stream according to the markless identifier to obtain a processed media data stream, and the markless processing is used to mask part or all of the media content corresponding to the media data in the media data stream; Send the processed media data stream to a recording server so that the recording server records the media content indicated by the processed media data stream.
2. The method according to claim 1, characterized in that The media data stream includes a plurality of data packets, and performing markless processing on the media data stream according to the markless identifier to obtain a processed media data stream includes: Determine a target data packet from the plurality of data packets, where the target data packet is the data packet that includes the markless identifier in the packet header field; Perform markless processing on the target data packet to obtain the processed media data stream, and the processed media data stream includes the target data packet that has undergone markless processing and other data packets in the plurality of data packets except the target data packet.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Send the processed media data stream to an artificial intelligence (AI) server, and the AI server is used to perform AI intelligent analysis on the media content indicated by the processed media data stream; Receive the result of AI intelligent analysis on the processed media data stream returned by the AI server, and the result is used to be displayed during the audio-video conference.
4. The method according to any one of claims 1 to 3, characterized in that, The media data stream includes one or more of the following: audio data stream, video data stream, shared media stream, message stream.
5. The method according to any one of claims 1 to 4, characterized in that, If the media data stream is a media stream based on the scalable video coding (SVC) mode, the computing device includes a media stream router (SFU); or, if the media data stream is a media stream based on the advanced video coding (AVC) mode, the computing device includes a media server.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: If it is determined that the media data stream includes a markless identifier, generate status information, and the status information is used to indicate that the media data stream sent by the first terminal needs to be subjected to markless processing; Send the status information to the other terminals so that the other terminals display specified information to the user, and the specified information is used to prompt the user that the media data stream sent by the first terminal needs to be subjected to markless processing.
7. The method according to claim 6, wherein If the status information is further used to indicate the type of media data stream in the media data stream sent by the first terminal that needs to be subjected to markless processing, the specified information is further used to prompt the user of the type of media data stream in the media data stream sent by the first terminal that needs to be subjected to markless processing.
8. A method for processing media data streams, characterized in that, Applied to a first client, the first client is a client of an audio-video conference running on a first terminal, and the method includes: Set a traceless application control in the conference interface of the audio-video conference. The traceless application control is used to provide a user with an application to perform traceless processing on the media data stream. The media data stream is data transmitted by the first terminal to other participating terminals during the audio-video conference, and the other participating terminals are used to display the media content corresponding to the media data stream. The traceless processing is used to mask the media content corresponding to some or all of the media data in the media data stream; In response to receiving a trigger operation on the traceless application control, add a traceless identifier to the media data stream; Send the media data stream with the traceless identifier added to a computing device. The computing device is used to perform traceless processing on the media data stream according to the traceless identifier to obtain a processed media data stream, and send the processed media data stream to a recording server, so that the recording server records the media content indicated by the processed media data stream.
9. The method according to claim 8, wherein Selection controls corresponding to different types of media content are also displayed in the conference interface of the audio-video conference; The adding a traceless identifier to the media data stream in response to receiving a trigger operation on the traceless application control includes: In response to receiving a trigger operation on the traceless application control and a selection operation on the selection control corresponding to the first type of media content, add the traceless identifier to the header fields of each data packet of the media data stream corresponding to the first type of media content in the media data stream.
10. A media data stream processing device, characterized in that, Applied to a computing device, the apparatus includes: An acquisition module, configured to acquire a media data stream. The media data stream is data transmitted by a first terminal to other participating terminals during an audio-video conference, and the other participating terminals are used to display the media content corresponding to the media data stream; A processing module, configured to, if the media data stream includes a traceless identifier, perform traceless processing on the media data stream according to the traceless identifier to obtain a processed media data stream. The traceless processing is used to mask the media content corresponding to some or all of the media data in the media data stream; A sending module, configured to send the processed media data stream to a recording server, so that the recording server records the media content indicated by the processed media data stream.
11. The device according to claim 10, characterized in that, The media data stream includes a plurality of data packets. The processing module is further configured to determine a target data packet from the plurality of data packets. The target data packet is the data packet that includes the traceless identifier in the packet header field; Perform traceless processing on the target data packet to obtain the processed media data stream. The processed media data stream includes the target data packet that has undergone traceless processing and the other data packets in the plurality of data packets except the target data packet.
12. The device according to claim 10 or 11, characterized in that, The sending module is further configured to send the processed media data stream to an artificial intelligence (AI) server, where the AI server is configured to perform AI intelligent analysis on the media content indicated by the processed media data stream; and receive the result of the AI intelligent analysis on the processed media data stream returned by the AI server, where the result is used for display in the audio and video conference.
13. The device according to any one of claims 10 to 12, characterized in that, The media data stream includes one or more of the following: audio data stream, video data stream, shared media stream, and message stream.
14. The device according to any one of claims 10 to 13, characterized in that, If the media data stream is a media stream based on the scalable video coding (SVC) mode, the computing device includes a media stream router (SFU); or, if the media data stream is a media stream based on the advanced video coding (AVC) mode, the computing device includes a media server.
15. The device according to any one of claims 10 to 14, characterized in that The device further includes: A generating module, configured to generate status information if it is determined that the media data stream includes a trace-free identifier, where the status information is used to indicate that the media data stream sent by the first terminal needs to be processed without a trace; The sending module is further configured to send the status information to the other terminals, so that the other terminals display specified information to the user, where the specified information is used to prompt the user that the media data stream sent by the first terminal needs to be processed without a trace.
16. The device according to claim 15, characterized in that, If the status information is further used to indicate the type of the media data stream that needs to be processed without a trace in the media data stream sent by the first terminal, the specified information is further used to prompt the user of the type of the media data stream that needs to be processed without a trace in the media data stream sent by the first terminal.
17. A media data stream processing device, characterized in that, Applied to a first client, where the first client is a client of an audio and video conference running on a first terminal, the device includes: A setting module, configured to set a trace-free application control in the conference interface of the audio and video conference, where the trace-free application control is used to provide the user with an application to process the media data stream without a trace, where the media data stream is data transmitted by the first terminal to other participating terminals in the audio and video conference, and the other participating terminals are used to display the media content corresponding to the media data stream, and the trace-free processing is used to mask part or all of the media content corresponding to the media data in the media data stream; An adding module, configured to add a trace-free identifier to the media data stream in response to receiving a trigger operation on the trace-free application control; A sending module, configured to send the media data stream with the trace-free identifier added to a computing device, where the computing device is configured to perform trace-free processing on the media data stream according to the trace-free identifier to obtain a processed media data stream, and send the processed media data stream to a recording server, so that the recording server records the media content indicated by the processed media data stream.
18. The device according to claim 17, wherein Selection controls corresponding to different types of media content are further displayed in the conference interface of the audio and video conference; The adding module is further configured to, in response to receiving a triggering operation on the traceless application control and a selection operation on the selection control corresponding to the media content of the first type, add the traceless identifier to a header field of each data packet of the media data stream corresponding to the media content of the first type in the media data stream.
19. A cluster of computing devices, characterized in that, It includes at least one computing device, and each computing device includes a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the media data stream processing method according to any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that, It includes computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the media data stream processing method according to any one of claims 1 to 9.
21. A computer program product, characterized in that, The computer program product includes instructions, and when the instructions are run by a computing device cluster, the computing device cluster is caused to execute the media data stream processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method and device for processing recorded content of audio and video conference
CN103475835A
Video conference content protection method, device, equipment and system
CN110012260A
Articulated naturality web conference recording method and system
CN110536100A
Video conference cloud recording method and device, electronic equipment and storage medium
CN117294805A
Systems and Methods for Recording and Storing Media Content
US20180241785A1