A multi-mode fusion gateway based on an integrated video platform
By utilizing deep learning and graph neural network technologies in the multi-mode fusion gateway, the problems of audio codec incompatibility and access control in terminal devices of the integrated video platform were solved, realizing audio compatibility mapping and flexible access control, thereby improving platform performance and security.
Patent Information
- Application Number
- CN202510139714.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-02-08
AI Technical Summary
The incompatibility of audio codec standards among different terminal devices in existing integrated video conferencing platforms and the difficulty in flexibly adjusting user role and permission management in complex conferencing scenarios lead to a decline in audio quality and an increase in security risks.
A multi-mode fusion gateway is adopted, which uses a deep learning model to identify the terminal audio codec standard, generates an adaptation matrix through a dynamic codec library to achieve audio compatibility mapping, and combines a graph neural network to build a cross-platform permission relationship graph to adjust user permissions in real time.
It improves audio transmission quality and reliability, simplifies access control processes, reduces security risks, and enhances the adaptability of the integrated video conferencing platform in complex environments and improves user experience.
Smart Images

Figure CN119966957B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication network technology, and more specifically, to a multi-mode converged gateway based on an integrated video platform. Background Technology
[0002] With the acceleration of globalization and the increasing demand for remote work, video conferencing systems have become an indispensable collaboration tool for modern enterprises and organizations. As the core of the next generation of video conferencing systems, integrated video conferencing platforms combine audio and video communication, instant messaging, file sharing, and other functions, providing users with a comprehensive remote collaboration experience. However, with the diversification of application scenarios and the increasing complexity of user needs, integrated video conferencing platforms face a series of challenges, including cross-platform compatibility, multimodal data processing, and dynamic access control, urgently requiring more advanced and intelligent technological solutions.
[0003] Currently, there are various integrated video conferencing platform solutions on the market, such as Zoom, Microsoft Teams, and Cisco Webex. These platforms perform well in basic functions such as audio and video transmission, screen sharing, and instant messaging, and are gradually introducing artificial intelligence technologies to enhance user experience; for example, some platforms use speech recognition technology to achieve real-time captioning, or use computer vision technology for virtual background replacement; at the same time, in order to adapt to different network environments, these platforms also adopt technologies such as adaptive bitrate encoding to optimize audio and video quality; in terms of security, most platforms provide role-based access control (RBAC) mechanisms to manage user permissions.
[0004] However, existing technologies still have some significant shortcomings. First, within the same integrated video conferencing platform, different terminal devices may use different audio codec standards, potentially leading to audio quality degradation or communication failures. This heterogeneous terminal audio codec compatibility issue directly impacts the quality and effectiveness of the meeting. Second, in complex meeting scenarios, existing technologies have failed to effectively implement flexible user role and permission management, potentially resulting in improper permission allocation, increased management difficulty, and security risks. These problems severely restrict the application and promotion of integrated video conferencing platforms in complex and diverse environments.
[0005] There are currently no effective solutions to the problems in the relevant technologies. Summary of the Invention
[0006] To address the problems in related technologies, this invention proposes a multi-mode fusion gateway based on an integrated video conferencing platform. This gateway effectively solves the problems of audio compatibility between different terminals within the platform and flexible user permission management. It also solves the compatibility issues caused by different audio codec standards used by different terminal devices within the video conferencing platform, as well as the difficulty in flexibly managing user role permissions in complex meeting scenarios, which increases management difficulty and security risks.
[0007] Therefore, the specific technical solution adopted by the present invention is as follows:
[0008] According to one aspect of the present invention, a multi-mode fusion gateway based on an integrated video platform is provided, the multi-mode fusion gateway based on the integrated video platform comprising:
[0009] The multi-modal data processing module is used to extract audio data, video data and auxiliary data in real time based on the multi-modal data stream of the integrated video platform using a data separation algorithm. For the extracted audio data, a deep learning model is used to identify the audio codec standards supported by each terminal, and an adaptation matrix is generated through a dynamic codec library to obtain the audio compatibility mapping relationship between terminals. The video data and auxiliary data are also preliminarily analyzed to establish a multi-modal data association index.
[0010] The user identity association module is used to convert audio streams between different codec standards in real time based on the audio compatibility mapping relationship between terminals. It also combines a multi-modal data association index and uses a multi-modal fusion algorithm to analyze the comprehensive characteristics of audio data, video data, and auxiliary data to establish an association model between user identity and multi-modal data stream.
[0011] The dynamic permission management module is used to construct a cross-platform permission relationship graph based on user identity and multi-modal data stream association model, combined with user role permissions of the access platform, and using semantic analysis technology and graph neural network. Through dynamic permission allocation algorithm, it adjusts the audio and video transmission permissions and interaction permissions of users across different platforms in real time.
[0012] Furthermore, the multi-modal data processing module, based on the multi-modal data stream of the integrated video platform, utilizes a data separation algorithm to extract audio data, video data, and auxiliary data in real time. For the extracted audio data, a deep learning model is used to identify the audio codec standards supported by each terminal, and an adaptation matrix is generated through a dynamic codec library to obtain the audio compatibility mapping relationship between terminals. Preliminary analysis is performed on the video data and auxiliary data, and a multi-modal data association index is established, including:
[0013] Based on the multi-modal data stream of the integrated video platform, data separation algorithms are used to obtain separated audio data, video data, and auxiliary data. Metadata analysis technology is then used to establish a preliminary index of the multi-modal data.
[0014] Based on the separated audio data, a deep learning model is used to obtain the audio codec standard information supported by each terminal, and a terminal audio codec capability descriptor is generated through a feature matching algorithm.
[0015] Based on the terminal audio codec capability descriptor, an audio compatibility adaptation matrix is obtained using a dynamic codec library, and an optimal audio compatibility mapping relationship between terminals is established through a matrix optimization algorithm.
[0016] Based on the separated video data and auxiliary data, video frame analysis algorithms and text parsing algorithms are used to obtain video keyframe information and auxiliary data feature information. The preliminary index of multi-modal data is then updated using the video keyframe information and auxiliary data feature information to obtain the multi-modal data association index.
[0017] Furthermore, when the user identity association module converts audio streams between different codec standards in real time based on the audio compatibility mapping relationship between terminals, and combines the multimodal data association index with the multimodal fusion algorithm to analyze the comprehensive characteristics of audio data, video data, and auxiliary data to establish an association model between user identity and multimodal data streams, it includes:
[0018] Based on the optimal audio compatibility mapping relationship between terminals, adaptive audio transcoding is used to generate cross-platform audio streams that are converted in real time.
[0019] Based on the multimodal data association index, a multimodal feature extraction algorithm is used to obtain the comprehensive feature vector of audio data, video data and auxiliary data, and a preliminary user identity feature model is generated through a deep neural network.
[0020] Based on a preliminary user identity feature model and a real-time converted cross-platform audio stream, user identity information is obtained using speaker recognition technology, and a dynamic correlation model between user identity and multi-modal data stream is established through a time-series analysis algorithm.
[0021] Furthermore, the dynamic permission management module, based on user identity and a multi-modal data stream association model, combined with user role permissions on the access platform, utilizes semantic analysis technology and graph neural networks to construct a cross-platform permission relationship graph, and adjusts users' audio and video transmission permissions and interaction permissions across different platforms in real time through a dynamic permission allocation algorithm, including:
[0022] Based on the user identity and multimodal data flow association model, combined with the user role permissions of the access platform, semantic analysis technology is used to obtain the permission feature vector of each platform;
[0023] Based on the permission feature vectors of each platform, a cross-platform permission relationship graph is generated using a graph neural network algorithm;
[0024] Based on a cross-platform permission relationship graph, and using a dynamic permission allocation algorithm, audio and video transmission permissions and interaction permission policies for users across different platforms are generated in real time.
[0025] The beneficial effects of this invention are as follows:
[0026] (1) Through innovative multi-modal data processing technology, the audio codec compatibility problem between different terminal devices in the integrated video platform is effectively solved. The gateway uses a deep learning model to identify the audio codec standards supported by each terminal and generates an adaptation matrix through a dynamic codec library to achieve optimal audio compatibility mapping between terminals. This method not only improves the quality and reliability of audio transmission, but also enhances the adaptability of the integrated video platform in complex environments. At the same time, by using a multi-modal fusion algorithm to analyze the comprehensive characteristics of audio, video and auxiliary data, more accurate user identification and data flow association are achieved, laying the foundation for access control.
[0027] (2) This invention also innovatively introduces a dynamic permission management mechanism based on semantic analysis technology and graph neural network. By constructing a permission relationship graph and combining it with a dynamic permission allocation algorithm, it realizes flexible management and real-time adjustment of user role permissions. This method not only simplifies the permission management process, but also significantly improves security and management efficiency. Especially in complex meeting scenarios, it can dynamically adjust the user's audio and video transmission permissions and interaction permissions according to the real-time situation, effectively preventing the risk of permission abuse and unauthorized access. It provides a comprehensive, intelligent and secure solution for integrated video conferencing platforms, greatly improving platform performance and user experience. Attached Figure Description
[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0029] Figure 1 This is a schematic diagram of a multi-mode fusion gateway based on an integrated video platform according to an embodiment of the present invention. Detailed Implementation
[0030] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention. The components in the drawings are not drawn to scale, and similar component symbols are generally used to represent similar components.
[0031] According to an embodiment of the present invention, a multi-mode fusion gateway based on an integrated video platform is provided.
[0032] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, according to an embodiment of the present invention, a multi-mode fusion gateway based on an integrated video platform is provided, the multi-mode fusion gateway based on the integrated video platform comprising:
[0033] Multi-modal data processing module 1 is used to extract audio data, video data and auxiliary data in real time based on the multi-modal data stream of the integrated video platform using a data separation algorithm. For the extracted audio data, a deep learning model is used to identify the audio codec standards supported by each terminal, and an adaptation matrix is generated through a dynamic codec library to obtain the audio compatibility mapping relationship between terminals. The video data and auxiliary data are also preliminarily analyzed to establish a multi-modal data association index.
[0034] User identity association module 2 is used to convert audio streams between different codec standards in real time according to the audio compatibility mapping relationship between terminals, and combine multi-modal data association index to analyze the comprehensive characteristics of audio data, video data and auxiliary data using multi-modal fusion algorithm to establish an association model between user identity and multi-modal data stream;
[0035] The dynamic permission management module 3 is used to construct a cross-platform permission relationship graph based on user identity and multi-modal data flow association model, combined with user role permissions of the access platform, using semantic analysis technology and graph neural network, and adjusting the user's audio and video transmission permissions and interaction permissions between different platforms in real time through dynamic permission allocation algorithm.
[0036] In one embodiment, the multi-modal data processing module 1 extracts audio data, video data, and auxiliary data in real time from the multi-modal data stream based on the integrated video platform using a data separation algorithm. For the extracted audio data, it uses a deep learning model to identify the audio codec standards supported by each terminal, and generates an adaptation matrix through a dynamic codec library to obtain the audio compatibility mapping relationship between terminals. When performing preliminary analysis on the video data and auxiliary data and establishing a multi-modal data association index, the module includes:
[0037] S11. Based on the multi-mode data stream of the integrated video platform, use the data separation algorithm to obtain the separated audio data, video data and auxiliary data, and establish a preliminary index of multi-mode data through metadata analysis technology;
[0038] Specifically, the real-time data separation algorithm employs a frequency domain separation method based on FFT (Fast Fourier Transform). First, the input multi-mode data stream is segmented into frames, each 20ms in length. Then, an FFT is performed on each frame to transform it into the frequency domain. In the frequency domain, the data is classified according to a preset frequency range (e.g., audio: 20Hz-20kHz, video: 0-4MHz). Finally, an inverse FFT is used to transform the classified frequency domain data back into the time domain, yielding the separated audio, video, and auxiliary data streams.
[0039] Specifically, metadata analysis technology uses regular expression matching to extract information such as timestamps, data types, and source terminal IDs from the packet header. This information is then stored in a relational database (such as SQLite), with the data table formatted as (timestamp, data_type, source_id, packet_id), thus forming a preliminary multi-modal data index.
[0040] S12. Based on the separated audio data, use a deep learning model to obtain the audio codec standard information supported by each terminal, and generate a terminal audio codec capability descriptor through a feature matching algorithm.
[0041] Specifically, the deep learning model employs a pre-trained ResNet-18 convolutional neural network. First, the audio data is converted into a Mel-spectrogram, which serves as the model's input. The last layer of the ResNet-18 model is replaced with a fully connected layer with N output nodes, where N is the number of predefined codec standards (e.g., N = 10, including common standards such as G.711, G.722, and Opus).
[0042] Specifically, the model was trained using the cross-entropy loss function and the Adam optimizer, with a learning rate of 0.001, a batch size of 64, and 100 training epochs. The training dataset contained labeled audio samples from various codec standards.
[0043] Specifically, the feature matching algorithm uses cosine similarity calculation. For each terminal, its audio data is processed through a trained ResNet-18 model to obtain an N-dimensional vector, representing the support probability of each codec standard. This vector is then compared with a predefined standard vector using cosine similarity calculation, and the codec standard with the highest similarity is the one supported by that terminal.
[0044] S13. Based on the terminal audio codec capability descriptor, use the dynamic codec library to obtain the audio compatibility adaptation matrix, and establish the optimal audio compatibility mapping relationship between terminals through the matrix optimization algorithm.
[0045] S14. Based on the separated video data and auxiliary data, use video frame analysis algorithm and text parsing algorithm to obtain video key frame information and auxiliary data feature information, and update the preliminary index of multi-mode data through video key frame information and auxiliary data feature information to obtain multi-mode data association index.
[0046] Specifically, based on the separated video data, a video keyframe extraction algorithm based on convolutional neural networks is employed. First, a pre-trained ResNet-50 model is used to extract features from the video frames. Then, the cosine similarity of features between adjacent frames is calculated; when the similarity is below a preset threshold, the frame is marked as a keyframe. The extracted keyframe information includes the frame number and the corresponding timestamp.
[0047] Specifically, for auxiliary data, the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm is used to extract text features. First, the text data is segmented and stop words are removed. Then, the TF-IDF value of each word is calculated, and the top N words with the highest TF-IDF values are selected as the feature words of the text.
[0048] Specifically, finally, the video keyframe information and auxiliary data feature information are added to the preliminary multi-modal data index built in S11. The updated data table format is: (timestamp, data_type, source_id, packet_id, key_frame_info, text_features). The key_frame_info field stores the keyframe sequence number and timestamp, and the text_features field stores the list of extracted text feature words, thus obtaining a more complete multi-modal data association index.
[0049] In one embodiment, based on the terminal audio codec capability descriptor, an audio compatibility adaptation matrix is obtained using a dynamic codec library, and an optimal audio compatibility mapping relationship between terminals is established through a matrix optimization algorithm, including:
[0050] S131. Based on the terminal audio codec capability descriptor, calculate the audio codec standard quality score and establish an audio compatibility adaptation matrix.
[0051] S132. Based on the audio compatibility adaptation matrix, establish a compatibility weighted graph; each node of the compatibility weighted graph represents an audio codec standard, and each edge of the compatibility weighted graph represents a conversion path between two audio codec standards;
[0052] S133. Using the Floyd algorithm, calculate the shortest path between any two nodes in the compatibility weighted graph to obtain the optimized set of transformation paths.
[0053] S134. Store the optimized conversion path set in a hash table to obtain the optimal audio compatibility mapping relationship between terminals;
[0054] The keys in the hash table include source format identifiers and destination format identifiers;
[0055] The source format identifier is a unique identifier of the audio codec standard from which the conversion begins, and the target format identifier is a unique identifier of the audio codec standard from which the conversion is targeted.
[0056] The hash table contains the optimal conversion path from the starting audio codec standard to the target audio codec standard.
[0057] In one embodiment, the expression for calculating the audio codec standard quality score is:
[0058]
[0059] In the formula, Q(i,j) is the conversion quality score from audio codec standard i to audio codec standard j; B i and B j L represents the bit rate of audio codec standard i and audio codec standard j, respectively; i and L j C represents the delays of audio codec standard i and audio codec standard j, respectively; ij T represents the compatibility coefficient between audio codec standard i and audio codec standard j; ij The average processing time required to convert audio codec standard i to audio codec standard j; T max α is the maximum processing time required to convert between any two audio codec standards; β is the bit rate impact factor; γ is the latency sensitivity coefficient; δ is the compatibility weight; and δ is the conversion efficiency coefficient.
[0060] In one embodiment, when the user identity association module 2 converts audio streams between different codec standards in real time according to the audio compatibility mapping relationship between terminals, and combines a multimodal data association index to analyze the comprehensive characteristics of audio data, video data, and auxiliary data using a multimodal fusion algorithm to establish an association model between user identity and multimodal data streams, the module includes:
[0061] S21. Based on the optimal audio compatibility mapping relationship between terminals, adaptive audio transcoding is used to generate a cross-platform audio stream with real-time conversion.
[0062] Specifically, based on the optimal audio compatibility mapping relationship between terminals, the source audio codec standard and the target audio codec standard are first determined. Then, the real-time target bitrate is calculated according to the adaptive audio transcoding expression. Next, a suitable audio codec (such as Opus, AAC, or AMR-WB) is selected, and the encoding parameters are dynamically adjusted according to the calculated target bitrate. Finally, the transcoded audio stream is packaged into RTP packets, and appropriate timestamps and sequence numbers are added to generate a real-time converted cross-platform audio stream.
[0063] S22. Based on the multimodal data association index, use the multimodal feature extraction algorithm to obtain the comprehensive feature vector of audio data, video data and auxiliary data, and generate a preliminary user identity feature model through a deep neural network.
[0064] S23. Based on the preliminary user identity feature model and the real-time converted cross-platform audio stream, user identity information is obtained by using speaker recognition technology, and a dynamic correlation model between user identity and multi-modal data stream is established through time series analysis algorithm.
[0065] In one embodiment, the expression for adaptive audio transcoding is:
[0066] R(t)=R base ×(1-ε×PLR(t))×(1+η×RTT(t));
[0067] In the formula, R(t) is the target bit rate at time t; Rbase is the base bit rate; ε is the adjustment coefficient; PLR(t) is the packet loss rate at time t; η is the delay compensation coefficient; and RTT(t) is the round-trip delay at time t.
[0068] In one embodiment, based on a multimodal data association index, a multimodal feature extraction algorithm is used to obtain a comprehensive feature vector of audio data, video data, and auxiliary data, and a preliminary user identity feature model is generated through a deep neural network, including:
[0069] S221. Based on the audio data in the multi-mode data association index, the audio feature vector is obtained through the Mel frequency cepstral coefficient algorithm;
[0070] S222. Based on the video data in the multi-modal data association index, the local binary pattern histogram face recognition algorithm is used to obtain the face feature descriptor, and the deep face feature vector is obtained through a pre-trained convolutional neural network.
[0071] S223. Based on the auxiliary data in the multi-modal data association index, the feature vector of the text data is obtained through the word frequency inverse document frequency algorithm. Combined with the audio feature vector and the deep face feature vector, a preliminary user identity feature model is generated using a deep neural network.
[0072] In one embodiment, based on a preliminary user identity feature model and a real-time converted cross-platform audio stream, user identity information is obtained using speaker recognition technology, and a dynamic correlation model between user identity and multi-modal data stream is established through a time-series analysis algorithm, including:
[0073] S231. Based on real-time conversion of cross-platform audio streams, feature extraction technology is used to obtain speaker feature vectors;
[0074] S232. Based on the speaker feature vector and the preliminary user identity feature model, obtain user identity information using the probabilistic linear discriminant analysis algorithm;
[0075] S233. Based on user identity information and multi-modal data association index, a dynamic association model between user identity and multi-modal data stream is generated using the sliding window time series analysis method.
[0076] Specifically, first, a fixed-size time window (e.g., 30 seconds) is defined. Within each time window, the system collects user behavioral characteristics, such as speech duration, video start time, and number of text messages. Then, a Long Short-Term Memory (LSTM) network is used to model these temporal features, obtaining a dynamic representation of user behavior. Finally, user identity information is combined with this dynamic representation to construct a graph structure, where nodes represent users and edges represent interactions between users. This graph structure is the dynamic association model between user identity and multimodal data flow.
[0077] In one embodiment, the dynamic permission management module 3, based on a user identity and multi-modal data stream association model, combined with user role permissions of the access platform, utilizes semantic analysis technology and graph neural networks to construct a cross-platform permission relationship graph, and adjusts users' audio and video transmission permissions and interaction permissions across different platforms in real time through a dynamic permission allocation algorithm, including:
[0078] S31. Based on the user identity and multi-modal data flow association model, combined with the user role permissions of the access platform, semantic analysis technology is used to obtain the permission feature vector of each platform.
[0079] Specifically, the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm is used to process user role permission descriptions. First, the permission description text is segmented and stop words are removed. Then, the TF-IDF value of each word is calculated. Finally, the top N words with the highest TF-IDF values are selected, and their TF-IDF values are combined into an N-dimensional vector, which serves as the permission feature vector for the platform.
[0080] S32. Based on the permission feature vectors of each platform, a cross-platform permission relationship graph is generated using a graph neural network algorithm;
[0081] Specifically, a Graph Convolutional Network (GCN) is used to construct a cross-platform permission relationship graph. Each platform is considered a node in the graph, and permission feature vectors are used as node features. The connection relationships and edge weights between nodes are determined by calculating the cosine similarity between permission feature vectors. Two layers of GCN are used for message passing and feature aggregation to generate the final cross-platform permission relationship graph.
[0082] S33. Based on the cross-platform permission relationship graph, and using a dynamic permission allocation algorithm, generate in real time the audio and video transmission permissions and interaction permission policies of users across different platforms.
[0083] Specifically, the dynamic permission allocation algorithm is implemented based on rule matching and score calculation methods. First, features of the two platform nodes are extracted according to their positions in the permission relationship graph. Then, the Euclidean distance between these two feature vectors is calculated as a reference score for permission adjustment. Finally, based on preset score thresholds and adjustment rules, the specific permission settings for the user on the target platform are determined, including the enabled or disabled status of various permissions such as audio, video, and text interaction.
[0084] In one embodiment, the user role permissions for accessing the platform include host permissions, ordinary participant permissions, listen-only permissions, guest permissions, and administrator permissions.
[0085] The permission feature vectors of each platform include audio permissions, video permissions, text interaction permissions, meeting control permissions, and data access permissions.
[0086] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A multi-mode fusion gateway based on an integrated video conferencing platform, characterized in that, This multi-mode convergence gateway based on the integrated video conferencing platform includes: The multi-modal data processing module is used to extract audio data, video data and auxiliary data in real time based on the multi-modal data stream of the integrated video platform using a data separation algorithm. For the extracted audio data, a deep learning model is used to identify the audio codec standards supported by each terminal, and an adaptation matrix is generated through a dynamic codec library to obtain the audio compatibility mapping relationship between terminals. The video data and auxiliary data are also preliminarily analyzed to establish a multi-modal data association index. The user identity association module is used to convert audio streams between different codec standards in real time based on the audio compatibility mapping relationship between terminals. It also combines a multi-modal data association index and uses a multi-modal fusion algorithm to analyze the comprehensive characteristics of audio data, video data, and auxiliary data to establish an association model between user identity and multi-modal data stream. The dynamic permission management module is used to construct a cross-platform permission relationship graph based on user identity and multi-modal data stream association model, combined with user role permissions of the access platform, and using semantic analysis technology and graph neural network. Through dynamic permission allocation algorithm, it adjusts the audio and video transmission permissions and interaction permissions of users across different platforms in real time.
2. The multi-mode fusion gateway based on an integrated video platform according to claim 1, characterized in that, The multi-modal data processing module, based on the multi-modal data stream of the integrated video platform, uses a data separation algorithm to extract audio data, video data, and auxiliary data in real time. For the extracted audio data, a deep learning model is used to identify the audio codec standards supported by each terminal, and an adaptation matrix is generated through a dynamic codec library to obtain the audio compatibility mapping relationship between terminals. When performing preliminary analysis on the video data and auxiliary data and establishing a multi-modal data association index, the module includes: Based on the multi-modal data stream of the integrated video platform, data separation algorithms are used to obtain separated audio data, video data, and auxiliary data. Metadata analysis technology is then used to establish a preliminary index of the multi-modal data. Based on the separated audio data, a deep learning model is used to obtain the audio codec standard information supported by each terminal, and a terminal audio codec capability descriptor is generated through a feature matching algorithm. Based on the terminal audio codec capability descriptor, an audio compatibility adaptation matrix is obtained using a dynamic codec library, and an optimal audio compatibility mapping relationship between terminals is established through a matrix optimization algorithm. Based on the separated video data and auxiliary data, video frame analysis algorithms and text parsing algorithms are used to obtain video keyframe information and auxiliary data feature information. The preliminary index of multi-modal data is then updated using the video keyframe information and auxiliary data feature information to obtain the multi-modal data association index.
3. A multi-mode fusion gateway based on an integrated video platform according to claim 2, characterized in that, The step of obtaining an audio compatibility adaptation matrix based on the terminal's audio codec capability descriptor using a dynamic codec library, and establishing the optimal audio compatibility mapping relationship between terminals through a matrix optimization algorithm includes: Based on the terminal audio codec capability descriptor, calculate the audio codec standard quality score and establish an audio compatibility adaptation matrix. A compatibility weighted graph is established based on the audio compatibility adaptation matrix; each node of the compatibility weighted graph represents an audio codec standard, and each edge of the compatibility weighted graph represents a conversion path between two audio codec standards. Using the Floyd algorithm, the shortest path between any two nodes in the compatibility weighted graph is calculated to obtain the optimized set of transformation paths; The optimized set of conversion paths is stored in a hash table to obtain the optimal audio compatibility mapping relationship between terminals; The keys in the hash table include a source format identifier and a target format identifier; The source format identifier is a unique identifier of the audio codec standard from which the conversion begins, and the target format identifier is a unique identifier of the audio codec standard from which the conversion is targeted. The values in the hash table are the optimal conversion paths from the starting audio codec standard to the target audio codec standard.
4. A multi-mode fusion gateway based on an integrated video platform according to claim 3, characterized in that, The expression for calculating the audio codec standard quality score is as follows: In the formula, Q(i,j) is the conversion quality score from audio codec standard i to audio codec standard j; B i and B j L represents the bit rate of audio codec standard i and audio codec standard j, respectively; i and L j C represents the delays of audio codec standard i and audio codec standard j, respectively; ij T represents the compatibility coefficient between audio codec standard i and audio codec standard j; ij The average processing time required to convert audio codec standard i to audio codec standard j; T max α is the maximum processing time required to convert between any two audio codec standards; β is the bit rate impact factor; γ is the latency sensitivity coefficient; δ is the compatibility weight; and δ is the conversion efficiency coefficient.
5. A multi-mode fusion gateway based on an integrated video platform according to claim 4, characterized in that, The user identity association module, when converting audio streams between different codec standards in real time based on the audio compatibility mapping relationship between terminals, and combining a multimodal data association index with a multimodal fusion algorithm to analyze the comprehensive characteristics of audio data, video data, and auxiliary data to establish an association model between user identity and multimodal data streams, includes: Based on the optimal audio compatibility mapping relationship between terminals, adaptive audio transcoding is used to generate cross-platform audio streams that are converted in real time. Based on the multimodal data association index, a multimodal feature extraction algorithm is used to obtain the comprehensive feature vector of audio data, video data and auxiliary data, and a preliminary user identity feature model is generated through a deep neural network. Based on a preliminary user identity feature model and a real-time converted cross-platform audio stream, user identity information is obtained using speaker recognition technology, and a dynamic correlation model between user identity and multi-modal data stream is established through a time-series analysis algorithm.
6. A multi-mode fusion gateway based on an integrated video platform according to claim 5, characterized in that, The expression for the adaptive audio transcoding is: R(t)=R base ×(1-ε×PLR(t))×(1+η×RTT(t)); In the formula, R(t) is the target bit rate at time t; Rbase is the base bit rate; ε is the adjustment coefficient; PLR(t) is the packet loss rate at time t; η is the delay compensation coefficient; and RTT(t) is the round-trip delay at time t.
7. A multi-mode fusion gateway based on an integrated video platform according to claim 5, characterized in that, The step of obtaining a comprehensive feature vector of audio data, video data, and auxiliary data based on a multimodal data association index and using a multimodal feature extraction algorithm, and generating a preliminary user identity feature model through a deep neural network includes: Based on the audio data in the multi-modal data association index, the audio feature vector is obtained through the Mel frequency cepstral coefficient algorithm; Based on the video data in the multimodal data association index, the local binary pattern histogram face recognition algorithm is used to obtain face feature descriptors, and deep face feature vectors are obtained through a pre-trained convolutional neural network. Based on auxiliary data in the multi-modal data association index, feature vectors of text data are obtained through the term frequency inverse document frequency algorithm. Combined with audio feature vectors and deep face feature vectors, a preliminary user identity feature model is generated using a deep neural network.
8. A multi-mode fusion gateway based on an integrated video platform according to claim 5, characterized in that, The method, based on a preliminary user identity feature model and a real-time converted cross-platform audio stream, utilizes speaker recognition technology to obtain user identity information and establishes a dynamic correlation model between user identity and multi-modal data streams through time-series analysis algorithms, including: Based on real-time converted cross-platform audio streams, feature extraction technology is used to obtain speaker feature vectors; Based on the speaker feature vector and the preliminary user identity feature model, the user identity information is obtained using the probabilistic linear discriminant analysis algorithm. Based on user identity information and multimodal data association index, a dynamic association model between user identity and multimodal data stream is generated using the sliding window time series analysis method.
9. A multi-mode fusion gateway based on an integrated video platform according to claim 8, characterized in that, The dynamic permission management module, based on user identity and a multi-modal data stream association model, combined with user role permissions of the access platform, utilizes semantic analysis technology and graph neural networks to construct a cross-platform permission relationship graph, and adjusts users' audio and video transmission permissions and interaction permissions across different platforms in real time through a dynamic permission allocation algorithm, including: Based on the user identity and multimodal data flow association model, combined with the user role permissions of the access platform, semantic analysis technology is used to obtain the permission feature vector of each platform; Based on the permission feature vectors of each platform, a cross-platform permission relationship graph is generated using a graph neural network algorithm; Based on a cross-platform permission relationship graph, and using a dynamic permission allocation algorithm, audio and video transmission permissions and interaction permission policies for users across different platforms are generated in real time.
10. A multi-mode fusion gateway based on an integrated video platform according to claim 9, characterized in that, The user role permissions of the access platform include host permissions, ordinary participant permissions, listen-only permissions, guest permissions, and administrator permissions; The permission feature vectors of each platform include audio permissions, video permissions, text interaction permissions, meeting control permissions, and data access permissions.
Citation Information
Patent Citations
Communication convergence data transmission system
CN111585944A
Multi-modal data processing method and device based on unified representation model, equipment and medium
CN119272229A