A multi-scene video conference security access method for IMS switching network

By building an intelligent communication service platform and combining AI detection, IoT management, and encryption technologies, we have resolved security threats to the IMS video conferencing system, improved system security and user experience, and achieved secure audio and image recognition.

CN118921355BActive Publication Date: 2025-10-17STATE GRID GRID GANSU ELECTRIC POWER CO QINGYANG POWER SUPPLY CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411301220.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-10-17
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

IMS-based video conferencing systems face security threats such as unauthorized access, data leakage, and denial of service attacks. Especially with the support of 5G and IoT technologies, traditional security protection measures are difficult to fully cover, increasing the complexity and difficulty of system security management.

Method used

Artificial intelligence is used for intelligent threat detection, IoT device security management, and encryption technology to protect data transmission. An intelligent communication service platform is built, combining the audio security access mechanism of speech recognition STT and text-to-speech TTS, as well as the image recognition mechanism based on deep learning, to improve security and reliability.

Benefits of technology

It effectively improves the security and reliability of video conferencing systems, provides secure and flexible communication services, enhances user experience and system protection capabilities, and achieves secure recognition of audio and images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118921355B_ABST
    Figure CN118921355B_ABST
Patent Text Reader

Abstract

The application provides a video conference security access method for IMS switching network in multiple scenes, comprising the following steps: step 1, an intelligent communication service platform is constructed based on an IP multimedia subsystem (IMS) ICT and IoT fusion architecture; step 2, an IP multimedia subsystem (IMS) network audio security access mechanism based on speech recognition (STT) and text-to-speech (TTS) is established in the intelligent communication service platform; and step 3, an IP multimedia subsystem (IMS) network image recognition mechanism based on deep learning is established in the platform. The application effectively improves the accessibility, security and user experience of the video conference and voice communication system, and realizes the security recognition function of the audio. The IP multimedia subsystem (IMS) network image recognition mechanism based on deep learning can effectively and securely recognize the accessed image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of information and communication technology, Internet of Things technology and artificial intelligence, and particularly relates to a video conference security access method for IMS switching network in multiple scenarios. BACKGROUND

[0002] In recent years, with the rapid development and wide application of new generation information and communication technology (ICT), Internet of Things (IoT) and artificial intelligence (AI), the potential of IMS network has been further released and enhanced. As a typical communication application scenario supported by IMS network, video conference system has played an important role in enterprise collaboration, distance education, online medical treatment and personal socialization and other fields.

[0003] It is found in research that the video conference system based on IP multimedia subsystem (IMS) has strong network adaptability, excellent audio and video transmission quality and flexible interaction function, which meets the efficient and convenient communication needs of users. However, with the continuous expansion of application scenarios and the rapid increase of user number, the video conference system faces security challenges from multiple levels of network, software and hardware. Unauthorized access, data leakage, denial of service attack (DDoS), channel attack and other security threats occur frequently, which brings great risks to user privacy protection and data security. In addition, with the deep integration of video conference system and Internet of Things (IoT) devices, the security boundary of video conference system has been greatly expanded. Each IoT device accessing the system may become a potential entry point for security attacks, bringing new challenges to security management. The diversity of IoT devices and the increase of system access points make it difficult for traditional security protection measures to be fully covered, increasing the complexity and difficulty of system security management. Therefore, in the context of the rapid development of information technology today, effectively protecting the security access and data protection of IMS-based video conference system is not only the need of technological development, but also the urgent need of market and users. In the future, with the continuous progress of technology and the increasing complexity of security threats, the security protection measures of video conference system also need to be updated and optimized to cope with new security challenges.

[0004] The integration of artificial intelligence technology brings revolutionary improvements to IMS-based video conferencing systems. Through AI technology, video conferencing systems can achieve automated meeting content management, intelligent identification of participants, real-time monitoring and optimization of meeting quality, and automatic detection and response to security threats, significantly improving the efficiency, experience, and security of video conferencing. In addition, the popularity of 5G technology provides more stable high-speed network connections for video conferencing, making it possible to have high-definition, ultra-high-definition, or even 360-degree panoramic video conferences, greatly improving the immersion and interactivity of meetings. At the same time, IoT technology enables video conferencing systems to seamlessly connect with more smart devices and sensors, enabling environmental control, intelligent assistance, and virtual interaction, further expanding the application fields of video conferencing. IP Multimedia Subsystem is a new generation of communication switching technology standard that promotes the integration of mobile and fixed networks, providing an architectural framework for implementing IP multimedia and voice services. IMS networks can integrate various communication services, including voice, video, and SMS, and achieve quality of service (QoS) guarantees and diverse application scenario requirements. By integrating these technologies with video conferencing services based on IMS, video conferencing will exhibit more diverse application scenarios and more intelligent and efficient service models, bringing more convenience and innovative experiences to people's work and life.

[0005] Based on the above background, in order to effectively improve the security protection capability of the video conference system, maintain the normal order of the meeting and the confidentiality of the meeting information, a video conference security access method for IMS switching network in multiple scenarios is needed. SUMMARY

[0006] The purpose of the invention: With the rapid development of information and communication technology, Internet of Things and artificial intelligence, video conferencing systems based on IP Multimedia Subsystem (IMS) are widely used. However, this widespread application also poses challenges such as frequent occurrence of unauthorized access, data leakage, and denial-of-service attacks (DDoS). This invention aims to solve the security access and data protection problems faced by video conference services in multiple scenarios of IMS networks supported by 5G and IoT technologies. In order to improve the security and service efficiency of video conferencing and ensure the safe, stable and efficient operation of the video conference system, this invention proposes a video conference security access method for IMS switching network in multiple scenarios, including intelligent threat detection using AI, IoT enhanced terminal security management, and encryption technology to protect data transmission.

[0007] The application specifically provides a video conference security access method for IMS switching network in multiple scenarios.

[0008] The method comprises the following steps:

[0009] Step 1, an intelligent communication service platform is constructed based on the ICT (Information and Communication Technology, ICT) and IoT (Internet of Things, IoT) fusion architecture of the IP multimedia subsystem IMS;

[0010] Step 2, an IP multimedia subsystem IMS network audio security access mechanism based on speech recognition STT (Speech-to-Text, STT) and text-to-speech TTS (Text-to-Speech, TTS) is established in the intelligent communication service platform;

[0011] Step 3, an IP multimedia subsystem IMS network image recognition mechanism based on deep learning is established in the platform.

[0012] Step 1 comprises that the intelligent communication service platform comprises an application layer, a network layer, a device layer and a service capability layer.

[0013] The application layer is responsible for comprehensively processing the interaction, service logic and service interface of various applications of users, including IMS applications, ICT applications, Tas applications and the like.

[0014] The network layer integrates multiple network protocols, connects multiple networks such as IMS, PSTN, PLMN and IP, realizes heterogeneous and multi-network data transmission and network communication, and ensures smooth and reliable transmission of data between different levels and different protocols; the network layer utilizes IMS technology to process call control, session management, network connection and data transmission, supports and integrates multiple network protocols and communication standards;

[0015] The device layer is responsible for sensing the environment and collecting data, and provides the required information for the application layer;

[0016] The service capability layer integrates ICT and IoT technologies, provides core service logic and service capability, and includes data processing, storage, analysis and execution of business rules;

[0017] The intelligent communication service platform is used to provide speech recognition STT and text-to-speech TTS functions, and users can control Internet of Things devices through the speech recognition STT function using a voice input device, and the application layer notifies users of Internet of Things sensing information through the text-to-speech TTS function.

[0018] In step 1, the intelligent communication service platform is connected with the CHT IoT intelligent platform, and the CHT IoT intelligent platform provides a set of open application programming interfaces.

[0019] Step 2 includes the following steps:

[0020] Step 2-1, converting the voice input of the user into text information through the speech recognition STT function;

[0021] Step 2-2, converting the text information back into voice output for the user through the text-to-speech TTS function;

[0022] Step 2-3, security authentication: in combination with a biometric recognition technology, the intelligent communication service platform can recognize and verify the voiceprint of the user, ensuring that only authorized users can access the system and preventing malicious information systems from accessing the system.

[0023] Step 2-1 includes:

[0024] Step 2-1-1, pre-processing the input voice signal, including removing noise and unnecessary information and performing standardization processing; the unnecessary information refers to parts in the voice signal that interfere with the speech recognition or other processing processes. For example: background noise (such as wind noise, keyboard typing), echo and reverberation, silent parts (parts without voice content), and non-verbal sounds (such as coughing, laughing), etc. The standardization processing refers to a series of conversions on the voice signal to facilitate subsequent processing. The standardization processing includes: amplitude normalization, adjusting the amplitude of the voice signal to fluctuate within a certain range, avoiding recognition difficulties due to excessively large or small signal intensity, to facilitate subsequent digital processing; spectral standardization, converting the voice signal into a spectral representation through a mel-frequency cepstral coefficient, and performing standardization processing on the spectrum to reduce recognition difficulties caused by differences in timbre or speaking speed of different speakers. All the technologies used in the pre-processing step are in the existing field of voice processing.

[0025] Step 2-1-2, feature extraction: extracting mel-frequency cepstral coefficients and mel filter bank features from the pre-processed voice signal, and performing standardization processing on the features;

[0026] The extraction of mel-frequency cepstral coefficients and mel filter bank features from the pre-processed voice signal specifically includes:

[0027] Extracting mel-frequency cepstral coefficients:

[0028] Frame and window, the preprocessed speech signal is divided into small segments, each frame is 20-40 milliseconds, and a Hamming window is applied on each frame to reduce spectral leakage;

[0029] Fast Fourier transform (FFT), fast Fourier transform is performed on each frame signal y(t), and the frequency domain signal X(t,w) is obtained, and the formula is:

[0030]

[0031] Wherein, w(t) is the window function at t time, which is used to localize the audio signal in each small time; e is a natural constant, w is frequency; dτ represents integration, and j is the imaginary unit;

[0032] The power spectrum P(t,w) of the frame signal y(t) at t time is calculated by the following formula:

[0033] P(t,w) = |X(t,w) 2 (2)

[0034] The obtained fast Fourier transform spectrum is input into a group of mel filter banks M(m,w), and the output obtained is the mel frequency energy spectrum

[0035] The mel frequency energy spectrum is logarithmically operated, and the obtained result is recorded as S(t,m):

[0036]

[0037] Using discrete cosine transform (DCT), the logarithmic mel energy spectrum is subjected to discrete cosine transform, and the first 13 coefficients are extracted to obtain a group of mel frequency cepstral coefficients;

[0038] Extracting mel filter bank features: directly taking the mel frequency energy spectrum As the mel filter bank feature;

[0039] The standardization processing of the features includes:

[0040] Mean normalization: calculate the mean of the mel filter bank feature data set, and subtract the mean from each feature, so that the mean of the feature data is zero;

[0041] Standard deviation normalization: calculate the standard deviation of the feature data set, and divide each feature by the standard deviation, so that the standard deviation of the feature data is one;

[0042] Step 2-1-3, establish a convolutional neural network acoustic model, and map the convolutional neural network acoustic model to the voice unit of the language;

[0043] The convolutional neural network acoustic model comprises an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer.

[0044] The input layer receives the features extracted in step 2-1-2, which are input into the convolutional neural network acoustic model in the form of a two-dimensional matrix, where one dimension represents the time step and the other dimension represents the frequency feature.

[0045] The convolutional layer performs convolutional operations on the input two-dimensional matrix through convolutional kernels to extract local time-frequency domain features and generate feature maps, where different convolutional kernels can capture features in different time and frequency ranges. The operation of the lth convolutional layer is represented as:

[0046]

[0047] wherein, is the kth feature map of the lth layer, is the convolutional kernel from the ith feature map to the kth feature map; is the bias term, and f is the activation function;

[0048] The pooling layer performs down-sampling on the feature maps generated by the convolutional layer through the max-pooling operation to reduce the spatial dimension of the features while retaining the most important information, generating a two-dimensional feature map.

[0049] The fully connected layer flattens the two-dimensional feature map output by the pooling layer into a one-dimensional vector and performs further feature abstraction, finally mapping the features to higher-level language features o k , and the formula is:

[0050]

[0051] wherein, is the jth element of the flattened feature vector output by the last convolutional layer; and represent the weights and biases of the fully connected layer, respectively, and g is the activation function;

[0052] The output layer uses the Softmax function to convert the final output of the convolutional neural network acoustic model into a probability distribution, and according to the probability distribution, selects the sound unit with the highest probability as the final prediction result, and the formula is:

[0053]

[0054] Step 2-1-4, establishing a neural network language model, learning the co-occurrence probability of words based on text data; the language model is a tool for predicting the probability of word occurrence in a sentence or text, responsible for understanding the vocabulary and grammatical structure in speech. The neural network language model refers to this model built using neural networks. Neural networks are algorithmic structures that mimic the way the human brain works, capable of learning complex patterns and relationships from large amounts of data. Co-occurrence probability refers to the probability of two or more words appearing together in a given text set. The invention uses an N-gram model to learn the co-occurrence probability of words based on a large amount of text data;

[0055] Step 2-1-5, the intelligent communication service platform also includes a decoder and a semantic analysis system, the decoder uses Bayesian decision theory to optimize search efficiency, combines the output of the convolutional neural network acoustic model and the neural network language model to generate the final text output; the semantic analysis system performs semantic analysis on text data through the neural network language model to check if the text language is legal and safe.

[0056] Step 2-3 includes: when caller A calls caller B, if the call state is responsive, the intelligent communication service platform converts speech into text information for analysis, the semantic analysis system will check the context of the conversation, if the result of the conversation is illegal, the semantic analysis system will notify the called party B; otherwise, the conversation ends.

[0057] In steps 2-1 and 2-2, data in the speech recognition STT and text-to-speech TTS processing process is encrypted, including:

[0058] Data encryption transmission, when data is transmitted from user equipment to the intelligent communication service platform, the transmission link is encrypted through the Transport Layer Security (TLS) protocol to prevent data from being eavesdropped and intercepted during transmission;

[0059] Data encryption storage, user voice and text data are encrypted and stored using the Advanced Encryption Standard (AES) to ensure data confidentiality;

[0060] End-to-end encryption, end-to-end encryption (E2EE) technology is used throughout the data transmission process to ensure that the entire process from user input to final output is carried out in an encrypted environment, only the communication parties can decrypt;

[0061] Data integrity verification, using Hash-based Message Authentication Code (HMAC) to verify the integrity of the data, to ensure that the data has not been tampered with during transmission;

[0062] These encryption technologies collectively provide comprehensive security for IP Multimedia Subsystem (IMS) systems based on speech recognition (STT) and text-to-speech (TTS).

[0063] Step 3 includes:

[0064] Step 3-1, image capture: capture image data containing legal information and illegal information in multimedia communication, build a local training dataset, and perform preliminary formatting processing on the dataset;

[0065] Step 3-2, data preprocessing: preprocess the image data captured in step 3-1, including size adjustment, format conversion, and normalization, to meet the input requirements of the YOLOv5 model. For size adjustment, adjust the size of the image to 640x640 pixels while maintaining the aspect ratio by padding the edges. For format conversion, convert all images to a three-channel RGB format. For normalization, convert the image data to floating-point type and divide by 255 to scale the image pixel values to the [0, 1] range. For real-time video stream data, frame extraction and buffer management are also required. For frame extraction, extract image frames from the video stream at a fixed frame rate of 5fps. For buffer management, use a circular buffer to manage frames to ensure that new frames do not overwrite frames that have not yet been processed.

[0066] Step 3-3, YOLOv5 detection: input the preprocessed image data into the trained YOLOv5 model for object recognition. The YOLOv5 model will predict objects and bounding boxes in the image in a single forward pass and classify each object.

[0067] Step 3-4, process and analyze the results output by the YOLOv5 model.

[0068] Step 3-3 includes:

[0069] The input of the YOLOv5 model uses Mosaic data enhancement and adaptive anchor box calculation, and the optimal anchor box value is adaptively calculated according to different training sets during each training. The different training sets include public data sets and constructed local training data sets. Different training sets mean that different data sets or different partitions of data sets (such as the partition of training sets and validation sets) can be used during each training to improve the generalization ability of the YOLOv5 model. The local data set refers to a data set constructed locally that is labeled for legal or illegal data, that is, the data set for training the YOLOv5 model constructed by capturing image data in step 3-1;

[0070] The focus module of the backbone network of the YOLOv5 model performs the following operations:

[0071] Get a value from every other pixel in an image to get four independent feature layers;

[0072] Four independent feature layers are stacked to convert width and height information into channel information, which increases the number of input channels by four times, from the original red, green and blue three-channel mode to 12 channels;

[0073] Perform a convolution operation on the obtained new image to obtain a two-fold downsampled feature map without information loss;

[0074] The Conv module in the YOLOv5 model includes the convolutional layer Conv2d, the normalization Batch Normalization and the activation function SiLU;

[0075] The C3 module in the YOLOv5 model is a residual structure that enables the entire network to achieve richer gradient combinations, while reducing the number of parameters and computations, and improving the training speed of the YOLOv5 model;

[0076] The SPPF (Spatial Pyramid Pooling Fast) module in the YOLOv5 model can expand the receptive field, allowing the YOLOv5 model to obtain global feature information;

[0077] The Neck in the YOLOv5 model uses a path aggregation network to fuse the deep semantic information extracted by the backbone network with the shallow detail information. It then performs upsampling through the UpSample operation, converting the W×H image into 2W×2H, where W and H represent the width and height of the image, respectively. It then performs the Concat operation, which means concatenation, which is to concatenate feature maps of the same size.

[0078] The output Prediction of the YOLOv5 model has three branches, which are used for the detection output of large, medium and small targets respectively.

[0079] Step 3-4 includes:

[0080] Step 3-4-1, Data Collection: Image capture modules integrated in various nodes of the IP Multimedia Subsystem (IMS) network are responsible for collecting incoming multimedia content. The multimedia content is subjected to preliminary screening, which combines manual and automated screening. Preliminary screening includes content type filtering, image quality detection, duplicate removal, and time window filtering. Content type filtering automatically filters out non-target content using content recognition algorithms. Image quality detection automatically detects and filters out content with substandard image quality using edge detection algorithms. Duplicate removal automatically identifies and deletes duplicate image frames. Time window filtering automatically filters data based on timestamps, processing only images within a specific time period. Relevant image data is tagged and transmitted to the central processing server in the intelligent communication service platform. The central processing server is responsible for centralized processing and management of data streams in the IMS network.

[0081] Step 3-4-2, Data Transmission: Image data is transmitted through the secure channel of the IP Multimedia Subsystem (IMS) network to the data processing center in the intelligent communication service platform. Encryption and integrity verification are performed during transmission.

[0082] Step 3-4-3, Real-time Processing: In the data processing center, the YOLOv5 model processes and analyzes real-time or near-real-time image data. "Near real-time" generally refers to a relatively short delay in data processing and response, but it is not completely "real-time." "Near real-time" processing typically involves a time delay from data collection to processing completion and result output, which can be quantified by the delay time. In the intelligent communication service platform, if the image data is transmitted to the data processing center and the YOLOv5 model can complete processing and return results within 100ms to 1 second, such processing speed is usually referred to as "near real-time."

[0083] Beneficial effects: The video conference security access method for IMS switching network multi-scenarios of the application effectively improves the security of the video conference system by using AI for intelligent threat detection, IoT for enhanced terminal security management, and encryption technology for data transmission protection, etc. strategies, provides a safe and reliable conference environment for users. The application is based on the ICT and IoT fusion architecture of the IP multimedia subsystem IMS, and constructs an intelligent communication service platform. The intelligent communication service platform integrates the communication capabilities of IMS and ICT and IoT technologies, and provides more flexible and intelligent communication services for users, which can meet the needs of various application scenarios. The IP multimedia subsystem IMS network audio security access mechanism based on speech recognition STT and text-to-speech TTS of the application effectively improves the accessibility, security and user experience of the video conference and voice communication system, and realizes the security recognition function of the audio. The IP multimedia subsystem IMS network image recognition mechanism based on deep learning of the application can effectively and safely recognize the accessed images. BRIEF DESCRIPTION OF DRAWINGS

[0084] The above and / or other aspects of the application will become more apparent by describing in detail the preferred embodiments thereof with reference to the attached drawings.

[0085] Figure 1 The system architecture diagram of the intelligent communication service platform of the application.

[0086] Figure 2 The access flowchart of the IMS network audio security access mechanism based on STT and TTS of the application.

[0087] Figure 3 The access flowchart of the malicious information system prevention system of the application.

[0088] Figure 4 The flowchart of the YOLOv5 detection of the application.

[0089] Figure 5 The structure diagram of the C3 module of the application.

[0090] Figure 6 The structure diagram of the SPPF module of the application. DETAILED DESCRIPTION

[0091] In the embodiments of the application, a video conference security access method for IMS switching network multi-scenarios is specifically provided, including a fusion network architecture based on IMS, an IMS network audio security access mechanism based on STT and TTS, and an IMS network image recognition mechanism based on deep learning.

[0092] 1. ICT and IoT fusion architecture based on IP multimedia subsystem IMS

[0093] Based on the IMS-based ICT and IoT fusion architecture, the application constructs an intelligent communication service platform. The platform integrates the communication capabilities of IMS and ICT and IoT technologies, provides more flexible and intelligent communication services for users, and supports the needs of various application scenarios.

[0094] With reference to Figure 1 The intelligent communication service platform provided by the application comprises an application layer, a service capability layer, a network layer, and a device layer. Figure 1 In the application, CS represents conference service, STT / TTS represents speech-to-text / text-to-speech service, IP represents Internet, MS represents media service, IMS represents IP multimedia subsystem, PLMN represents public land mobile network, TAS represents telephone AP service, PSTN represents public switched telephone network, and CRS represents call record service.

[0095] The application layer is responsible for processing user interaction and service logic, providing application programs and service interfaces, and realizing specific intelligent services and application functions; wherein IMS APs are application program interfaces for realizing and managing multimedia services based on IMS networks, ICT APs are application program interfaces for managing and integrating information communication technology services, and IoT APs are application program interfaces for connecting and managing Internet of Things devices.

[0096] The service capability layer integrates ICT and IoT technologies, provides core service logic and service capabilities, and provides required services through processing of application layer requests.

[0097] The network layer is responsible for data transmission and network communication, and ensures smooth and reliable transmission of data between different layers; the network layer utilizes IMS technology to process communication tasks such as call control, session management, network connection, and data transmission, and supports multiple network protocols and communication standards; wherein PSTN is a public switched telephone network, which is a circuit switching network for voice communication, and PLMN refers to a public land mobile network, which provides mobile communication support and allows voice and data communication through a cellular network.

[0098] The device layer comprises hardware devices that interact with the intelligent communication service platform, and is responsible for data collection and transmission, and provides required information for the application layer.

[0099] In order to support users to develop IMS-based IoT services, the intelligent communication service platform is connected with the CHT IoT intelligent platform. The CHT IoT intelligent platform provides a set of open application programming interfaces (APIs), called Tas APIs. With the Tas APIs, users can independently develop customized applications to meet their specific needs. Through the intelligent communication service platform, network audio security access and other business functions can be realized, and the diversified open interfaces can add various business needs.

[0100] 2. IMS network audio security access mechanism based on STT and TTS is established on the intelligent communication service platform. The IMS network audio security access mechanism based on STT and TTS can effectively improve the accessibility, security and user experience of video conference and voice communication systems. Figure 2 The access flowchart of the IMS network audio security access mechanism based on STT and TTS is shown, and the specific implementation steps of the mechanism are as follows:

[0101] Step 1: Speech recognition (STT), which converts the user's voice input into text information. Through deep learning and neural network technology, the STT system can accurately understand and transcribe the user's voice instructions or queries. The speech recognition STT process includes the following steps:

[0102] Preprocessing of input speech signal, including removing noise and unnecessary information, and standardizing processing.

[0103] Feature extraction, according to the embodiments of the present application, the Mel frequency cepstral coefficients and Mel filter bank features are extracted from the preprocessed speech signal, and the features are standardized;

[0104] Specifically, let y(t) be the speech signal, and the short-time Fourier transform (STFT) be:

[0105]

[0106] Where w(t) is a window function used to localize the audio signal in each small time period.

[0107] The power spectrum P is calculated by the modulus square of the STFT:

[0108] P(t,w) = |X(t,w) 2 (2)

[0109] Then the power spectrum P is converted to the Mel spectrum S, which is obtained by applying the Mel filter bank to the power spectrum and taking the logarithm:

[0110]

[0111] where M(m,w) is a mel filter bank used to convert frequencies to a mel scale, making the spectrum more consistent with human auditory perception.

[0112] Standardization of extracted features includes: mean normalization, calculating the mean of the feature dataset, and subtracting each feature from the mean, so that the mean of the feature data is zero; standard deviation normalization, calculating the standard deviation of the feature dataset, and dividing each feature by the standard deviation, so that the standard deviation of the feature data is one.

[0113] The convolutional neural network acoustic model is established, and the Faster-RCNN algorithm is used to establish the acoustic model, which is used to understand the extracted audio features and map them to the sound units of language.

[0114] For a specific convolutional layer, its operation can be represented by the following formula:

[0115]

[0116] where, is the kth feature map of the lth layer, is the convolution kernel from the ith feature map to the kth feature map. is the bias term, and f is the activation function, such as ReLU;

[0117] The feature map is further processed by the fully connected layer to obtain the final classification result:

[0118]

[0119] where, is the jth element of the flattened feature vector output by the last convolutional layer. and represent the weights and biases of the fully connected layer, respectively, and g is the activation function, usually softmax function in the output layer for multi-class classification:

[0120]

[0121] The neural network language model is established, and the N-gram model is used to learn the co-occurrence probability of words based on a large amount of text data;

[0122] The decoder uses Bayesian decision theory to optimize search efficiency, combining the outputs of the acoustic model and the language model to generate the final text output.

[0123] Step 2: Text-to-Speech (TTS), convert text information back to speech output to the user. This allows the system to interact with the user in the form of voice, providing feedback or executing commands.

[0124] Step 3: Security authentication, combined with biometric technology such as voiceprint recognition, the system can identify and verify the user's voiceprint to realize the user's identity, ensuring that only authorized users can access the system. Combined with password, biometric technology, electronic token and other authentication methods, further improve the security of IMS system.

[0125] Referring to Figure 3 , the specific access process of the malicious information system is as follows:

[0126] When caller A calls the callee B. If the call state is responsive, the STT server converts the voice into text information for analysis;

[0127] The semantic analysis system will check the context of the conversation. If the result of the conversation is illegal, the semantic analysis system will notify the callee B; otherwise, the conversation ends.

[0128] Step 4: Encrypted communication, to ensure the security of data transmission in the process of STT and TTS processing, encrypt the data to prevent data from being intercepted or tampered with during transmission.

[0129] 3. IMS network image recognition mechanism based on deep learning

[0130] The IMS network image recognition system based on YOLO includes the following steps:

[0131] Step 1: Capture images

[0132] Capture image data containing legal and illegal information in the IMS network, construct a local training data set, and perform preliminary formatting processing on the data set to facilitate subsequent image recognition analysis.

[0133] Step 2: Data preprocessing

[0134] Preprocessing steps for captured image data, including size adjustment, format conversion, normalization, etc., to meet the input requirements of the YOLO model. For real-time video stream data, frame extraction and buffer management are also required.

[0135] For size adjustment, the size of the image is adjusted to 640x640 pixels, and the aspect ratio of the image is maintained by padding the edges during adjustment;

[0136] For format conversion, convert all image formats to three-channel RGB format;

[0137] For normalization, convert the image data to floating-point type and divide by 255, so that the pixel values of the input YOLOv5 model are scaled to the range [0, 1];

[0138] For frame extraction, image frames are extracted from the video stream at a fixed frame rate, where the frame rate is set to 5 fps;

[0139] For buffer management, a circular buffer is used to manage frames, ensuring that new frames do not overwrite frames that have not yet been processed during processing;

[0140] Step 3: YOLOv5 detection

[0141] The preprocessed image data is fed into the trained YOLOv5 model for object recognition. Referring to Figure 4 , the YOLOv5 detection process includes:

[0142] The input (Input) uses Mosaic data enhancement and adaptive anchor box calculation. The best anchor box value is calculated adaptively for different training sets during each training.

[0143] The focus module of the backbone network operates as follows:

[0144] A value is obtained every other pixel from an image, resulting in four independent feature layers. These four independent feature layers are stacked to convert width and height information into channel information, increasing the number of input channels by 4 times, from the original RGB three-channel mode to 12 channels. Convolution operations are performed on the new image to obtain a two-fold down-sampled feature map without information loss. The Conv module is composed of a convolution layer (Conv2d), normalization (Batch Normalization, BN), and an activation function (SiLU). Figure 5 The structure diagram of the C3 module is shown. The C3 module is a residual structure that allows the entire network to achieve more rich gradient combinations, while reducing the amount of parameters and calculations, and improving the model training speed. Figure 6 The structure diagram of the SPPF module is shown. The SPPF module can expand the receptive field, allowing the model to obtain global feature information.

[0145] The neck (Neck) uses a path aggregation network to fuse the deep semantic information extracted by the backbone network with the shallow detail information. The UpSample operation is used for up-sampling, converting an image of size WxH to 2Wx2H. The Concat operation is used for concatenation, concatenating feature maps of the same size.

[0146] The output (Prediction) of the three branches is used for the detection and output of large, medium, and small targets, respectively.

[0147] 4. Result processing

[0148] The output of the YOLO model needs further analysis and processing. For example, triggering alarms, logging, or adjusting network configurations based on the recognition results. In addition, the system also needs to implement the visualization of the results, so that network administrators can directly observe and evaluate the image recognition results.

[0149] According to the embodiments of the present application, in the IMS network, the processing of image data stream includes the following key steps:

[0150] 1) Data collection

[0151] The image capture module integrated in each node of the IMS network is responsible for collecting incoming multimedia content. These contents are preliminarily screened, and the relevant image data is marked and transmitted to the central processing server.

[0152] 2) Data transmission

[0153] Image data is transmitted to the data processing center through the secure channel of the IMS network. It is necessary to ensure the security and integrity of the data during transmission, which may involve encryption and integrity verification mechanisms.

[0154] 3) Real-time processing

[0155] In the data processing center, the YOLO model analyzes the real-time or near real-time image data. The system needs to be optimized to support high concurrency processing, ensuring efficient operation even during network traffic peaks.

[0156] By deploying YOLO-based image recognition mechanisms in the IMS network, image content in network traffic can be monitored in real time, potential security threats, illegal content, or abnormal behavior can be identified, thereby strengthening the security protection and management control of the IMS network, and ensuring the legality and security of communication content.

[0157] The present application provides a video conference security access method for IMS switching network in multiple scenarios. There are many methods and ways to realize this technical solution. The above description is only the preferred embodiment of the present application. It should be pointed out that for ordinary technical personnel in this technical field, without departing from the principle of the present application, some improvements and refinements can be made, which should be regarded as the protection scope of the present application. The components not explicitly described in the embodiment can be realized by existing technology.

Claims

1. A method for secure access to video conferencing in multiple scenarios in an IMS switching network, characterized in that: The following steps are involved: Step 1: Building an intelligent communication service platform based on the ICT and IoT convergence architecture of the IP Multimedia Subsystem (IMS). The intelligent communication service platform includes an application layer, a network layer, a device layer, and a service capability layer. The application layer is responsible for comprehensively processing the interaction, business logic and service interface of various user applications; The network layer integrates multiple network protocols, connects IMS, PSTN, PLMN and IP networks, realizes heterogeneous and multi-network data transmission and network communication, and ensures data transmission between different layers and different protocols; the network layer uses IMS technology to handle call control, session management, network connection and data transmission, and supports and integrates network protocols and communication standards; The device layer is responsible for sensing the environment and collecting data, providing information to the application layer; The business capability layer integrates ICT and IoT technologies, providing business logic and service capabilities, including data processing, storage, analysis, and execution of business rules. The intelligent communication service platform is used to provide speech recognition (STT) and text-to-speech (TTS) functions. Users can use voice input devices to control IoT devices through the speech recognition (STT) function. The application layer notifies users of IoT sensor information through the text-to-speech (TTS) function. Connecting the intelligent communication service platform to the CHT IoT intelligent platform, which provides a set of open application programming interfaces; Step 2: establishing an IP multimedia subsystem IMS network audio security access mechanism based on speech recognition STT and text-to-speech TTS in the intelligent communication service platform; Step 2 includes the following steps: Step 2-1, converting the user's voice input into text information through the voice recognition STT function; Step 2-2, converting the text information back into speech and outputting it to the user through the text-to-speech (TTS) function; Step 2-3, security authentication: In combination with biometric technology, the intelligent communication service platform can identify and verify the user's voiceprint to ensure that only authorized users can access the system; Step 3: Establish an IP Multimedia Subsystem IMS network image recognition mechanism based on deep learning in the platform.

2. The method according to claim 1, characterized in that Step 2-1 includes: Step 2-1-1, pre-processing the input speech signal, including removing noise and unnecessary information, and performing normalization; Step 2-1-2, feature extraction: extract Mel-frequency cepstral coefficients and Mel filter bank features from the preprocessed speech signal, and standardize the features; The step of extracting Mel-frequency cepstral coefficients and Mel-filter bank features from the preprocessed speech signal specifically includes: Extract Mel-frequency cepstral coefficients: Framing and windowing: the preprocessed speech signal is divided into small segments and a Hamming window is applied to each frame to reduce spectral leakage. Perform fast Fourier transform on each frame signal y(t) to obtain the frequency domain signal X(t,w), the formula is: Where w(t) is the window function at time t, which is used to localize the audio signal in each small time segment; e is a natural constant, w is the frequency; dτ represents the integral, and j is the imaginary unit; The power spectrum P(t,w) of the frame signal y(t) at time t is calculated using the following formula: P(t,w)=|X(t,w)| 2 (2) The obtained fast Fourier transform spectrum is input into a set of Mel filter banks M(m,w), and the output is the Mel frequency energy spectrum Perform logarithmic operation on the Mel frequency energy spectrum and record the result as S(t,m): Using discrete cosine transform, the logarithmic Mel energy spectrum is subjected to discrete cosine transform to obtain a set of Mel frequency cepstrum coefficients; Extract Mel filter bank features: directly convert the Mel frequency energy spectrum as Mel filter bank features; The standardization process of the features includes: Mean normalization: Calculate the mean of the Mel filter bank feature data set and subtract the mean from each feature to make the mean of the feature data zero; Standard deviation normalization: Calculate the standard deviation of the feature data set and divide each feature by the standard deviation to make the standard deviation of the feature data equal to one; Step 2-1-3: Establish a convolutional neural network acoustic model and map the convolutional neural network acoustic model to the sound units of the language; The convolutional neural network acoustic model includes an input layer, a convolution layer, a pooling layer, a fully connected layer and an output layer; The input layer receives the features extracted in step 2-1-2, and the features are input into the convolutional neural network acoustic model in the form of a two-dimensional matrix, where one dimension represents the time step and the other dimension represents the frequency feature; The convolution layer performs a convolution operation on the input two-dimensional matrix through the convolution kernel, extracts local time-frequency domain features, and generates a feature map, where different convolution kernels can capture features in different time and frequency ranges; the operation of the first convolution layer is expressed as: in, is the kth feature map of the lth layer, It is the convolution kernel from the i-th feature map to the k-th feature map; is the bias term, f is the activation function; The pooling layer downsamples the feature map generated by the convolutional layer through the maximum pooling operation, reducing the spatial dimension of the feature while retaining the most important information to generate a two-dimensional feature map; The fully connected layer flattens the two-dimensional feature map output by the pooling layer into a one-dimensional vector, and further abstracts the features, and finally maps the features to higher-level language features. k , the formula is: in, is the jth element of the flattened feature vector output by the last convolutional layer; and They represent the weight and bias of the fully connected layer respectively, and g is the activation function; The output layer uses the Softmax function to convert the final output of the convolutional neural network acoustic model into a probability distribution. According to the probability distribution, the sound unit with the highest probability is selected as the final prediction result. The formula is: Step 2-1-4: Build a neural network language model to learn the co-occurrence probability of words based on text data; Step 2-1-5, the intelligent communication service platform also includes a decoder and a semantic analysis system. The decoder uses Bayesian decision theory to optimize search efficiency and combines the output of the convolutional neural network acoustic model and the neural network language model to generate the final text output; the semantic analysis system performs semantic analysis on the text data through the neural network language model to check whether the text language is legal and safe.

3. The method according to claim 2, characterized in that Step 2-3 includes: when caller A calls callee B, if the call status is responding, the intelligent communication service platform converts the voice into text information for analysis, and the semantic analysis system checks the context of the conversation. If the result of the conversation is illegal, the semantic analysis system notifies callee B; otherwise, the conversation ends.

4. The method according to claim 3, characterized in that In steps 2-1 and 2-2, data in the speech recognition (STT) and text-to-speech (TTS) processing processes is encrypted, including: Data encryption transmission: When data is transmitted from user devices to the intelligent communication service platform, the transmission link is encrypted through the transport layer security protocol; Data encryption storage, using advanced encryption standards to encrypt and store users' voice and text data; End-to-end encryption uses end-to-end confidentiality technology throughout the entire data transmission process, ensuring that the entire process from user input to final output is carried out in an encrypted environment and only the communicating parties can decrypt it; Data integrity verification uses hash message authentication codes to verify the integrity of data to ensure that the data has not been tampered with during transmission.

5. The method according to claim 4, characterized in that Step 3 includes: Step 3-1, capturing images: capturing image data containing legal and illegal information in multimedia communications, building a local training dataset, and performing preliminary formatting on the dataset; Step 3-2, data preprocessing: Preprocess the image data captured in step 3-1, including resizing, format conversion, and normalization to meet the input requirements of the YOLOv5 model. For resizing, resize the image to 640x640 pixels, and maintain the aspect ratio of the image by padding the edges during adjustment; for format conversion, convert the formats of all images to a three-channel RGB format; for normalization, convert the image data to a floating-point type and divide it by 255 so that the image pixel values ​​input to the YOLOv5 model are scaled to the range of [0,1]; for real-time video stream data, frame extraction and buffer management are also required; for frame extraction, extract image frames from the video stream at a fixed frame rate, where the frame rate is set to 5fps; for buffer management, use a circular buffer to manage frames to ensure that new frames do not overwrite unprocessed frames during processing; Step 3-3, YOLOv5 detection: The preprocessed image data is fed into the trained YOLOv5 model for object recognition. The YOLOv5 model predicts the objects and bounding boxes in the image and classifies each object in a single forward pass. Step 3-4: Process and analyze the results output by the YOLOv5 model.

6. The method according to claim 5, characterized in that Step 3-3 includes: The input of the YOLOv5 model uses Mosaic data augmentation and adaptive anchor box calculation. During each training, the optimal anchor box value is adaptively calculated based on different training sets; the different training sets include public datasets and constructed local training datasets. The focus module of the backbone network of the YOLOv5 model performs the following operations: Get a value from every other pixel in an image to get four independent feature layers; Four independent feature layers are stacked to convert width and height information into channel information, which increases the number of input channels by four times, from the original red, green and blue three-channel mode to 12 channels; Perform a convolution operation on the obtained new image to obtain a two-fold downsampled feature map without information loss; The Conv module in the YOLOv5 model includes the convolutional layer Conv2d, the normalization Batch Normalization and the activation function SiLU; The C3 module in the YOLOv5 model is a residual structure; The SPPF module in the YOLOv5 model can expand the receptive field, allowing the YOLOv5 model to obtain global feature information; The Neck in the YOLOv5 model uses a path aggregation network to fuse the deep semantic information extracted by the backbone network with the shallow detail information. It then performs upsampling through the UpSample operation, converting the W×H image into 2W×2H, where W and H represent the width and height of the image, respectively. It then performs the Concat operation, which means concatenation, which is to concatenate feature maps of the same size. The output Prediction of the YOLOv5 model has three branches, which are used for the detection output of large, medium and small targets respectively.

7. The method according to claim 6, characterized in that Steps 3-4 include: Step 3-4-1, Data Collection: The image capture module integrated into each node of the IP Multimedia Subsystem (IMS) network is responsible for collecting incoming multimedia content. The multimedia content undergoes preliminary screening, and the relevant image data is marked and transmitted to the central processing server in the intelligent communication service platform. The central processing server is responsible for centrally processing and managing the data flow in the IMS network; Step 3-4-2, data transmission: The image data is transmitted to the data processing center in the intelligent communication service platform through the secure channel of the IP Multimedia Subsystem (IMS) network, and encryption and integrity verification are performed during the transmission process; Step 3-4-3, real-time processing: In the data processing center, the YOLOv5 model processes and analyzes real-time or near real-time image data.

Citation Information

Patent Citations

  • Smart Internet security platform based on big data technology

    CN114896305A

  • Intelligent conference system based on Internet of Things

    CN116545790A