Intelligent conference room gateway comprehensive control method
By using intelligent conference room gateways and multimodal interaction technology, device linkage strategies are dynamically generated, solving the problems of complex operation and insufficient adaptability of traditional conference room equipment control systems, and achieving efficient device linkage and improved user satisfaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-03-13
AI Technical Summary
Traditional conference room equipment control systems are complex to operate, cannot adapt to the needs of meetings of different sizes and types, and lack multimodal data fusion and adaptive optimization mechanisms, resulting in low equipment linkage efficiency and operational errors.
By establishing digital models of multiple IoT devices and using a smart conference room gateway for联动 configuration, combined with edge computing and multimodal interaction, multi-source data is collected in real time. Incremental evolutionary algorithms and self-evolving linkage engines are used to dynamically generate device linkage strategies, forming a closed-loop control process.
It significantly improves the accuracy of user intent recognition, optimizes device linkage strategies, reduces meeting preparation time, enhances user satisfaction and device compatibility, and achieves comprehensive scene awareness and device linkage.
Smart Images

Figure CN121664639A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital information transmission technology, and in particular to a comprehensive control method for intelligent conference room gateways. Background Technology
[0002] With the development of modern information technology, especially the rapid innovation of Internet of Things (IoT) technology, enterprises and institutions have increasingly higher requirements for the functionality of conference room systems. Early intelligent conference room systems mainly relied on preset fixed linkage rules, such as "automatically turn on the projector and dim the lights when the meeting starts." However, this static strategy cannot adapt to the needs of meetings of different sizes and types, nor can it be personalized according to the behavior habits of the participants. With the development of IoT technology, networking of conference room equipment has become possible, but the lack of effective multimodal data fusion and adaptive optimization mechanisms means that equipment linkage remains at the level of "simple on / off".
[0003] Traditional conference room equipment control systems have many limitations: projectors, audio equipment, lighting, air conditioning, and other equipment operate independently, and the operation methods are complex and varied. Participants need to spend a lot of time debugging the equipment before the meeting starts, which is not only inefficient, but also prone to equipment failure and operational errors, seriously affecting the quality of the meeting.
[0004] Therefore, it is necessary to provide a new integrated control method for intelligent conference room gateways to solve the above-mentioned technical problems. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a comprehensive control method for intelligent conference room gateways. The intelligent conference room gateway integrated control method provided by this invention includes the following steps: S1. Establish multiple digital models with unique correspondences to different IoT devices in the conference room, bind the multiple digital models in the central control command library to the gateway, and connect to remote control for linkage configuration through API interface. The linkage configuration includes the preset linkage strategy inside the device. S2. Based on the preset linkage strategy, monitor the control command calls of the digital models of all IoT devices in the conference room and the control commands of the central control command library, and link the control of the conference room equipment. S3. Real-time collection of multi-source data in the conference room through edge computing devices. The multi-source data in the conference room includes user behavior data, device status data, and environmental parameter data. S4. The multimodal interaction and intent understanding center analyzes the user's multimodal input in real time, generates intent tags, and uses the intent tags as a component of the multi-source data of the conference room. The intent tags are input into the self-evolving linkage engine, and the incremental evolutionary algorithm is used to learn the collected multi-source data of the conference room online, dynamically generating and optimizing the device linkage strategy. The intent tags include intent type, user satisfaction weight and device operation priority. S5. Continuously monitor the status of IoT devices in the conference room and user feedback data, and dynamically adjust the preset linkage strategy through the edge decision engine to form a closed-loop control process.
[0006] Furthermore, the incremental evolutionary algorithm includes: S401. Initialize the population. Randomly generate an initial population. The initial population includes multiple candidate solutions. Each candidate solution represents a combination of coded IoT device linkage parameters in the conference room. S402. Calculate the fitness value for each individual in the current population, where the fitness value is calculated based on a weighted average of energy efficiency, user satisfaction, and device compatibility. S403. Dynamically adjust the population size based on the fitness value of each individual in the current population and the real-time meeting requirements; S404. Dynamically optimize algorithm parameters based on data characteristics; S405. An incremental population update strategy is adopted, combining newly generated offspring with their parents and selecting the one with the highest fitness. Individuals form a new generation of population.
[0007] Furthermore, the self-evolving linkage engine includes: The behavior graph construction module is used to construct user behavior feature vectors based on historical meeting data, including the frequency of participants speaking, the frequency of device operation, and the rate of change of environmental parameters. The adaptive strategy generation module is used to dynamically generate device linkage strategies based on an incremental evolutionary algorithm. The strategies include lighting adjustment, air conditioning temperature and projector status. The cross-meeting scenario migration module is used to intelligently migrate historical strategies from different meeting types to new scenarios, extracting common features of meeting scenarios through comparative learning methods.
[0008] Furthermore, the contrastive learning method includes: Construct positive and negative sample pairs, using different scenarios of the same meeting type as positive samples and different meeting types as negative samples; Semantic features of the meeting scene are extracted through self-supervised learning, and feature similarity is calculated using a Siamese network structure. Transfer weights are calculated based on feature similarity to transfer the feature representations of the source task to the target task.
[0009] Furthermore, through real-time parsing of user multimodal input by the multimodal interaction and intent understanding center, intent tags are generated, specifically including the following: Obtain the raw multimodal data input by the user; The raw data from different modalities are transformed into a unified embedding vector, and features are fused through a cross-modal attention mechanism; Based on the fused multimodal features, user intent is identified and structured intent labels are generated.
[0010] Furthermore, step S403, which adjusts the population size according to the needs of the real-time meeting, includes the following steps: Real-time number of people is counted by facial recognition or infrared sensors, the intensity of discussion is judged by voice emotion analysis, and the need for device linkage is extracted by the frequency of use of IoT devices. Use time series models to predict changes in meeting demand over a future period and adjust the population size in advance. Population size can be adjusted via a smart gateway to avoid cloud latency.
[0011] Furthermore, step S404, which dynamically optimizes the algorithm parameters based on data characteristics, includes the following steps: Data characteristics are categorized into user demand characteristics, equipment load characteristics, and environmental stability characteristics. Based on the feature classification results, parameter adjustment rules are formulated. When a feature value exceeds a preset threshold, parameter adjustment is triggered. Parameters are linearly adjusted based on the rate of change of real-time data feature values to achieve dynamic optimization.
[0012] Furthermore, calculating feature similarity using the Siamese network structure includes the following steps: Step 1: Construct the Siamese network structure: Use two completely symmetrical sub-networks that share the same weight parameters. Use a convolutional neural network as the image encoder and a recurrent neural network as the text encoder. Input pairs of samples (such as image and text pairs) and output two low-dimensional feature vectors. The two inputs are passed through the encoder with shared weights to obtain the two feature vectors. and ; Step 2: Define a similarity measurement function based on the feature vectors; Step 3: Define the loss function.
[0013] Furthermore, constructing positive and negative sample pairs specifically includes: Voice data: Different noise additions and speed variations of the same conference segment are used as positive samples, while other segments are used as negative samples; Video data: Positive and negative samples are constructed using similarity such as that between adjacent frames; Text data: Sentence fragments from the same meeting minutes are used as positive samples, and randomly shuffled sentences are used as negative samples.
[0014] Furthermore, the adjustment rule design is divided into adjustments based on user demand characteristics, device load characteristics, and environmental stability characteristics.
[0015] Compared with related technologies, the intelligent conference room gateway integrated control method provided by the present invention has the following beneficial effects: 1. This invention uses a multimodal interaction and intent understanding center to analyze user multimodal input in real time, generate intent tags, and input the intent tags as part of the multi-source data of the conference room into a self-evolving linkage engine. An incremental evolutionary algorithm is used to learn the collected multi-source data of the conference room online, dynamically generate and optimize the device linkage strategy, and use the Siamese network structure to calculate feature similarity, which significantly improves the accuracy of user intent recognition.
[0016] 2. This invention employs an incremental evolutionary algorithm to dynamically optimize device linkage strategies based on real-time data. It significantly improves key indicators such as user intent recognition accuracy, device linkage strategy optimization efficiency, meeting preparation time, and user satisfaction, providing a new technological path for the future development of smart conference rooms.
[0017] 3. This invention monitors device status and user feedback data in real time, dynamically adjusts linkage strategies, and continuously optimizes strategy parameters and rules based on historical data and real-time feedback. It integrates user behavior, device status, and environmental parameters to achieve comprehensive scene perception and reduce meeting preparation time. Attached Figure Description
[0018] Figure 1 A flowchart illustrating the integrated control method for an intelligent conference room gateway provided by the present invention; Figure 2 The flowchart of the incremental evolutionary algorithm provided by this invention is shown below. Figure 3 This is a structural block diagram of the self-evolving linkage engine provided by the present invention; Figure 4 A flowchart illustrating the comparative learning method provided by this invention; Figure 5 The flowchart for generating intent tags provided by this invention. Detailed Implementation
[0019] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0020] Please refer to the following: Figure 1 , Figure 2 , Figure 3 , Figure 4as well as Figure 5 ,in, Figure 1 A flowchart illustrating the integrated control method for an intelligent conference room gateway provided by the present invention; Figure 2 The flowchart of the incremental evolutionary algorithm provided by this invention is shown below. Figure 3 This is a structural block diagram of the self-evolving linkage engine provided by the present invention; Figure 4 A flowchart illustrating the comparative learning method provided by this invention; Figure 5 The flowchart for generating intent tags provided by this invention.
[0021] In the specific implementation process, such as Figure 1 As shown, the integrated control method for the smart conference room gateway includes the following steps: S1. Establish multiple digital models with unique correspondences to different IoT devices in the conference room, bind the multiple digital models in the central control command library to the gateway, and connect to remote control for linkage configuration through API interface. The linkage configuration includes the preset linkage strategy inside the device. S2. Based on the preset linkage strategy, monitor the control command calls of the digital models of all IoT devices in the conference room and the control commands of the central control command library, and link the control of the conference room equipment. S3. Real-time collection of multi-source data in the conference room through edge computing devices. The multi-source data in the conference room includes user behavior data, device status data, and environmental parameter data. S4. Through multimodal interaction and intent understanding center, the user's multimodal input is analyzed in real time to generate intent tags. The intent tags are used as a component of the multi-source data of the conference room and input into the self-evolving linkage engine. An incremental evolutionary algorithm is used to learn online from the collected multi-source data of the conference room, dynamically generate and optimize the device linkage strategy. The intent tags include intent type, user satisfaction weight and device operation priority. S5 continuously monitors the status of IoT devices in the conference room and user feedback data, and dynamically adjusts preset linkage strategies through the edge decision engine to form a closed-loop control process.
[0022] refer to Figure 2 As shown, the incremental evolutionary algorithm includes: S401. Initialize the population. Randomly generate an initial population. The initial population includes multiple candidate solutions. Each candidate solution represents a combination of coded IoT device linkage parameters in the conference room. S402. Calculate the fitness value for each individual in the current population, where the fitness value is calculated based on a weighted average of energy efficiency, user satisfaction, and device compatibility. S403. Dynamically adjust the population size based on the fitness value of each individual in the current population and the real-time meeting requirements; S404. Dynamically optimize algorithm parameters based on data characteristics; S405. An incremental population update strategy is adopted, combining newly generated offspring with their parents and selecting the one with the highest fitness. Individuals form a new generation of population.
[0023] The fitness value is calculated using the following formula: in, For energy efficiency, For user satisfaction, For device compatibility, , and The weighting coefficients and fitness values provide a quantitative evaluation standard for each device linkage strategy, serving as the basis for subsequent selection, optimization, and adjustment. In a specific implementation process, step S403, adjusting the population size according to the needs of the real-time meeting, includes the following steps: Real-time number of people is counted by facial recognition or infrared sensors, the intensity of discussion is judged by voice emotion analysis, and the need for device linkage is extracted by the frequency of use of IoT devices. Use time series models to predict changes in meeting demand over a future period and adjust the population size in advance. Population size can be adjusted via a smart gateway to avoid cloud latency.
[0024] It is worth noting that the population size was adjusted as follows: If the average fitness value is low and the fitness value variance is large: increase the population size (increase population diversity and enhance exploration ability). If the average fitness value is high and the variance of the fitness value is small: reduce the population size (increase the convergence speed and reduce the consumption of computational resources). If the fitness value growth rate decreases: appropriately increase the population size to avoid premature convergence.
[0025] In a specific implementation process, step S404, which dynamically optimizes the algorithm parameters based on data characteristics, includes the following: Data characteristics are categorized into user demand characteristics, equipment load characteristics, and environmental stability characteristics. Based on the feature classification results, parameter adjustment rules are formulated. When a feature value exceeds a preset threshold (e.g., user command frequency > 5 times / second), parameter adjustment is triggered. Parameters are linearly adjusted based on the rate of change of real-time data feature values to achieve dynamic optimization.
[0026] The parameter adjustment rules include: Parameter classification: Population size ( ): Controls the number of candidate solutions in the evolutionary algorithm; Fitness weights ( , , Energy efficiency User satisfaction Weight of device compatibility ; Crossover probability ( ) and the probability of mutation ( ): This affects the ability to explore and develop genetic algorithms; Simulation evaluation threshold ( ): The upper limit of response time or resource utilization of the lightweight simulation module.
[0027] It should be noted that the adjustment rules are designed in three ways: based on user demand characteristics, based on device load characteristics, and based on environmental stability characteristics. The adjustment rules are designed based on user needs, device load characteristics, and environmental stability characteristics. Based on user needs characteristics, the following rules apply: Rule 1: When user commands are concentrated on "lighting adjustment", increase the population size. +20%), increasing strategy diversity; Rule 2: When the user satisfaction weight ( When the probability of mutation is high, the mutation probability decreases. -10%), reduce randomness to stabilize the strategy.
[0028] Based on device load characteristics, the following rules apply: Rule 3: When the number of concurrent operations on a device > 5, reduce the crossover probability. -15%) to avoid resource overload; Rule 4: When resource utilization (CPU > 80%), dynamically reduce the population size. -30%), reducing computational burden.
[0029] Based on environmental stability characteristics, the following rules apply: Rule 5: When temperature and humidity fluctuations exceed 10%, increase the weight of "environmental compatibility" in the fitness assessment. +20%), prioritizing optimization of environment adaptation strategies; The rules are sorted by urgency and scope of impact.
[0030] It should be noted that the reference Figure 3 As shown, the self-evolving linkage engine includes: The behavior graph construction module is used to construct user behavior feature vectors based on historical meeting data, including the frequency of participants speaking, the frequency of device operation, and the rate of change of environmental parameters. The adaptive strategy generation module is used to dynamically generate device linkage strategies based on an incremental evolutionary algorithm. The strategies include lighting adjustment, air conditioning temperature and projector status. The cross-meeting scenario migration module is used to intelligently migrate historical strategies from different meeting types to new scenarios, extracting common features of meeting scenarios through comparative learning methods.
[0031] To further clarify, refer to Figure 4 As shown, contrastive learning methods include: Construct positive and negative sample pairs, using different scenarios of the same meeting type as positive samples and different meeting types as negative samples; Semantic features of the meeting scene are extracted through self-supervised learning, and feature similarity is calculated using a Siamese network structure. Transfer weights are calculated based on feature similarity to transfer the feature representations of the source task to the target task.
[0032] Constructing positive and negative sample pairs specifically includes: Voice data: Different noise additions and speed variations of the same conference segment are used as positive samples, while other segments are used as negative samples; Video data: Positive and negative samples are constructed using similarity such as that between adjacent frames; Text data: Sentence fragments from the same meeting minutes are used as positive samples, and randomly shuffled sentences are used as negative samples.
[0033] The semantic features extracted from the meeting scene through self-supervised learning include the following: 1. Data preprocessing: Preprocess the audio data, video data, and text data separately; Speech data: Extract MFCC time-frequency features, add noise and speed up the speech to generate positive and negative samples; Video data: Extract RGB frames and perform data enhancement such as cropping and scaling on the video; Text data: Clean meeting minutes, remove stop words and punctuation marks, and construct sentence-level or fragment-level semantic units; 2. Training tasks: Train the speech-text alignment task, the temporal prediction task, and the multimodal reconstruction task respectively; Speech-text alignment task: Input the speech signal and the corresponding text embedding vector into the CLAP model for comparison and learning; Temporal prediction task: The Transformer model can be used to predict the next sentence or action in meeting minutes to capture the temporal dynamics of meeting scenarios; Multimodal reconstruction task: Reconstructing input data using a multimodal autoencoder to integrate a joint representation of speech, video, and text; 3. Model Training Model architecture: Convolutional neural networks are used to extract speech features, deep learning models for video processing (such as TimeSformer) are used to extract video features, and BERT models are used to extract text features. Loss functions include contrastive loss, reconstruction loss, and temporal prediction loss. Contrastive loss is used for contrastive learning tasks, reconstruction loss is used for reconstruction tasks, and temporal prediction loss is used for temporal prediction tasks. Training: Training was performed using large-scale unlabeled conference data; 4. Feature Extraction After pre-training is complete, the model parameters are frozen, and the embedding vectors of the intermediate layers are extracted as semantic features. For example, the vectors of BERT are extracted as the global semantic representation of the meeting transcript.
[0034] It should be noted that the reference Figure 5 As shown, the intent labels generated by the multimodal interaction and intent understanding center, which analyzes the user's multimodal input in real time, include the following: Obtain the raw multimodal data input by the user; The raw data from different modalities are transformed into a unified embedding vector, and features are fused through a cross-modal attention mechanism; Based on the fused multimodal features, user intent is identified and structured intent labels are generated.
[0035] Calculating feature similarity using the Siamese network structure involves the following steps: Step 1: Construct the Siamese network structure: Use two completely symmetrical sub-networks that share the same weight parameters. Use a convolutional neural network as the image encoder and a recurrent neural network as the text encoder. Input pairs of samples (such as image and text pairs) and output two low-dimensional feature vectors. The two inputs are passed through the encoder with shared weights to obtain the two feature vectors. and ; Step 2: Define a similarity measurement function based on the feature vectors. European distance: ; in, The dimension index of the feature vector, i.e., the first dimension of the feature vector. One element, This represents the total number of dimensions of the feature vectors, i.e., the length of the feature vectors. Cosine similarity: ; Step 3: Define the loss function Comparative losses: in, For tags, The index of the sample pair, i.e., the first... One sample pair, The total number of training sample pairs, The distance between sample pairs, The boundary value is used; the contrastive loss aims to minimize the distance between samples of the same class and maximize the distance between samples of different classes. Triple loss: in, As anchor point, As a positive sample, For negative samples, For boundary values, anchor point Compared with positive samples The distance between them anchor point With negative samples The distance between them. According to embodiments of the present invention, a computing device that can be used to implement the above method includes a processor and a memory; The processor can be a multi-core processor or include multiple processors. In some embodiments, the processor may include a general-purpose main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a digital signal processor (DSP), etc. In some embodiments, the processor may be implemented using custom circuitry, such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).
[0036] Memory can include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM can store static data or instructions required by the processor or other modules of the computer. Permanent storage devices can be read-write storage devices. Permanent storage devices can be non-volatile storage devices that retain stored instructions and data even when the computer is powered off. In some embodiments, permanent storage devices use mass storage devices (e.g., magnetic or optical disks, flash memory) as permanent storage devices. In other embodiments, permanent storage devices can be removable storage devices (e.g., floppy disks, optical drives). System memory can be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. System memory can store some or all of the instructions and data required by the processor during operation. Furthermore, memory can include any combination of computer-readable storage media, including various types of semiconductor memory chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and disks and / or optical disks can also be used. In some implementations, the memory may include removable storage devices that are readable and / or writable, such as laser discs (CDs), read-only digital versatile optical discs (e.g., DVD-ROMs, dual-layer DVD-ROMs), read-only Blu-ray discs, ultra-high density optical discs, flash memory cards (e.g., SD cards, mini SD cards, Micro-SD cards, etc.), magnetic floppy disks, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or via wired connections.
[0037] It should be understood that, unless otherwise expressly stated herein, there is no strict order restriction on the execution of the above steps, and these steps may be executed in other orders. Moreover, at least some steps in the processes involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0038] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or basic characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects. The scope of the invention is defined by the appended claims rather than the foregoing description. Therefore, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.
[0039] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment includes only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A comprehensive control method for an intelligent conference room gateway, characterized in that, Includes the following steps: S1. Establish multiple digital models with unique correspondences to different IoT devices in the conference room, bind the multiple digital models in the central control command library to the gateway, and connect to remote control for linkage configuration through API interface. The linkage configuration includes the preset linkage strategy inside the device. S2. Based on the preset linkage strategy, monitor the control command calls of the digital models of all IoT devices in the conference room and the control commands of the central control command library, and link the control of the conference room equipment. S3. Real-time collection of multi-source data in the conference room through edge computing devices. The multi-source data in the conference room includes user behavior data, device status data, and environmental parameter data. S4. The multimodal interaction and intent understanding center analyzes the user's multimodal input in real time, generates intent tags, and uses the intent tags as a component of the multi-source data of the conference room. The intent tags are input into the self-evolving linkage engine, and the incremental evolutionary algorithm is used to learn the collected multi-source data of the conference room online, dynamically generating and optimizing the device linkage strategy. The intent tags include intent type, user satisfaction weight and device operation priority. S5. Continuously monitor the status of IoT devices in the conference room and user feedback data, and dynamically adjust the preset linkage strategy through the edge decision engine to form a closed-loop control process.
2. The intelligent conference room gateway integrated control method according to claim 1, characterized in that, The incremental evolutionary algorithm includes: S401. Initialize the population. Randomly generate an initial population. The initial population includes multiple candidate solutions. Each candidate solution represents a combination of coded IoT device linkage parameters in the conference room. S402. Calculate the fitness value for each individual in the current population, where the fitness value is calculated based on a weighted average of energy efficiency, user satisfaction, and device compatibility. S403. Dynamically adjust the population size based on the fitness value of each individual in the current population and the real-time meeting requirements; S404. Dynamically optimize algorithm parameters based on data characteristics; S405. An incremental population update strategy is adopted, combining newly generated offspring with their parents and selecting the one with the highest fitness. Individuals form a new generation of population.
3. The intelligent conference room gateway integrated control method according to claim 1, characterized in that, The self-evolving linkage engine includes: The behavior graph construction module is used to construct user behavior feature vectors based on historical meeting data, including the frequency of participants speaking, the frequency of device operation, and the rate of change of environmental parameters. The adaptive strategy generation module is used to dynamically generate device linkage strategies based on an incremental evolutionary algorithm. The strategies include lighting adjustment, air conditioning temperature and projector status. The cross-meeting scenario migration module is used to intelligently migrate historical strategies from different meeting types to new scenarios, extracting common features of meeting scenarios through comparative learning methods.
4. The intelligent conference room gateway integrated control method according to claim 3, characterized in that, The contrastive learning method includes: Construct positive and negative sample pairs, using different scenarios of the same meeting type as positive samples and different meeting types as negative samples; Semantic features of the meeting scene are extracted through self-supervised learning, and feature similarity is calculated using a Siamese network structure. Transfer weights are calculated based on feature similarity to transfer the feature representations of the source task to the target task.
5. The intelligent conference room gateway integrated control method according to claim 1, characterized in that, The process of generating intent labels by real-time parsing of user multimodal input through a multimodal interaction and intent understanding center includes the following steps: Obtain the raw multimodal data input by the user; The raw data from different modalities are transformed into a unified embedding vector, and features are fused through a cross-modal attention mechanism; Based on the fused multimodal features, user intent is identified and structured intent labels are generated.
6. The intelligent conference room gateway integrated control method according to claim 2, characterized in that, Step S403, which adjusts the population size according to the needs of the real-time meeting, includes the following steps: Real-time number of people is counted by facial recognition or infrared sensors, the intensity of discussion is judged by voice emotion analysis, and the demand for device linkage is extracted by the frequency of use of IoT devices. Use time series models to predict changes in meeting demand over a future period and adjust the population size in advance. Population size can be adjusted via a smart gateway to avoid cloud latency.
7. The intelligent conference room gateway integrated control method according to claim 2, characterized in that, Step S404, which dynamically optimizes the algorithm parameters based on data characteristics, includes the following steps: Data characteristics are categorized into user demand characteristics, equipment load characteristics, and environmental stability characteristics. Based on the feature classification results, parameter adjustment rules are formulated. When a feature value exceeds a preset threshold, parameter adjustment is triggered. Parameters are linearly adjusted based on the rate of change of real-time data feature values to achieve dynamic optimization.
8. The intelligent conference room gateway integrated control method according to claim 4, characterized in that, Calculating feature similarity using the Siamese network structure includes the following: Step 1: Construct the Siamese network structure: Use two completely symmetrical sub-networks that share the same weight parameters. Use a convolutional neural network as the image encoder and a recurrent neural network as the text encoder. Input pairs of input samples and output two low-dimensional feature vectors. The two inputs are passed through the encoder with shared weights to obtain two feature vectors. and ; Step 2: Define a similarity measurement function based on the feature vectors; Step 3: Define the loss function.
9. The intelligent conference room gateway integrated control method according to claim 4, characterized in that, Constructing positive and negative sample pairs specifically includes: Voice data: Different noise additions and speed variations of the same conference segment are used as positive samples, while other segments are used as negative samples; Video data: Positive and negative samples are constructed using the similarity of adjacent frames; Text data: Sentence fragments from the same meeting minutes are used as positive samples, and randomly shuffled sentences are used as negative samples.
10. The intelligent conference room gateway integrated control method according to claim 7, characterized in that, The adjustment rules are designed to be based on user demand characteristics, device load characteristics, and environmental stability characteristics.